BEGIN:VCALENDAR
PRODID:-//planitpurple.northwestern.edu//iCalendar Event//EN
VERSION:2.0
CALSCALE:GREGORIAN
METHOD:PUBLISH
CLASS:PUBLIC
BEGIN:VTIMEZONE
TZID:America/Chicago
TZURL:http://tzurl.org/zoneinfo-outlook/America/Chicago
X-LIC-LOCATION:America/Chicago
BEGIN:DAYLIGHT
TZOFFSETFROM:-0600
TZOFFSETTO:-0500
TZNAME:CDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0500
TZOFFSETTO:-0600
TZNAME:CST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
SEQUENCE:0
DTSTART;TZID=America/Chicago:20260826T190000
DTEND;TZID=America/Chicago:20260826T200000
DTSTAMP:20260827T030553Z
SUMMARY:Wenxuan Shi PhD Prospectus August 26: Understanding the Capabilities and Limits of Autonomous Cybersecurity Reasoning Systems
UID:644649@northwestern.edu
TZID:America/Chicago
DESCRIPTION:Cybersecurity Reasoning Systems (CRSs) combine large language models with tools\, program-analysis techniques\, execution feedback\, and control logic to discover and repair vulnerabilities in real software. Recent systems have demonstrated increasingly strong capabilities\, but a persistent gap remains between what these systems appear to achieve and our understanding of why they work. Aggregate success rates on fixed benchmarks reveal little about which components are responsible for these capabilities\, whether those components remain useful as foundation models improve\, or whether a system can recognize objectives that cannot be achieved. These questions are further complicated by benchmark issues such as answer leakage\, environment failures\, incomplete oracles\, and unintended solution shortcuts.  This thesis examines the capabilities and limits of autonomous cybersecurity reasoning systems\, with a particular focus on the mechanisms that give rise to those capabilities. In this talk\, I will summarize my prior work on automated vulnerability discovery and repair\, including our work on CRSs developed for the Artificial Intelligence Cyber Challenge (AIxCC). I will discuss how these systems combine language-model reasoning with fuzzing\, program analysis\, execution feedback\, and other specialized components\, and what their behavior reveals about the sources of end-to-end capability.  I will then discuss broader questions that emerge from this work: which system components remain valuable as foundation models improve\, how benchmark design affects the conclusions we draw about system capability\, and whether these systems can recognize when a task cannot be completed under specified constraints. Together\, this work aims to move beyond aggregate success rates toward a more concrete understanding of how cybersecurity reasoning systems work\, how their capabilities generalize\, and where their boundaries lie.\n\nWebcast Link: https://northwestern.zoom.us/j/6192600817
LOCATION:Online
TRANSP:OPAQUE
URL:https://planitpurple.northwestern.edu/event/644649
CREATED:20260820T050000Z
STATUS:CONFIRMED
LAST-MODIFIED:20260820T050000Z
PRIORITY:0
BEGIN:VALARM
TRIGGER:-PT10M
ACTION:DISPLAY
DESCRIPTION:Reminder
END:VALARM
END:VEVENT
END:VCALENDAR