BEGIN:VCALENDAR
PRODID:-//planitpurple.northwestern.edu//iCalendar Event//EN
VERSION:2.0
CALSCALE:GREGORIAN
METHOD:PUBLISH
CLASS:PUBLIC
BEGIN:VTIMEZONE
TZID:America/Chicago
TZURL:http://tzurl.org/zoneinfo-outlook/America/Chicago
X-LIC-LOCATION:America/Chicago
BEGIN:DAYLIGHT
TZOFFSETFROM:-0600
TZOFFSETTO:-0500
TZNAME:CDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0500
TZOFFSETTO:-0600
TZNAME:CST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
SEQUENCE:0
DTSTART;TZID=America/Chicago:20260730T110000
DTEND;TZID=America/Chicago:20260730T120000
DTSTAMP:20260730T004708Z
SUMMARY:Sign Is All You Need - An ECE Seminar with Professor Lei Ying
UID:643564@northwestern.edu
TZID:America/Chicago
DESCRIPTION:Please join us in the Electrical and Computer Engineering Department at the Technological Institute for an hour-long seminar with Professor Lei Ying of the University of Michigan.   Abstract: Reward inference (learning a reward model from human preferences) is a critical intermediate step in Preference-based Reinforcement Learning (PbRL)\, such as Reinforcement Learning from Human Feedback (RLHF) for fine-tuning Large Language Models (LLMs). In practice\, reward inference faces fundamental challenges such as distribution shift\, reward model overfitting\, and problem misspecification. An alternative approach is direct policy optimization without reward inference\, such as Direct Preference Optimization (DPO)\, which offers a much simpler pipeline but only works in the bandit setting or for deterministic MDPs. This talk introduces new algorithms for stochastic MDPs and general preference models (link functions). The key idea is a sign-based policy perturbation approach based on zeroth-order optimization. We will discuss its applications to unknown link functions and federated RLHF.     Bio: Lei Ying is a Professor in the Electrical Engineering and Computer Science Department at the University of Michigan\, Ann Arbor. He is an IEEE Fellow and an Editor-at-Large for the IEEE/ACM Transactions on Networking. His research focuses on the interplay between complex stochastic systems and big data\, including reinforcement learning\, large-scale communication/computing systems for big-data processing\, private data marketplaces\, and large-scale graph mining.   Location: Tech L440
LOCATION:Technological Institute\, L440\, 2145 Sheridan Road\, Evanston\, IL 60208
TRANSP:OPAQUE
URL:https://planitpurple.northwestern.edu/event/643564
CREATED:20260722T050000Z
STATUS:CONFIRMED
LAST-MODIFIED:20260722T195407Z
PRIORITY:0
BEGIN:VALARM
TRIGGER:-PT10M
ACTION:DISPLAY
DESCRIPTION:Reminder
END:VALARM
END:VEVENT
END:VCALENDAR