BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//Memento EPFL//
BEGIN:VEVENT
SUMMARY:IEM Distinguished Lecturers Seminar: Nearly Minimax Optimal Reinfo
 rcement Learning for Linear Markov Decision Processes
DTSTART:20231027T131500
DTEND:20231027T140000
DTSTAMP:20260926T151744Z
UID:a06250f90e5f85b55c388836ceaca20d6bae1b57ec59f332e3261dd0
CATEGORIES:Conferences - Seminars
DESCRIPTION:Prof. Quanquan Gu\nUniversity of California\, Los Angeles\, US
 A\nThe seminar will take place in ELA 2 and will be simultaneously broadc
 asted in the main auditorium in Neuchâtel Campus (MC A1 272). \n\nCoffee
  and cookies will be served at 13:00 before the seminar\, in front of the 
 two auditoriums. \n\nAbstract\nHow to make reinforcement learning (RL) ef
 ficient with large state and action spaces has been a central research pro
 blem in the RL community. A widely used approach is function approximation
 \, which approximates the value function in RL with a predefined function 
 class for efficient exploration and exploitation. In this talk\, I will fo
 cus on RL with linear function approximation. For episodic time-inhomogene
 ous linear Markov decision processes (linear MDPs) whose transition dynami
 c can be parameterized as a linear function of a given feature mapping\, I
  will present the first computationally efficient algorithm that achieves 
 the nearly minimax optimal regret \\tilde{O}(d\\sqrt{H^3K})\, where d is t
 he dimension of the feature mapping\, H is the planning horizon\, and K is
  the number of episodes. Our algorithm is based on a weighted linear regre
 ssion scheme with a carefully designed weight\, which depends on a novel v
 ariance estimator that (1) directly estimates the variance of the optimal 
 value function\, (2) monotonically decreases with respect to the number of
  episodes to ensure a better estimation accuracy\, and (3) uses a rare-swi
 tching policy to update the value function estimator to control the comple
 xity of the estimated value function class. Our work provides a complete a
 nswer to optimal RL with linear MDPs\, and the developed algorithm and the
 oretical tools may be of independent interest.\n\nThis is a joint work wit
 h Jiafan He\, Heyang Zhao and Dongruo Zhou.\n\n\nBio\nQuanquan Gu is an As
 sociate Professor of Computer Science at UCLA. His research is in the area
  of artificial intelligence and machine learning\, with a focus on develop
 ing and analyzing nonconvex optimization algorithms for machine learning t
 o understand large-scale\, dynamic\, complex\, and heterogeneous data and 
 building the theoretical foundations of deep learning and reinforcement le
 arning. He received his Ph.D. degree in Computer Science from the Universi
 ty of Illinois at Urbana-Champaign in 2014. He is a recipient of the Sloan
  Research Fellowship\, NSF CAREER Award\, Simons Berkeley Research Fellows
 hip among other industrial research awards.
LOCATION:ELA 2 https://plan.epfl.ch/?room==ELA%202 https://epfl.zoom.us/j/
 66144079474
STATUS:CONFIRMED
END:VEVENT
END:VCALENDAR
