BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//Memento EPFL//
BEGIN:VEVENT
SUMMARY:A new Approach to Gradient-based Convex and Non-convex Optimizatio
 n: Beyond Lipschitz Gradients
DTSTART:20230428T110000
DTEND:20230428T120000
DTSTAMP:20260916T125858Z
UID:5e3a0947cae26753d29c3eba8973156a268028109f3d31829462bff3
CATEGORIES:Conferences - Seminars
DESCRIPTION:ProfessorAli Jadbabaie\, Massachusetts institute of Technology
 \, USA\nAbstract\nClassical textbook analysis of convex and non-convex opt
 imization problems often  assumes that gradients are Lipschitz continuous
 \, thereby mostly restricting the analysis to objective functions bounded
  by quadratics. In a recent work[1]\, we generalized this standard smoothn
 ess to the setting where the Lipschitz constant of the gradient (or altern
 atively norm of the Hessian when it exists) is bounded by a positive\, aff
 ine function of the gradient norm. The class of such functions is far more
  general than quadratics and contains (univariate) polynomials and exponen
 tial functions as special cases. In this talk\, we revisit this generalize
 d smoothness condition and use it to  develop a simple yet powerful analy
 sis technique that allows us to obtain much stronger results for both conv
 ex and non-convex optimization. This new approach is based on carefully bo
 unding the gradient norm along the optimization trajectory. Using this\, w
 e show that in the convex setting\, gradient descent (GD) converges under 
 a very general smoothness condition: all that is needed is for  the Hessi
 an to be bounded by a continuous function of the gradient norm.  In the n
 on-convex setting\, when the local smoothness  (or Hessian norm) is a sub
 -quadratic function of the gradient norm\, we show that GD\, SGD with cons
 tant stepsize\, and Adam [a popular adaptive approach used in Deep Learnin
 g] provably converge under very general settings. For SGD\, we show that u
 nder this more general smoothness condition\,  the standard assumption of
  bounded noise in previous works can be relaxed to merely boundedness of t
 he  noise variance. In the case of Adam\, we show convergence without ass
 uming the Lipschitzness of the objective function. Finally\, we propose a 
 variance-reduced version of Adam with an accelerated convergence rate that
  is tight. For all the above-mentioned settings and algorithms\, we prove 
 convergence rates that match the state of the art without restrictive assu
 mptions. [Joint work with PhD student Haochuan Li]\n\n[1] Zhang\, J.\, He\
 , T.\, Sra\, S.\, & Jadbabaie\, A. Why Gradient Clipping Accelerates Train
 ing: A Theoretical Justification for Adaptivity. In International Confere
 nce on Learning Representations.ICLR 2020\n\n\nBiography\n\nAli Jadbabaie 
 is the JR East Professor and Head of the Department of Civil and Environme
 ntal Engineering at Massachusetts Institute of Technology (MIT)\, Cambridg
 e\, MA\, USA\, where he is also a core faculty in the Institute for Data\,
  Systems\, and Society (IDSS) and a Principal Investigator with the Labora
 tory for Information and Decision Systems. Previously\, he served as the D
 irector of the Sociotechnical Systems Research Center and as the Associate
  Director of the IDSS\, MIT\, which he helped found in 2015. He is the co-
 creator of the Social and Engineering Systems PhD program within IDSS\, an
  interdisciplinary PhD program that combines information sciences and soci
 al sciences to solve complex societal problems in different domains.\nHe r
 eceived a B.S. degree with High Honors in electrical engineering with a fo
 cus on control systems from the Sharif University of Technology\, an M.S. 
 degree in electrical and computer engineering from the University of New M
 exico\, and a Ph.D. degree in control and dynamical systems from the Calif
 ornia Institute of Technology. He was a Postdoctoral Scholar at Yale Unive
 rsity before joining the faculty at the University of Pennsylvania\, where
  he was subsequently promoted through the ranks and held the Alfred Fitler
  Moore Professorship in network science in the Department of Electrical an
 d Systems Engineering.   He is a recipient of a US National Science Found
 ation Career Development Award\, an US Office of Naval Research Young Inve
 stigator Award\, the O. Hugo Schuck Best Paper Award from the American Aut
 omatic Control Council\, and the George S. Axelby Best Paper Award from th
 e IEEE Control Systems Society. He has been a senior author of several stu
 dent best paper awards\, in several conference including ACC\, IEEE CDC an
 d IEEE ICASSP. He is an IEEE fellow\, and the recipient of a Vannevar Bush
  Fellowship from the Office of Secretary of Defense.  His research intere
 sts are broadly in systems theory\, decision theory and control\,  optimi
 zation theory\, as well as multiagent systems\, collective decision making
  in social and economic networks\, and computational social science. He ha
 s served as the inaugural Editor-in-Chief of the IEEE Transactions on netw
 ork Science and Engineering\, an interdisciplinary journal sponsored by se
 veral IEEE societies\, and as an associate editor for the Informs journal 
  Operations Research.\n 
LOCATION:ME C2 405 https://plan.epfl.ch/?room==ME%20C2%20405
STATUS:CONFIRMED
END:VEVENT
END:VCALENDAR
