BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//Memento EPFL//
BEGIN:VEVENT
SUMMARY:EPFL CIS – RIKEN AIP Seminar Series by Prof. Taiji Suzuki\, The 
 University of Tokyo and AIP-RIKEN\, Japan
DTSTART:20211020T100000
DTEND:20211020T110000
DTSTAMP:20260916T212925Z
UID:d2a93e0f46cb94997b6494ccc6f24d3db8e81d8bda2591fb50d1bbc3
CATEGORIES:Conferences - Seminars
DESCRIPTION:Taiji Suzuki\nGet your Zoom link:\nhttps://www.epfl.ch/researc
 h/domains/cis/center-for-intelligent-systems-cis/events/riken/prof-taiji-s
 uzuki/\n\nTitle: Optimization theories of neural networks with its statist
 ical perspective\n\nAbstract: In this talk\, I discuss some optimization t
 heories of deep learning and its impact on generalization ability. First\,
  I present a deep learning optimization framework based on a noisy gradien
 t descent in an infinite dimensional Hilbert space (gradient Langevin dyna
 mics)\, and show generalization error and excess risk bounds for the solut
 ion obtained by the optimization procedure. The proposed framework can dea
 l with finite and infinite width networks simultaneously unlike existing o
 ne such as neural tangent kernel and mean field analysis. It can be shown 
 that deep learning can avoid the curse of dimensionality in a teacher-stud
 ent setting\, and eventually achieve better excess risk than kernel method
 s.\nNext\, I present a particle type optimization technique of two layer n
 eural network in the mean field regime. The proposed method\, called parti
 cle dual averaging (PDA)\, generalizes the dual averaging method in a fini
 te dimensional convex optimization to the optimization over probability di
 stributions\, and is justified by quantitative global convergence theory. 
 In addition to that\, I present a stochastic dual coordinate ascent versio
 n of PDA. Unlike PDA\, it can achieve an exponential convergence in terms 
 of the number of the outer loops.\nFinally\, (if I have time\,) I will dis
 cuss the generalization error of preconditioned ridgeless regression in th
 e overparameterized regime. In particular\, I will discuss the optimal pre
 conditioner for both the bias and variance and how it depends on label noi
 se and shape of the signal.\n\nBio: Taiji Suzuki is currently an Associate
  Professor in the Department of Mathematical Informatics at the University
  of Tokyo. He also serves as the team leader of “Deep learning theory”
  team in AIP-RIKEN.\nHe received his Ph.D. degree in information science a
 nd technology from the University of Tokyo in 2009. He has a broad researc
 h interest in statistical learning theory on deep learning\, kernel method
 s and sparse estimation\, and stochastic optimization for large-scale mach
 ine learning problems. He served as area chairs of premier conferences suc
 h as NeurIPS\, ICML\, ICLR\, AISTATS and a program chair of ACML.\nHe rece
 ived the Outstanding Paper Award at ICLR in 2021\, the MEXT Young Scientis
 ts’ Prize\, Outstanding Achievement Award in 2017 from the Japan Statist
 ical Society\, Outstanding Achievement Award in 2016 from the Japan Societ
 y for Industrial and Applied Mathematics\, and Best Paper Award in 2012 fr
 om IBISML.\n 
LOCATION:By Zoom https://www.epfl.ch/research/domains/cis/center-for-intel
 ligent-systems-cis/events/riken/prof-taiji-suzuki/
STATUS:CONFIRMED
END:VEVENT
END:VCALENDAR
