BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//Memento EPFL//
BEGIN:VEVENT
SUMMARY:On the role of Weight Decay in modern Deep Learning
DTSTART:20230905T140000
DTEND:20230905T160000
DTSTAMP:20261001T205317Z
UID:8d8f7277ae80869df69b60fc50eaaed220fb93baaff4fe9adbab8a3b
CATEGORIES:Conferences - Seminars
DESCRIPTION:Francesco D'Angelo\nEDIC candidacy exam\nExam president: Prof.
  Martin Jaggi\nThesis advisor: Prof. Nicolas Flammarion\nCo-examiner: Prof
 . Florent Krzakala\n\nAbstract\nThis manuscript presents recent insights i
 nto the\nrole of Weight Decay in modern deep learning. The experiments\nin
  the work of Zhang et al. [1] challenge the traditional\nview of Weight De
 cay as a capacity constraint and posit the\nnecessity of explicit regulari
 zation for good generalization. This\nprompts a need to explore Weight Dec
 ay’s impact on training\ndynamics and optimization. Li and Arora [2] sho
 w that under\nBatch Normalization\, there exists an equivalence between th
 e\ntrajectory in function space of SGD with Weight Decay and\nthat of SGD 
 with an exponentially increasing learning rate. This\nresult challenges th
 e conventional optimization understanding\nand highlights the confounding 
 effect that can stem from the deployment\nof normalization layers. Finally
 \, the work of Li et al. [3]\nintroduces an SDE framework in which the int
 eraction between\nlearning rate schedules\, Weight Decay and Batch Normali
 zation\nis jointly studied. Their analysis unveils the existence an intrin
 sic\nlearning rate parameter\, which controls the speed of learning\nand t
 he equilibrium distribution in function space. This defies the\nwidespread
  belief that large initial learning rates are essential for\ngood generali
 zation. We conclude the manuscript with a proposal\ndescribing the next st
 eps towards explaining the role of Weight\nDecay in present-day deep learn
 ing.\n\nBackground papers\n\n	Understanding deep learning requires rethink
 ing generalization\n	An Exponential Learning Rate Schedule for Deep Learni
 ng\n	Reconciling Modern Deep Learning with Traditional Optimization Analys
 es: The Intrinsic Learning Rate\n
LOCATION:BC 229 https://plan.epfl.ch/?room==BC%20229
STATUS:CONFIRMED
END:VEVENT
END:VCALENDAR
