BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//Memento EPFL//
BEGIN:VEVENT
SUMMARY:Theoretical analysis of multiple descents in neural networks throu
 gh random matrix theory
DTSTART:20220825T140000
DTEND:20220825T160000
DTSTAMP:20260916T085834Z
UID:8e4311eef47bbef1b1ac30367d2d14dd296070a6fc4d640f1ffce51f
CATEGORIES:Conferences - Seminars
DESCRIPTION:Anastasia Remizova\nEDIC candidacy exam\n\nExam president: Pro
 f. Olivier Lévêque\nThesis advisor: Prof. Nicolas Macris\nCo-examiner: D
 r. Yanina Shkel\n\nAbstract\nWhile artificial neural networks are wildly s
 uccessful\, theoretical understanding of them is quite limited. One of the
  oddities is the phenomenon of double descent: increasing model complexity
  results first in a decrease in test risk\, then in an increase due to ove
 rfitting\, and\, lastly\, in a decrease again in the over-parametrized reg
 ime. These observations lead to a line of research focusing on exploring d
 ouble descent dynamics in the models approximating neural networks.\nThis 
 write-up focuses on three different works relevant to this area. The first
  one presents a range of experiments showcasing double descent in deep net
 works in various settings. The second one introduces a method to compute t
 he spectra of a large block Gaussian matrix. This result is used in the th
 ird work which analyzes the test risk by describing asymptotics of a 2-lay
 er neural network with the Neural Tangent Kernel regression and discovers 
 cases of multiple descent.\n\nBackground papers\n\n	  Nakkiran\, P.\, Kap
 lun\, G.\, Bansal\, Y.\, Yang\, T.\, Barak\, B.\, & Sutskever\, I. (2020\,
  April). Deep Double Descent: Where Bigger Models and More Data Hurt. In 
 8th International Conference on Learning Representations\,{ICLR} 2020.\n	h
 ttps://par.nsf.gov/servlets/purl/10204300 Base paper (w/o supplementary) -
  12 pages.\n	 \n	Far\, R. R.\, Oraby\, T.\, Bryc\, W.\, & Speicher\, R. S
 pectra of large block matrices https://mast.queensu.ca/~speicher/papers/bl
 ock.pdf\n	Chapters 1\, 2\, 3\, two examples from chapter 5 (e.g. 5.1 and 5
 .2)\, 6 - around 22 pages. Chapter 4 contains a proof of the theorem\, kno
 wledge of the proof is not required to understand other parts of the pap
 er.\n	 \n	Adlam\, B.\, & Pennington\, J. (2020). The Neural Tangent Kerne
 l in High Dimensions: Triple Descent and a Multi-Scale Theory of Generaliz
 ation. arXiv preprint arXiv:2008.06786.https://arxiv.org/abs/2008.06786 .
 Base paper + chapters S1\, S2\, and S3 from supplementary - 19 pages.\n\n
 \n
LOCATION:BC 129 https://plan.epfl.ch/?room==BC%20129
STATUS:CONFIRMED
END:VEVENT
END:VCALENDAR
