Reproducing Double Descent: Why 300 Features Beat 39 on the Same 40 Data Points
A minimal least-squares experiment reproduces double descent: test error peaks near the interpolation threshold, then drops 28x as parameters keep growing.
A minimal least-squares experiment reproduces double descent: test error peaks near the interpolation threshold, then drops 28x as parameters keep growing.
A 500-run simulation shows recursive synthetic training can shrink variance by up to 81% in 40 generations; mixing in 10% real data cuts that to under 2%.
Breaking down the paper that changed deep learning, in plain English.