Nina Panickssery comments on My Criticism of Singular Learning Theory

Nina Panickssery 19 Nov 2023 22:31 UTC
LW: 21 AF: 10
5
AF
suppose we have a 500,000-degree polynomial, and that we fit this to 50,000 data points. In this case, we have 450,000 degrees of freedom, and we should by default expect to end up with a function which generalises very poorly. But when we train a neural network with 500,000 parameters on 50,000 MNIST images, we end up with a neural network that generalises well. Moreover, adding more parameters to the neural network will typically make generalisation better, whereas adding more parameters to the polynomial is likely to make generalisation worse.
Only tangentially related, but your intuition about polynomial regression is not quite correct. A large range of polynomial regression learning tasks will display double descent where adding more and more higher degree polynomials consistently improves loss past the interpolation threshold.

Examples from here:

Relevant paper
- Lucius Bushnaq 20 Nov 2023 6:41 UTC
  7 points
  6
  Parent
  IIRC this is probably the case for a broad range of non-NN models. I think the original Double Descent paper showed it for random Fourier features.
  
  My current guess is that NN architectures are just especially affected by this, due to having even more degenerate behavioral manifolds, ranging very widely from tiny to large RLCTs.
- Joar Skalse 20 Nov 2023 10:14 UTC
  LW: 2 AF: 1
  0
  AF Parent
  That’s interesting, thank you for this!
- rotatingpaguro 20 Nov 2023 0:57 UTC
  1 point
  0
  Parent
  I’m trying to get a quick intuition of this. I’ve not read the papers.
  My attempt:
  - On a compact domain, any function can be uniformly approximated by a polynomial (Weierstrass)
  - Powers explode quickly, so you need many terms to make a nice function with a power series, to correct the high powers at the edges
  - As the domain gets larger, it is more difficult to make the approximation
  So the relevant question is: how does the degree at training phase transition change with domain size, domain dimensionality, and Fourier series decay rate?
  Does this make sense?