LawrenceC comments on Multi-Component Learning and S-Curves

LawrenceC 1 Dec 2022 20:57 UTC
LW: 1 AF: 1
0
AF
Well, I’d keep everything in log space and do the whole thing with log_sum_exp for numerical stability, but yeah.

EDIT: e.g. something like:
import torch.nn.functional as F
def cross_entropy_loss(Z, C):
return -torch.sum(F.log_softmax(Z) * C)
- Adam Jermyn 1 Dec 2022 22:20 UTC
  LW: 1 AF: 1
  0
  AF Parent
  Erm do C and Z have to be valid normalized probabilities for this to work?
  - LawrenceC 2 Dec 2022 7:17 UTC
    LW: 1 AF: 1
    0
    AF Parent
    C needs to be probabilities, yeah. Z can be any vector of numbers. (You can convert C into probabilities with softmax)
    - Adam Jermyn 2 Dec 2022 21:08 UTC
      LW: 1 AF: 1
      0
      AF Parent
      So indeed with cross-entropy loss I see two plateaus! Here’s rank 2:
      (note that I’ve offset the loss to so that equality of Z and C is zero loss)
      I have trouble getting rank 10 to find the zero-loss solution:
      But the phenomenology at full rank is unchanged: