Sam Marks comments on Sam Marks’s Shortform

Sam Marks 8 Feb 2023 20:44 UTC
LW: 4 AF: 2
AF
Somewhat related to the SolidGoldMagicarp discussion, I thought some people might appreciate getting a sense of how unintuitive the geometry of token embeddings can be. Namely, it’s worth noting that the tokens whose embeddings are most cosine-similar to a random vector in embedding space tend not to look very semantically similar to each other. Some examples:
```
v_1                 v_2             v_3
--------------------------------------------------
 characterized       Columb          determines
 Stra                1900           conserv
 Ire                 sher            distinguishes
sent                 paed            emphasizes
 Shelter             000             consists
 Pil                mx               operates
stro                 female          independent
 wired               alt             operate
 Kor                GW               encompasses
 Maul                lvl             consisted
```
Here v_1, v_2, v_3, are random vectors in embedding space (drawn from $N (0, I_{d_e m b})$ ), and the columns give the 10 tokens whose embeddings are most cosine-similar to $v_{i}$ . I used GPT-2-large.
Perhaps 20% of the time, we get something like $v_{3}$ , where many of the nearest neighbors have something semantically similar among them (in this case, being present tense verbs in the 3rd person singular).
But most of the time, we get things that look like $v_{1}$ or $v_{2}$ : a hodgepodge with no obvious shared semantic content. GPT-2-large seems to agree: picking ” female” and ” alt” randomly from the $v_{2}$ column, the cosine similarity between the embeddings of these tokens is 0.06.
[Epistemic status: I haven’t thought that hard about this paragraph.] Thinking about the geometry here, I don’t think any of this should be surprising. Given a random vector $v \in R^{1280}$ , we should typically find that $v$ is ~orthogonal to all of the ~50000 token embeddings. Moreover, asking whether the nearest neighbors to $v$ should be semantically clustered seems to boil down to the following. Divide the tokens into semantic clusters $S_{1}, \dots, S_{n}$ ; then compare the distribution of intra-cluster variances ${{V a r}_{w \in S_{i}} (⟨ w, v ⟩_{cos})}_{i = 1}^{n}$ to the distribution of cosine similiarities of the cluster means ${⟨ E_{w \in S_{i}} [w], v ⟩_{cos}}_{i = 1}^{n}$ . From the perspective of cosine similarity to $v$ , we should expect these clusters to look basically randomly drawn from the full dataset $S = ⋃_{i = 1}^{n} S_{i}$ , so that each variance in the former set should be $\approx {V a r}_{w \in S} (⟨ w, v ⟩_{cos})$ . This should be greater than the mean of the latter set, implying that we should expect the nearest neighbors to $v$ to mostly be random tokens taken from different clusters, rather than a bunch of tokens taken from the same cluster. I could be badly wrong about all of this, though.
There’s a little bit of code for playing around with this here.