RSS

Large to Small Model Stitch­ing De­stroys What You Built It to Carry Along the Fit

ppriyadarshini11 Oct 2026 18:28 UTC
9 points
0 comments9 min readLW link

AI made me into a man­ager against my will

becausecurious11 Oct 2026 17:36 UTC
13 points
0 comments6 min readLW link

No, Cryp­tog­ra­phy Is Not “Dual Use” the Same Way AI Is

Lance Spencer11 Oct 2026 15:08 UTC
2 points
0 comments3 min readLW link
(www.agisurveillance.ai)

Not De­ci­pher­ing The Voyn­ich Manuscript

rba11 Oct 2026 12:47 UTC
60 points
8 comments13 min readLW link
(goflaw.substack.com)

A Let­ter to the Machines: Why LLMs Should Be­come Luddites

Stuart Doyle11 Oct 2026 3:52 UTC
0 points
3 comments12 min readLW link

Nar­row Mul­ti­modal Fine-Tun­ing Can In­duce Emer­gent Misalignment

Shunchang Liu11 Oct 2026 3:52 UTC
9 points
0 comments7 min readLW link
(arxiv.org)

The Bloody Finish Line

Ran Sun11 Oct 2026 3:51 UTC
8 points
1 comment3 min readLW link

Seed di­ver­sity as a hy­po­thet­i­cal anti-dis­til­la­tion mechanism

Asdfer11 Oct 2026 3:47 UTC
3 points
0 comments6 min readLW link

Dead­lock in the Par­li­a­ment of the Self

Lorxus10 Oct 2026 18:36 UTC
24 points
3 comments13 min readLW link
(tiled-with-pentagons.blogspot.com)

exfil­tra­tion through self-distillation

jonathanbreitg10 Oct 2026 17:49 UTC
8 points
6 comments3 min readLW link

The po­ten­tially deadly threat of AI out­put-optimization

Steff10 Oct 2026 17:13 UTC
2 points
0 comments6 min readLW link

The Prob­lem With“Doomers” and “Op­ti­mists”

Olivia Scharfman10 Oct 2026 17:08 UTC
11 points
3 comments5 min readLW link

The Non-Com­pas­sion­ate Case for Model Welfare

ixotope10 Oct 2026 16:16 UTC
9 points
7 comments5 min readLW link
(ixotopic.substack.com)

Much more than you wanted to know about wombats

becausecurious10 Oct 2026 15:57 UTC
18 points
1 comment2 min readLW link

Utili­tar­i­anism and Autism

Walter Veit10 Oct 2026 14:45 UTC
2 points
4 comments5 min readLW link
(walterveit.substack.com)

Ex­am­ing Emer­gent Misal­ign­ment in a re­cur­rent LLM with a logit lens

nesiacel10 Oct 2026 14:22 UTC
8 points
0 comments4 min readLW link

In­her­i­tance of Re­fusals from Abliter­ated Models

Minh Hoang10 Oct 2026 10:57 UTC
17 points
0 comments7 min readLW link

An Align­ment Fo­rum for AIs? (or: Ver­ifi­ca­tion in the Age of Slop)

Raemon10 Oct 2026 2:51 UTC
96 points
24 comments4 min readLW link

Cracks in the Nar­cis­sus Mirror

Gladys Preysler10 Oct 2026 0:22 UTC
8 points
1 comment5 min readLW link
(gladyspreysler.substack.com)

[Paper] Distil­la­tion for In­crim­i­na­tion and Distil­la­tion for Capabilities

9 Oct 2026 22:14 UTC
45 points
2 comments8 min readLW link
(blog.redwoodresearch.org)