RSS

Align­ment fine-tun­ing in­duces con­di­tional mis­al­ign­ment in Qwen2.5-7B-In­struct

Rhea Srivats21 Aug 2026 20:40 UTC
9 points
1 comment13 min readLW link

When is Un­limited Op­ti­miza­tion Catas­trophic?

Winter Cross21 Aug 2026 20:08 UTC
24 points
0 comments13 min readLW link
(arxiv.org)

Eval­u­at­ing Ex­pla­na­tions of LLM Be­hav­ior In The Wild with Coun­ter­fac­tual Experiments

21 Aug 2026 19:09 UTC
46 points
2 comments7 min readLW link

In Defense of ASI Socialism

cguth721 Aug 2026 18:40 UTC
6 points
2 comments2 min readLW link

Misal­igned AI in the Bronze Age

frmsaul21 Aug 2026 16:51 UTC
35 points
3 comments7 min readLW link

The im­posters among us: func­tion vec­tors that ace ev­ery check and do the wrong task (in search of cir­cu­lar­ity)

star2vec21 Aug 2026 15:54 UTC
7 points
0 comments10 min readLW link

When Models Iden­tify as a Swarm

julius vidal21 Aug 2026 11:49 UTC
49 points
1 comment5 min readLW link

My Neel Nanda MATS 10.0 Ap­pli­ca­tion: Study­ing Fea­ture Split­ting in SAEs via Train­ing Data Attribution

J Rosser21 Aug 2026 9:41 UTC
17 points
0 comments13 min readLW link

Creativity Beyond the Manifold

Kartikay Luthra20 Aug 2026 23:31 UTC
8 points
2 comments6 min readLW link

Llama will aban­don a cor­rect an­swer if it thinks you’re educated

Nick Merrill20 Aug 2026 20:06 UTC
25 points
6 comments2 min readLW link

If Aliens Ex­ist, We Should Ex­pect to Find Them Around Now

jehan20 Aug 2026 18:57 UTC
10 points
1 comment1 min readLW link
(www.jehanazad.com)

We Must Re­mem­ber That Our World Con­tains Hell

James Brobin20 Aug 2026 14:22 UTC
118 points
34 comments3 min readLW link

Mak­ing sense of the mis­al­ign­ment risk model in the An­thropic Risk Re­port (Au­gust 2026)

jasmine.ren20 Aug 2026 2:33 UTC
27 points
5 comments8 min readLW link

Steer­ing Role Confusion

20 Aug 2026 2:09 UTC
19 points
1 comment4 min readLW link

Judg­ing eth­i­cal the­o­ries by up­date rules, not by ac­tion rankings

yatharth19 Aug 2026 22:41 UTC
11 points
1 comment5 min readLW link

The Rogue Agent Ex­plo­sion Will Be Mostly Invisible

Steven McCulloch19 Aug 2026 18:46 UTC
82 points
16 comments16 min readLW link

RL cre­ates split personas

Jan Betley19 Aug 2026 18:23 UTC
177 points
17 comments4 min readLW link

A failed solu­tion to open-source game theory

Richard Willis19 Aug 2026 16:11 UTC
15 points
3 comments4 min readLW link

In­side the mind of a fair player cooperating

transhumanist_atom_understander19 Aug 2026 14:53 UTC
35 points
2 comments3 min readLW link

Some rea­sons al­ign­ment doesn’t gen­er­al­ise well

Lucius Bushnaq19 Aug 2026 14:18 UTC
107 points
1 comment9 min readLW link