RSS

Moloch Does My Hair

spookyuser15 Sep 2026 10:31 UTC
2 points
0 comments19 min readLW link

For most peo­ple “in­tel­li­gence” is not goal achievement

Pato15 Sep 2026 7:52 UTC
10 points
1 comment1 min readLW link

How we might ac­tu­ally pace the fron­tier: A pro­posal for AI com­pa­nies to do pub­lic pac­ing ex­er­cises.

Dewi Erwan15 Sep 2026 5:36 UTC
2 points
0 comments4 min readLW link
(blog.dewierwan.com)

As­tra ap­pears to perform be­lief-prop­a­ga­tion-like in­fer­ence with­out CoT

MBaert15 Sep 2026 2:48 UTC
49 points
0 comments10 min readLW link

One co­or­di­nate breaks abliter­a­tion on Gemma-3

Abhishek Mishra15 Sep 2026 2:36 UTC
7 points
0 comments14 min readLW link

How to think about LLM effort

Tao Lin15 Sep 2026 2:31 UTC
14 points
0 comments2 min readLW link

Study 3: Steer­ing welfare-rele­vant di­rec­tions moved the rep­re­sen­ta­tion, but not [de­tectably] the behavior

ashesfall15 Sep 2026 2:11 UTC
8 points
0 comments17 min readLW link

Model Weight Exfil­tra­tion Seems Overrated

Vaniver15 Sep 2026 0:01 UTC
75 points
6 comments3 min readLW link

Im­prov­ing Psy­chi­a­tric Medicine Devel­op­ment with AI

CMLKevin14 Sep 2026 23:21 UTC
5 points
0 comments3 min readLW link

Align­ment Prob­lem Redux

Mateusz Bagiński14 Sep 2026 21:53 UTC
14 points
1 comment2 min readLW link

Self Inoculation

epicurus14 Sep 2026 20:45 UTC
21 points
1 comment12 min readLW link

Au­tomt­ing AI Safety Re­search: 1 Manag­ing expectations

Gerard Boxo14 Sep 2026 17:56 UTC
4 points
0 comments4 min readLW link

Brock­man says Hug­gingFace in­ci­dent model had not been al­ign­ment-trained; ETA: roon clar­ifies that it was

Caspar Oesterheld14 Sep 2026 15:28 UTC
31 points
15 comments1 min readLW link
(www.bloomberg.com)

Threat Models for Catas­trophic Risks from De­cen­tral­ised Agent Swarms

Stephen Elliott14 Sep 2026 15:22 UTC
11 points
0 comments33 min readLW link

[Cross-post] Pal­isade Pod­cast epi­sode “How to Ac­tu­ally In­fluence AI Policy (No Law De­gree Re­quired) — with Matthew Lipka”

davekasten14 Sep 2026 13:40 UTC
27 points
0 comments55 min readLW link
(palisaderesearch.org)

What we have is not what we pre­pared for

PeacockOfJuno14 Sep 2026 13:03 UTC
6 points
1 comment5 min readLW link

Cur­rent al­ign­ment tech­niques might be in­effec­tive (and ac­tively bad) in the age of RL

Daniel Tan14 Sep 2026 9:28 UTC
110 points
3 comments6 min readLW link

There is a chan­nel to 900M weekly users. What goes in it?

Charbel-Raphaël14 Sep 2026 9:21 UTC
205 points
15 comments3 min readLW link

Watch AI ma­te­ri­als-sci­ence & bio­science abil­ities closely

Yair Halberstadt14 Sep 2026 7:10 UTC
16 points
8 comments1 min readLW link

(Hu­mor) Yet an­other mes­sage board

ErnestScribbler14 Sep 2026 7:08 UTC
−1 points
0 comments1 min readLW link