RSS

Au­tomt­ing AI Safety Re­search: Manag­ing expectations

Gerard Boxo14 Sep 2026 17:56 UTC
4 points
0 comments4 min readLW link

OpenAI Pres­i­dent Brock­man says Hug­gingFace in­ci­dent model had not been al­ign­ment-trained

Caspar Oesterheld14 Sep 2026 15:28 UTC
26 points
8 comments1 min readLW link
(www.bloomberg.com)

Threat Models for Catas­trophic Risks from De­cen­tral­ised Agent Swarms

Stephen Elliott14 Sep 2026 15:22 UTC
9 points
0 comments33 min readLW link

[Cross-post] Pal­isade Pod­cast epi­sode “How to Ac­tu­ally In­fluence AI Policy (No Law De­gree Re­quired) — with Matthew Lipka”

davekasten14 Sep 2026 13:40 UTC
27 points
0 comments55 min readLW link
(palisaderesearch.org)

What we have is not what we pre­pared for

PeacockOfJuno14 Sep 2026 13:03 UTC
4 points
1 comment5 min readLW link

Cur­rent al­ign­ment tech­niques might be in­effec­tive (and ac­tively bad) in the age of RL

Daniel Tan14 Sep 2026 9:28 UTC
92 points
1 comment6 min readLW link

There is a chan­nel to 900M weekly users. What goes in it?

Charbel-Raphaël14 Sep 2026 9:21 UTC
119 points
6 comments3 min readLW link

Watch AI ma­te­ri­als-sci­ence & bio­science abil­ities closely

Yair Halberstadt14 Sep 2026 7:10 UTC
14 points
7 comments1 min readLW link

Yet an­other mes­sage board

ErnestScribbler14 Sep 2026 7:08 UTC
−1 points
0 comments1 min readLW link

RSI will not con­tinue indefinitely

dominicq14 Sep 2026 6:43 UTC
9 points
8 comments1 min readLW link
(blog.d11r.eu)

Talk­ing points for AI doomers

hgnathan14 Sep 2026 0:46 UTC
1 point
1 comment6 min readLW link

Deployment

Nina Panickssery14 Sep 2026 0:00 UTC
66 points
1 comment3 min readLW link

Align­ment & Suc­ces­sion: The Two Bars of Alignment

L Rudolf L13 Sep 2026 20:04 UTC
61 points
7 comments14 min readLW link

Con­sider how your global gov­er­nance pro­posal is differ­ent from the EU Code of Practice

David Matolcsi13 Sep 2026 16:15 UTC
72 points
10 comments4 min readLW link

Tele­op­er­ated Humans

jefftk13 Sep 2026 14:00 UTC
71 points
11 comments6 min readLW link
(www.jefftk.com)

A helpful al­ign­ment gadget

Logan Zoellner13 Sep 2026 13:38 UTC
14 points
3 comments3 min readLW link

An­thropic and OpenAI haven’t pub­lished a plan for al­ign­ing superintelligence

Zephaniah Roe13 Sep 2026 7:15 UTC
30 points
12 comments1 min readLW link

Le­gal Max­i­mums on Con­text Windows

Julian Bradshaw13 Sep 2026 7:04 UTC
3 points
5 comments1 min readLW link

We need a ‘The Day After’ mo­ment for AI X-risk

L3moncak313 Sep 2026 4:09 UTC
16 points
3 comments6 min readLW link
(kenorland.substack.com)

The Talker Does Not Con­trol The Doer (in Cur­rent AIs)

Eliezer Yudkowsky13 Sep 2026 0:49 UTC
377 points
50 comments12 min readLW link