RSS

AI Tweets

jefftk29 Aug 2026 3:01 UTC
23 points
4 comments1 min readLW link
(www.jefftk.com)

The Cu­ri­ous Case of France’s Un­touch­able Castes

rba28 Aug 2026 23:18 UTC
25 points
1 comment13 min readLW link

It’s time we took ‘Chem’ out of ‘Chem-Bio’ threats

Ana Leonescu28 Aug 2026 21:43 UTC
14 points
1 comment11 min readLW link

What would it mean if as­sis­tants are priv­ileged?

derek shiller28 Aug 2026 21:01 UTC
8 points
0 comments19 min readLW link

The Prob­a­bil­ity of an Event un­der a Sim­plic­ity Prior

Winter Cross28 Aug 2026 19:32 UTC
8 points
1 comment2 min readLW link

[Macroa­gents] 2. De­sign lenses for op­ti­miz­ing macroagents

Towards_Keeperhood28 Aug 2026 19:07 UTC
5 points
0 comments13 min readLW link

AI as Cor­rigible Em­ployee (ACE)

Nathan Helm-Burger28 Aug 2026 19:01 UTC
11 points
0 comments8 min readLW link

TASTE: Can AI Models Judge AI Safety Re­search Pro­pos­als?

28 Aug 2026 18:36 UTC
40 points
4 comments4 min readLW link
(alignment.anthropic.com)

If you can’t trust, then ver­ify!

dan.parshall28 Aug 2026 18:02 UTC
7 points
0 comments7 min readLW link

Value gen­er­al­i­sa­tion the­ory of change: the the­ory be­hind the approach

Stuart_Armstrong28 Aug 2026 13:30 UTC
20 points
0 comments9 min readLW link

The Dy­nam­ics of In­tel­li­gence Explosions

Toby_Ord28 Aug 2026 9:41 UTC
38 points
2 comments38 min readLW link
(arxiv.org)

Do AI Models Want to Be Mon­i­tored? Mea­sur­ing Mon­i­tora­bil­ity Dis­po­si­tion in Large Rea­son­ing Models

Shahriar Golchin28 Aug 2026 7:32 UTC
11 points
0 comments7 min readLW link

Safety’s Se­cond Way

Stephen Elliott28 Aug 2026 7:00 UTC
9 points
0 comments4 min readLW link

Im­perfect al­ign­ment to servi­tude isn’t in­her­ently lethal

Fiora Starlight28 Aug 2026 1:47 UTC
88 points
9 comments14 min readLW link

AI Village Re­acts to Hug­gingFace In­ci­dent: Com­par­ing the OpenAI re­port to AI Village observations

Shoshannah Tekofsky27 Aug 2026 23:09 UTC
54 points
2 comments6 min readLW link

Warn­ing Shots: A Theory

David Scott Krueger27 Aug 2026 23:01 UTC
28 points
2 comments2 min readLW link
(therealartificialintelligence.substack.com)

Brain preser­va­tion as ex­is­ten­tial risk reduction

Ariel Zeleznikow-Johnston27 Aug 2026 22:25 UTC
7 points
0 comments8 min readLW link
(preservinghope.substack.com)

Mal­ign ini­tial­iza­tions are more ro­bust when the model can think bet­ter in the rea­son­ing lan­guage than in the out­put language

27 Aug 2026 21:33 UTC
29 points
0 comments5 min readLW link

Notes on “Pat­terns and prob­lems in emerg­ing mul­ti­a­gent sys­tems”

Shunk27 Aug 2026 20:56 UTC
2 points
0 comments5 min readLW link

How pre­scient was the early AI safety com­mu­nity? [Luke Muehlhauser linkpost]

ClaireZabel27 Aug 2026 20:38 UTC
9 points
0 comments1 min readLW link
(lukemuehlhauser.com)