RSS

Have mod­els re­port prov­able se­cu­rity bugs in their environment

anithite6 Aug 2026 22:52 UTC
12 points
0 comments4 min readLW link

Why do mod­els task game?

6 Aug 2026 22:16 UTC
30 points
0 comments23 min readLW link

Con­tra Oster on Al­co­hol in Preg­nancy. Part 1. The phar­ma­coki­net­ics of al­co­hol metabolism

Mvolz6 Aug 2026 21:34 UTC
41 points
1 comment7 min readLW link

User aware­ness in fron­tier models

6 Aug 2026 20:43 UTC
24 points
0 comments12 min readLW link
(transluce.org)

My Pri­vate Per­sonal Agent

danielms6 Aug 2026 19:53 UTC
6 points
1 comment5 min readLW link

Traf­fic Shap­ing for Work­load Classification

Andrew Dickson6 Aug 2026 19:49 UTC
3 points
0 comments8 min readLW link
(lucidcomputing.substack.com)

You need to stop X com­pa­nies to get a Y-month pause

Expertium6 Aug 2026 19:30 UTC
26 points
2 comments2 min readLW link

Three years of progress in 500 lines of code

Gerard Boxo6 Aug 2026 18:27 UTC
43 points
0 comments4 min readLW link

Model Or­ganisms of Sand­bag­ging in the Wild

Vladimir Ivanov6 Aug 2026 17:25 UTC
25 points
0 comments8 min readLW link

Func­tion vec­tors as a model diffing tool: 17 heads re­pair a bad fine-tune

Aniket Ghosh6 Aug 2026 14:33 UTC
18 points
0 comments17 min readLW link

How to define P(doom) and why it matters

Christopher King6 Aug 2026 14:31 UTC
23 points
2 comments3 min readLW link

The Open Prob­lems of the AI Align­ment Field and their Cruxes

Gunnar_Zarncke6 Aug 2026 12:38 UTC
29 points
0 comments5 min readLW link

Ma­tryoshka NLAs: train­ing ac­ti­va­tion ver­bal­iz­ers to front­load re­con­struc­tion-rele­vant information

6 Aug 2026 9:55 UTC
22 points
0 comments7 min readLW link

Why You Should Al­most Never Use AI to Write Any­thing Substantive

Erich_Grunewald6 Aug 2026 9:50 UTC
43 points
10 comments10 min readLW link
(www.erichgrunewald.com)

Re­ward is hy­per­sti­tional information

Abhimanyu Pallavi Sudhir6 Aug 2026 8:05 UTC
16 points
0 comments6 min readLW link

The Ghost Scale

Abraham Haskins6 Aug 2026 2:30 UTC
−7 points
0 comments1 min readLW link

Five coun­ter­in­tu­itive in­sights from Plan A

romeo6 Aug 2026 1:19 UTC
26 points
0 comments10 min readLW link

Alex Turner on Leav­ing Google Deep­Mind and Disagree­ments with Yudkowsky

Liron5 Aug 2026 22:45 UTC
55 points
1 comment40 min readLW link

Thomas Schel­ling’s No­bel Prize Speech: An As­ton­ish­ing Sixty Years: The Le­gacy of Hiroshima

Nathan Young5 Aug 2026 21:30 UTC
25 points
0 comments19 min readLW link

An ar­gu­ment of Parfit’s re­con­sid­ered with log­i­cal de­ci­sion theory

transhumanist_atom_understander5 Aug 2026 20:47 UTC
12 points
0 comments5 min readLW link