RSS

You Can’t Iter­ate to Trust­wor­thy AI Code Without Understanding

ronbodkin17 Aug 2026 23:00 UTC
8 points
0 comments9 min readLW link

What gives you away: how LLMs form opinions of you

Cat McGee17 Aug 2026 20:59 UTC
20 points
4 comments6 min readLW link

Eval­u­at­ing Chain-of-Thought Mon­i­tora­bil­ity is Still an Open Prob­lem: Com­ments on OpenAI’s Mon­i­tora­bil­ity Evals

17 Aug 2026 19:45 UTC
15 points
1 comment18 min readLW link

Weird Re-To­k­eniza­tion, Sym­me­tries and Com­pres­sion: Re­search Agenda

17 Aug 2026 19:43 UTC
12 points
1 comment16 min readLW link

Fork Around and Find Out Part 3: In­ter­pret­ing the knight auditor

dl2717 Aug 2026 18:11 UTC
15 points
0 comments5 min readLW link

Con­nect to your fu­ture selves

PatrickDFarley17 Aug 2026 14:57 UTC
10 points
0 comments11 min readLW link

Study Up­date: Does post-train­ing quan­ti­za­tion change welfare-rele­vant in­di­ca­tors in open-weight lan­guage mod­els?

ashesfall16 Aug 2026 21:11 UTC
7 points
0 comments7 min readLW link

Q2.5 2026 Timelines Up­date: Uplift and Revenue

16 Aug 2026 19:00 UTC
54 points
15 comments11 min readLW link
(blog.aifutures.org)

Will There Be an AI Hege­mon? A Men­tal Model for AI Power Concentration

simeon_c16 Aug 2026 18:41 UTC
21 points
6 comments3 min readLW link
(simeoncampos.substack.com)

The Dooms­day Ar­gu­ment is Rea­son­able and Mostly Points to Longevity

Josh Snider16 Aug 2026 17:14 UTC
27 points
5 comments6 min readLW link

Three thoughts on civil­i­sa­tional handoff

Cleo Nardo16 Aug 2026 17:12 UTC
56 points
2 comments2 min readLW link
(clattubato.substack.com)

Should Less Wrong add sub­ti­tles?

Chris_Leong16 Aug 2026 9:05 UTC
29 points
8 comments1 min readLW link

Are ques­tions al­lowed on LessWrong?

yatharth16 Aug 2026 7:04 UTC
13 points
7 comments2 min readLW link

Does Diffu­sionGemma do la­tent rea­son­ing?

16 Aug 2026 4:22 UTC
36 points
0 comments9 min readLW link

Kimi likes causal de­ci­sion the­ory more af­ter RL in twin pris­oner’s dilemmas

oakhu15 Aug 2026 22:31 UTC
104 points
41 comments6 min readLW link

What if Pa­ram­e­ter Up­dates were Text?

DaemonicSigil15 Aug 2026 20:06 UTC
16 points
0 comments11 min readLW link

How To Catch a Distil­led Model

15 Aug 2026 19:09 UTC
12 points
8 comments9 min readLW link

Learn­ing new facts can change LLM behaviour

Richard Juggins15 Aug 2026 13:46 UTC
33 points
2 comments14 min readLW link
(www.workingthroughai.com)

Rerun­ning AI safety pa­pers on ev­ery fron­tier re­lease would be pretty easy and valuable

15 Aug 2026 5:14 UTC
105 points
7 comments4 min readLW link
(secondlookresearch.com)

Me­taphilos­o­phy II: Em­piri­cal Flywheels

interstice15 Aug 2026 2:08 UTC
18 points
7 comments5 min readLW link
(thermontology.com)