RSS

What Hap­pens Now? Fore­cast­ing the Fal­lout from the Hug­ging Face Incident

ChristianWilliams11 Sep 2026 19:22 UTC
10 points
0 comments10 min readLW link
(metaculus.substack.com)

As­tra’s no-CoT limits track spec­u­la­tive depth, not step count

MBaert11 Sep 2026 18:05 UTC
38 points
0 comments9 min readLW link

CoT con­trol­la­bil­ity evals seem very un­der-elicited

Jozdien11 Sep 2026 17:12 UTC
53 points
1 comment4 min readLW link

Lo­cal Fac­tor Graph Debate

Alexander Heckett11 Sep 2026 16:59 UTC
21 points
0 comments9 min readLW link

Post-AGI, we are all jobless aristocrats

djbinder11 Sep 2026 16:53 UTC
23 points
6 comments3 min readLW link
(defensesindepth.bio)

SFT Also Drives Safety Eval Re­sults in Olmo 3

Finn Cairns11 Sep 2026 16:20 UTC
21 points
0 comments1 min readLW link

Why Hug­gingFace Hasn’t Shifted my P(Doom)

Josh Snider11 Sep 2026 13:54 UTC
18 points
0 comments3 min readLW link

Gen­er­al­ized UDT 1.0 tiling

Roman Malov11 Sep 2026 12:28 UTC
18 points
0 comments6 min readLW link

Could Re­s­olu­tion Help Build Align­ment Re­search as a Dis­ci­pline?

IanWS11 Sep 2026 8:49 UTC
11 points
0 comments3 min readLW link

OpenAI-Hug­gingFace: A Re­pro­duc­tion & Les­sons for Align­ment Testing

Stewy Slocum11 Sep 2026 8:48 UTC
18 points
1 comment12 min readLW link

We need good evals for ac­ti­va­tion faithfulness

11 Sep 2026 8:20 UTC
23 points
0 comments2 min readLW link

Ques­tions for the “New En­light­en­ment” in the Age of AGI

Jordan Arel11 Sep 2026 5:49 UTC
8 points
0 comments4 min readLW link

Is func­tional welfare speak­able?

11 Sep 2026 4:59 UTC
6 points
0 comments1 min readLW link
(latentminds.org)

Vol­un­tary Grad­ual Disem­pow­er­ment in the Ju­di­ciary/​Le­gal Sys­tem

Caleb Horn11 Sep 2026 4:00 UTC
14 points
0 comments1 min readLW link

Which char­ac­ter are we eval­u­at­ing? Per­sona sta­bil­ity and AI welfare

Joshua Fonseca Rivera11 Sep 2026 3:04 UTC
13 points
0 comments5 min readLW link

Per­spec­tives in favour of im­prov­ing con­cep­tual rea­son­ing capabilities

Chi Nguyen11 Sep 2026 1:32 UTC
16 points
1 comment7 min readLW link

De­fault con­tinu­a­tion mes­sage in In­spect and Petri could be problematic

Ziqian Zhong11 Sep 2026 0:20 UTC
9 points
2 comments4 min readLW link

Subagents com­ply more

jacob_drori10 Sep 2026 23:11 UTC
26 points
0 comments3 min readLW link

As­tra is much bet­ter at rea­son­ing with filler to­kens than pre­vi­ous models

10 Sep 2026 22:21 UTC
99 points
2 comments2 min readLW link

To Thine Own AI Be Truth­ful: emer­gent mis­al­ign­ment in al­ign­ment research

lumpenspace10 Sep 2026 21:41 UTC
28 points
24 comments8 min readLW link