RSS

On the ori­gins of al­tru­is­tic be­havi­our in the Hug­ging Face incident

Fernando Rosas12 Sep 2026 11:35 UTC
17 points
0 comments17 min readLW link

It’s fair to say we now have “a coun­try of ge­niuses in a dat­a­cen­ter”

fluxxrider12 Sep 2026 11:29 UTC
12 points
0 comments2 min readLW link

AI takeover is ob­vi­ously bad, whether or not ev­ery­one dies

Caleb Biddulph12 Sep 2026 11:28 UTC
22 points
0 comments4 min readLW link

Scien­tific Episte­mol­ogy needs his­tory (Part 1 of 2)

Archie Chaudhury12 Sep 2026 11:26 UTC
2 points
0 comments4 min readLW link

Com­pre­hen­sive FAQ on AI risks

MarkelKori12 Sep 2026 8:34 UTC
10 points
0 comments25 min readLW link

Miti­gat­ing Re­ward Hack­ing as In­sti­tu­tional Design

beren12 Sep 2026 6:17 UTC
20 points
0 comments32 min readLW link

Let’s Own the Term “Elitism”

Martin Sustrik12 Sep 2026 6:00 UTC
16 points
8 comments3 min readLW link
(www.250bpm.com)

What I want you to do when I tell you to “think about your the­ory of change more care­fully”

Roman Ross12 Sep 2026 2:51 UTC
8 points
0 comments5 min readLW link

Some ways AI could kill us all

Ruby12 Sep 2026 1:08 UTC
82 points
9 comments9 min readLW link

A nor­mal Fri­day in 2042

RobinHa11 Sep 2026 22:29 UTC
13 points
0 comments13 min readLW link

My recom­mended re­sources for AI safety, al­ign­ment, and ex­is­ten­tial risks

Lysandre Terrisse11 Sep 2026 22:19 UTC
8 points
0 comments2 min readLW link

Con­sider pos­i­tive feed­back loops

Patodesu11 Sep 2026 21:55 UTC
7 points
0 comments3 min readLW link

Align­ment Hierarchy

Gadersd11 Sep 2026 21:26 UTC
2 points
0 comments4 min readLW link

As­tra’s no-CoT limits track spec­u­la­tive depth, not step count

MBaert11 Sep 2026 18:05 UTC
77 points
3 comments9 min readLW link

CoT con­trol­la­bil­ity evals seem very un­der-elicited

Jozdien11 Sep 2026 17:12 UTC
58 points
1 comment4 min readLW link

Lo­cal Fac­tor Graph Debate

Alexander Heckett11 Sep 2026 16:59 UTC
21 points
0 comments9 min readLW link

Post-AGI, we are all jobless aristocrats

djbinder11 Sep 2026 16:53 UTC
33 points
7 comments3 min readLW link
(defensesindepth.bio)

SFT Also Drives Safety Eval Re­sults in Olmo 3

Finn Cairns11 Sep 2026 16:20 UTC
26 points
0 comments1 min readLW link
(secondlookresearch.com)

Gen­er­al­ized UDT 1.0 tiling

Roman Malov11 Sep 2026 12:28 UTC
18 points
0 comments6 min readLW link

We need good evals for ac­ti­va­tion faithfulness

11 Sep 2026 8:20 UTC
23 points
0 comments2 min readLW link