RSS

Com­men­tBench: Can Models Match Hu­man Com­ments on AI Safety Posts?

19 Sep 2026 14:27 UTC
13 points
0 comments8 min readLW link

The AI race is already multipolar

June Jimenez19 Sep 2026 13:51 UTC
14 points
0 comments4 min readLW link

The DT Ca­lyx space probe

DecoRayder19 Sep 2026 9:59 UTC
1 point
0 comments12 min readLW link

Learn­ings from a week in the wet lab

michaelwaves19 Sep 2026 6:55 UTC
18 points
0 comments4 min readLW link

Gem­ini had its first break­out: Google claims it is not mis­al­ign­ment?

Young Jae Koh19 Sep 2026 2:34 UTC
9 points
0 comments1 min readLW link

Koop­man The­ory and Metaethics

bruberu19 Sep 2026 2:28 UTC
10 points
4 comments15 min readLW link

My Cur­rent Model of What Hap­pened to Elon Musk

KatSpartz18 Sep 2026 23:15 UTC
10 points
8 comments5 min readLW link

Pre­train­ing data, not ver­ifi­a­bil­ity, is why LLMs are es­pe­cially good at math (and cod­ing)

Steven Byrnes18 Sep 2026 17:08 UTC
106 points
39 comments2 min readLW link

Stop­gap Mea­sures to Ad­dress Im­me­di­ate AI Se­cu­rity Threats

18 Sep 2026 16:52 UTC
32 points
1 comment12 min readLW link
(blog.controlai.org)

Per­sua­sion Un­der­min­ing Con­trol: Can AI Talk its Way Out of Hu­man Con­trol?

18 Sep 2026 16:34 UTC
11 points
0 comments7 min readLW link
(www.far.ai)

The J-Space De­bate, Agent Swarms, and Pac­ing Fron­tier AI—Digi­tal Minds Newslet­ter #4

18 Sep 2026 16:09 UTC
9 points
0 comments37 min readLW link
(digitalminds.substack.com)

My Reflec­tions Towards the Path to Greatness

Jv Thunder18 Sep 2026 15:33 UTC
4 points
1 comment5 min readLW link

An­nounc­ing For­mal Ver­ifi­ca­tion at RESI (The In­sti­tute for Re­spon­si­ble Su­per­in­tel­li­gence)

Adam Chlipala18 Sep 2026 14:32 UTC
13 points
0 comments6 min readLW link

Col­lec­tive Epistemics: Nap­kin Math on In­de­pen­dent Errors

Jonas Hallgren18 Sep 2026 13:42 UTC
26 points
1 comment7 min readLW link

The Game is Set for a Tar­geted Memetic At­tack on the AI Safety Community

keltan18 Sep 2026 5:43 UTC
112 points
16 comments1 min readLW link

The Horse

Character#273618 Sep 2026 2:52 UTC
58 points
3 comments3 min readLW link

Deep re­cur­rent mod­els are less ro­bustly CoT-mon­i­torable than nor­mal CoT mod­els in a toy setting

18 Sep 2026 2:47 UTC
85 points
1 comment11 min readLW link

Ma­chine in­tel­li­gence and the death of hu­man expression

Girard Dorney18 Sep 2026 2:10 UTC
6 points
0 comments6 min readLW link
(extinctiondesk.substack.com)

The Cost of Utopias (a Dia­log)

WillPetillo18 Sep 2026 1:57 UTC
16 points
2 comments17 min readLW link

Two Axes of Align­ment: A Frame­work for Ro­bust Su­per­in­tel­li­gence Alignment

Arihant Gadgade18 Sep 2026 0:59 UTC
7 points
0 comments4 min readLW link