RSS

How to An­swer a Ques­tion Without An­swer­ing The Question

Kabir Kumar13 Aug 2026 17:06 UTC
12 points
0 comments1 min readLW link

How My Stu­dents Think About AI

dvd13 Aug 2026 16:56 UTC
22 points
0 comments12 min readLW link

Au­to­mated al­ign­ment runs are hard to study!

13 Aug 2026 15:04 UTC
36 points
0 comments9 min readLW link

Longter­mism Seems Like A Religion

James Brobin13 Aug 2026 14:43 UTC
−4 points
1 comment3 min readLW link

Defense Against the De­cep­tive Arts

Kabir Kumar13 Aug 2026 14:21 UTC
26 points
3 comments1 min readLW link

Some prob­lems in de­ci­sion the­ory in­cor­rectly pre­con­di­tion on policy.

Canaletto13 Aug 2026 12:04 UTC
13 points
0 comments3 min readLW link

Pat­terns and prob­lems in emerg­ing mul­ti­a­gent sys­tems (An­thropic, Fron­tier Red Team)

Julian Bradshaw13 Aug 2026 4:10 UTC
40 points
0 comments1 min readLW link
(www.anthropic.com)

The Descen­ders and the Ab­sorbers: Two Per­spec­tives on Deep Learning

larry-dial13 Aug 2026 4:02 UTC
11 points
0 comments3 min readLW link

Free will is like temperature

Optimization Process12 Aug 2026 23:32 UTC
34 points
4 comments1 min readLW link

Every re­ward-hacked policy I tested trig­gered the OPE alarm — and so did my best hon­est one

JulesRoussel0112 Aug 2026 22:51 UTC
−5 points
0 comments48 min readLW link

Un­block­ing AI’s Con­tinual Learn­ing: Hints From How Hu­mans Learn

nimakeivan12 Aug 2026 18:28 UTC
12 points
4 comments10 min readLW link

Find­ing the Seams of Perception

jimmy12 Aug 2026 17:33 UTC
15 points
0 comments2 min readLW link
(beneathpsychology.com)

In­tro­duc­ing the Con­cep­tual Rea­son­ing Index

12 Aug 2026 17:08 UTC
72 points
7 comments5 min readLW link
(alignment.anthropic.com)

One at­ten­tion head car­ries knight forks in a chess trans­former, and here’s a new toolkit that found it.

dl2712 Aug 2026 16:48 UTC
8 points
0 comments1 min readLW link
(github.com)

The Psy­chol­ogy of Cope: Ra­tion­al­ity Is Not Re­v­ersed Irrationality

Chris_Leong12 Aug 2026 12:20 UTC
25 points
11 comments3 min readLW link

De­mon Safety

Ben Pace12 Aug 2026 11:55 UTC
47 points
4 comments1 min readLW link

The Clo­sure of the In­ter­net (Re­search Linkpost)

Dean Valentine (lc)12 Aug 2026 8:22 UTC
22 points
4 comments1 min readLW link
(arctotherium.substack.com)

AI swarms are start­ing to pose in­di­rect takeover risk

12 Aug 2026 5:05 UTC
116 points
5 comments10 min readLW link

Did the al­ign­ment com­mu­nity un­der­es­ti­mate its power?

StanislavKrym12 Aug 2026 2:56 UTC
18 points
0 comments10 min readLW link

The Age of Pluribus: One Con­sul­tant for Everyone

Dorothy Gale12 Aug 2026 1:39 UTC
9 points
0 comments4 min readLW link