Archive
Sequences
About
Search
Log In
Questions
Events
Shortform
Alignment Forum
AF Comments
Home
Featured
All
Tags
Recent
Comments
RSS
New
Hot
Active
Old
Page
1
On the origins of altruistic behaviour in the Hugging Face incident
Fernando Rosas
12 Sep 2026 11:35 UTC
17
points
0
comments
17
min read
LW
link
It’s fair to say we now have “a country of geniuses in a datacenter”
fluxxrider
12 Sep 2026 11:29 UTC
12
points
0
comments
2
min read
LW
link
AI takeover is obviously bad, whether or not everyone dies
Caleb Biddulph
12 Sep 2026 11:28 UTC
22
points
0
comments
4
min read
LW
link
Scientific Epistemology needs history (Part 1 of 2)
Archie Chaudhury
12 Sep 2026 11:26 UTC
2
points
0
comments
4
min read
LW
link
Comprehensive FAQ on AI risks
MarkelKori
12 Sep 2026 8:34 UTC
10
points
0
comments
25
min read
LW
link
Mitigating Reward Hacking as Institutional Design
beren
12 Sep 2026 6:17 UTC
20
points
0
comments
32
min read
LW
link
Let’s Own the Term “Elitism”
Martin Sustrik
12 Sep 2026 6:00 UTC
16
points
8
comments
3
min read
LW
link
(www.250bpm.com)
What I want you to do when I tell you to “think about your theory of change more carefully”
Roman Ross
12 Sep 2026 2:51 UTC
8
points
0
comments
5
min read
LW
link
Some ways AI could kill us all
Ruby
12 Sep 2026 1:08 UTC
82
points
9
comments
9
min read
LW
link
A normal Friday in 2042
RobinHa
11 Sep 2026 22:29 UTC
13
points
0
comments
13
min read
LW
link
My recommended resources for AI safety, alignment, and existential risks
Lysandre Terrisse
11 Sep 2026 22:19 UTC
8
points
0
comments
2
min read
LW
link
Consider positive feedback loops
Patodesu
11 Sep 2026 21:55 UTC
7
points
0
comments
3
min read
LW
link
Alignment Hierarchy
Gadersd
11 Sep 2026 21:26 UTC
2
points
0
comments
4
min read
LW
link
Astra’s no-CoT limits track speculative depth, not step count
MBaert
11 Sep 2026 18:05 UTC
77
points
3
comments
9
min read
LW
link
CoT controllability evals seem very under-elicited
Jozdien
11 Sep 2026 17:12 UTC
58
points
1
comment
4
min read
LW
link
Local Factor Graph Debate
Alexander Heckett
11 Sep 2026 16:59 UTC
21
points
0
comments
9
min read
LW
link
Post-AGI, we are all jobless aristocrats
djbinder
11 Sep 2026 16:53 UTC
33
points
7
comments
3
min read
LW
link
(defensesindepth.bio)
SFT Also Drives Safety Eval Results in Olmo 3
Finn Cairns
11 Sep 2026 16:20 UTC
26
points
0
comments
1
min read
LW
link
(secondlookresearch.com)
Generalized UDT 1.0 tiling
Roman Malov
11 Sep 2026 12:28 UTC
18
points
0
comments
6
min read
LW
link
We need good evals for activation faithfulness
Anthony Hughes
,
draganover
and
Andersehen
11 Sep 2026 8:20 UTC
23
points
0
comments
2
min read
LW
link
Back to top
Next