Archive
Sequences
About
Search
Log In
Questions
Events
Shortform
Alignment Forum
AF Comments
Home
Featured
All
Tags
Recent
Comments
RSS
New
Hot
Active
Old
Page
1
Moloch Does My Hair
spookyuser
15 Sep 2026 10:31 UTC
2
points
0
comments
19
min read
LW
link
For most people “intelligence” is not goal achievement
Pato
15 Sep 2026 7:52 UTC
10
points
1
comment
1
min read
LW
link
How we might actually pace the frontier: A proposal for AI companies to do public pacing exercises.
Dewi Erwan
15 Sep 2026 5:36 UTC
2
points
0
comments
4
min read
LW
link
(blog.dewierwan.com)
Astra appears to perform belief-propagation-like inference without CoT
MBaert
15 Sep 2026 2:48 UTC
49
points
0
comments
10
min read
LW
link
One coordinate breaks abliteration on Gemma-3
Abhishek Mishra
15 Sep 2026 2:36 UTC
7
points
0
comments
14
min read
LW
link
How to think about LLM effort
Tao Lin
15 Sep 2026 2:31 UTC
14
points
0
comments
2
min read
LW
link
Study 3: Steering welfare-relevant directions moved the representation, but not [detectably] the behavior
ashesfall
15 Sep 2026 2:11 UTC
8
points
0
comments
17
min read
LW
link
Model Weight Exfiltration Seems Overrated
Vaniver
15 Sep 2026 0:01 UTC
75
points
6
comments
3
min read
LW
link
Improving Psychiatric Medicine Development with AI
CMLKevin
14 Sep 2026 23:21 UTC
5
points
0
comments
3
min read
LW
link
Alignment Problem Redux
Mateusz Bagiński
14 Sep 2026 21:53 UTC
14
points
1
comment
2
min read
LW
link
Self Inoculation
epicurus
14 Sep 2026 20:45 UTC
21
points
1
comment
12
min read
LW
link
Automting AI Safety Research: 1 Managing expectations
Gerard Boxo
14 Sep 2026 17:56 UTC
4
points
0
comments
4
min read
LW
link
Brockman says HuggingFace incident model had not been alignment-trained; ETA: roon clarifies that it was
Caspar Oesterheld
14 Sep 2026 15:28 UTC
31
points
15
comments
1
min read
LW
link
(www.bloomberg.com)
Threat Models for Catastrophic Risks from Decentralised Agent Swarms
Stephen Elliott
14 Sep 2026 15:22 UTC
11
points
0
comments
33
min read
LW
link
[Cross-post] Palisade Podcast episode “How to Actually Influence AI Policy (No Law Degree Required) — with Matthew Lipka”
davekasten
14 Sep 2026 13:40 UTC
27
points
0
comments
55
min read
LW
link
(palisaderesearch.org)
What we have is not what we prepared for
PeacockOfJuno
14 Sep 2026 13:03 UTC
6
points
1
comment
5
min read
LW
link
Current alignment techniques might be ineffective (and actively bad) in the age of RL
Daniel Tan
14 Sep 2026 9:28 UTC
110
points
3
comments
6
min read
LW
link
There is a channel to 900M weekly users. What goes in it?
Charbel-Raphaël
14 Sep 2026 9:21 UTC
205
points
15
comments
3
min read
LW
link
Watch AI materials-science & bioscience abilities closely
Yair Halberstadt
14 Sep 2026 7:10 UTC
16
points
8
comments
1
min read
LW
link
(Humor) Yet another message board
ErnestScribbler
14 Sep 2026 7:08 UTC
−1
points
0
comments
1
min read
LW
link
Back to top
Next