Archive
Sequences
About
Search
Log In
Questions
Events
Shortform
Alignment Forum
AF Comments
Home
Featured
All
Tags
Recent
Comments
RSS
New
Hot
Active
Old
Page
1
Have models report provable security bugs in their environment
anithite
6 Aug 2026 22:52 UTC
12
points
0
comments
4
min read
LW
link
Why do models task game?
aditya singh
,
Neel Nanda
and
Senthooran Rajamanoharan
6 Aug 2026 22:16 UTC
30
points
0
comments
23
min read
LW
link
Contra Oster on Alcohol in Pregnancy. Part 1. The pharmacokinetics of alcohol metabolism
Mvolz
6 Aug 2026 21:34 UTC
41
points
1
comment
7
min read
LW
link
User awareness in frontier models
Ziqian Zhong
and
jsteinhardt
6 Aug 2026 20:43 UTC
24
points
0
comments
12
min read
LW
link
(transluce.org)
My Private Personal Agent
danielms
6 Aug 2026 19:53 UTC
6
points
1
comment
5
min read
LW
link
Traffic Shaping for Workload Classification
Andrew Dickson
6 Aug 2026 19:49 UTC
3
points
0
comments
8
min read
LW
link
(lucidcomputing.substack.com)
You need to stop X companies to get a Y-month pause
Expertium
6 Aug 2026 19:30 UTC
26
points
2
comments
2
min read
LW
link
Three years of progress in 500 lines of code
Gerard Boxo
6 Aug 2026 18:27 UTC
43
points
0
comments
4
min read
LW
link
Model Organisms of Sandbagging in the Wild
Vladimir Ivanov
6 Aug 2026 17:25 UTC
25
points
0
comments
8
min read
LW
link
Function vectors as a model diffing tool: 17 heads repair a bad fine-tune
Aniket Ghosh
6 Aug 2026 14:33 UTC
18
points
0
comments
17
min read
LW
link
How to define P(doom) and why it matters
Christopher King
6 Aug 2026 14:31 UTC
23
points
2
comments
3
min read
LW
link
The Open Problems of the AI Alignment Field and their Cruxes
Gunnar_Zarncke
6 Aug 2026 12:38 UTC
29
points
0
comments
5
min read
LW
link
Matryoshka NLAs: training activation verbalizers to frontload reconstruction-relevant information
loops
and
ceselder
6 Aug 2026 9:55 UTC
22
points
0
comments
7
min read
LW
link
Why You Should Almost Never Use AI to Write Anything Substantive
Erich_Grunewald
6 Aug 2026 9:50 UTC
43
points
10
comments
10
min read
LW
link
(www.erichgrunewald.com)
Reward is hyperstitional information
Abhimanyu Pallavi Sudhir
6 Aug 2026 8:05 UTC
16
points
0
comments
6
min read
LW
link
The Ghost Scale
Abraham Haskins
6 Aug 2026 2:30 UTC
−7
points
0
comments
1
min read
LW
link
Five counterintuitive insights from Plan A
romeo
6 Aug 2026 1:19 UTC
26
points
0
comments
10
min read
LW
link
Alex Turner on Leaving Google DeepMind and Disagreements with Yudkowsky
Liron
5 Aug 2026 22:45 UTC
55
points
1
comment
40
min read
LW
link
Thomas Schelling’s Nobel Prize Speech: An Astonishing Sixty Years: The Legacy of Hiroshima
Nathan Young
5 Aug 2026 21:30 UTC
25
points
0
comments
19
min read
LW
link
An argument of Parfit’s reconsidered with logical decision theory
transhumanist_atom_understander
5 Aug 2026 20:47 UTC
12
points
0
comments
5
min read
LW
link
Back to top
Next