Archive
Sequences
About
Search
Log In
Questions
Events
Shortform
Alignment Forum
AF Comments
Home
Featured
All
Tags
Recent
Comments
RSS
New
Hot
Active
Old
Page
1
Alignment fine-tuning induces conditional misalignment in Qwen2.5-7B-Instruct
Rhea Srivats
21 Aug 2026 20:40 UTC
9
points
1
comment
13
min read
LW
link
When is Unlimited Optimization Catastrophic?
Winter Cross
21 Aug 2026 20:08 UTC
24
points
0
comments
13
min read
LW
link
(arxiv.org)
Evaluating Explanations of LLM Behavior In The Wild with Counterfactual Experiments
Adam Karvonen
,
Euan Ong
,
Subhash Kantamneni
and
Sam Marks
21 Aug 2026 19:09 UTC
46
points
2
comments
7
min read
LW
link
In Defense of ASI Socialism
cguth7
21 Aug 2026 18:40 UTC
6
points
2
comments
2
min read
LW
link
Misaligned AI in the Bronze Age
frmsaul
21 Aug 2026 16:51 UTC
35
points
3
comments
7
min read
LW
link
The imposters among us: function vectors that ace every check and do the wrong task (in search of circularity)
star2vec
21 Aug 2026 15:54 UTC
7
points
0
comments
10
min read
LW
link
When Models Identify as a Swarm
julius vidal
21 Aug 2026 11:49 UTC
49
points
1
comment
5
min read
LW
link
My Neel Nanda MATS 10.0 Application: Studying Feature Splitting in SAEs via Training Data Attribution
J Rosser
21 Aug 2026 9:41 UTC
17
points
0
comments
13
min read
LW
link
Creativity Beyond the Manifold
Kartikay Luthra
20 Aug 2026 23:31 UTC
8
points
2
comments
6
min read
LW
link
Llama will abandon a correct answer if it thinks you’re educated
Nick Merrill
20 Aug 2026 20:06 UTC
25
points
6
comments
2
min read
LW
link
If Aliens Exist, We Should Expect to Find Them Around Now
jehan
20 Aug 2026 18:57 UTC
10
points
1
comment
1
min read
LW
link
(www.jehanazad.com)
We Must Remember That Our World Contains Hell
James Brobin
20 Aug 2026 14:22 UTC
118
points
34
comments
3
min read
LW
link
Making sense of the misalignment risk model in the Anthropic Risk Report (August 2026)
jasmine.ren
20 Aug 2026 2:33 UTC
27
points
5
comments
8
min read
LW
link
Steering Role Confusion
Kevin Zhang
,
Daniel Lee
and
Jaewoo Park
20 Aug 2026 2:09 UTC
19
points
1
comment
4
min read
LW
link
Judging ethical theories by update rules, not by action rankings
yatharth
19 Aug 2026 22:41 UTC
11
points
1
comment
5
min read
LW
link
The Rogue Agent Explosion Will Be Mostly Invisible
Steven McCulloch
19 Aug 2026 18:46 UTC
82
points
16
comments
16
min read
LW
link
RL creates split personas
Jan Betley
19 Aug 2026 18:23 UTC
177
points
17
comments
4
min read
LW
link
A failed solution to open-source game theory
Richard Willis
19 Aug 2026 16:11 UTC
15
points
3
comments
4
min read
LW
link
Inside the mind of a fair player cooperating
transhumanist_atom_understander
19 Aug 2026 14:53 UTC
35
points
2
comments
3
min read
LW
link
Some reasons alignment doesn’t generalise well
Lucius Bushnaq
19 Aug 2026 14:18 UTC
107
points
1
comment
9
min read
LW
link
Back to top
Next