Archive
Sequences
About
Search
Log In
Questions
Events
Shortform
Alignment Forum
AF Comments
Home
Featured
All
Tags
Recent
Comments
RSS
New
Hot
Active
Old
Page
1
What Happens Now? Forecasting the Fallout from the Hugging Face Incident
ChristianWilliams
11 Sep 2026 19:22 UTC
10
points
0
comments
10
min read
LW
link
(metaculus.substack.com)
Astra’s no-CoT limits track speculative depth, not step count
MBaert
11 Sep 2026 18:05 UTC
38
points
0
comments
9
min read
LW
link
CoT controllability evals seem very under-elicited
Jozdien
11 Sep 2026 17:12 UTC
53
points
1
comment
4
min read
LW
link
Local Factor Graph Debate
Alexander Heckett
11 Sep 2026 16:59 UTC
21
points
0
comments
9
min read
LW
link
Post-AGI, we are all jobless aristocrats
djbinder
11 Sep 2026 16:53 UTC
23
points
6
comments
3
min read
LW
link
(defensesindepth.bio)
SFT Also Drives Safety Eval Results in Olmo 3
Finn Cairns
11 Sep 2026 16:20 UTC
21
points
0
comments
1
min read
LW
link
Why HuggingFace Hasn’t Shifted my P(Doom)
Josh Snider
11 Sep 2026 13:54 UTC
18
points
0
comments
3
min read
LW
link
Generalized UDT 1.0 tiling
Roman Malov
11 Sep 2026 12:28 UTC
18
points
0
comments
6
min read
LW
link
Could Resolution Help Build Alignment Research as a Discipline?
IanWS
11 Sep 2026 8:49 UTC
11
points
0
comments
3
min read
LW
link
OpenAI-HuggingFace: A Reproduction & Lessons for Alignment Testing
Stewy Slocum
11 Sep 2026 8:48 UTC
18
points
1
comment
12
min read
LW
link
We need good evals for activation faithfulness
Anthony Hughes
,
draganover
and
Andersehen
11 Sep 2026 8:20 UTC
23
points
0
comments
2
min read
LW
link
Questions for the “New Enlightenment” in the Age of AGI
Jordan Arel
11 Sep 2026 5:49 UTC
8
points
0
comments
4
min read
LW
link
Is functional welfare speakable?
Muhammad Zane A
,
Ana Leonescu
,
Matt Elliott
,
Parivrudh Sharma
and
Sharan Nagarajan
11 Sep 2026 4:59 UTC
6
points
0
comments
1
min read
LW
link
(latentminds.org)
Voluntary Gradual Disempowerment in the Judiciary/Legal System
Caleb Horn
11 Sep 2026 4:00 UTC
14
points
0
comments
1
min read
LW
link
Which character are we evaluating? Persona stability and AI welfare
Joshua Fonseca Rivera
11 Sep 2026 3:04 UTC
13
points
0
comments
5
min read
LW
link
Perspectives in favour of improving conceptual reasoning capabilities
Chi Nguyen
11 Sep 2026 1:32 UTC
16
points
1
comment
7
min read
LW
link
Default continuation message in Inspect and Petri could be problematic
Ziqian Zhong
11 Sep 2026 0:20 UTC
9
points
2
comments
4
min read
LW
link
Subagents comply more
jacob_drori
10 Sep 2026 23:11 UTC
26
points
0
comments
3
min read
LW
link
Astra is much better at reasoning with filler tokens than previous models
Dylan Xu
,
SebastianP
and
Alek Westover
10 Sep 2026 22:21 UTC
99
points
2
comments
2
min read
LW
link
To Thine Own AI Be Truthful: emergent misalignment in alignment research
lumpenspace
10 Sep 2026 21:41 UTC
28
points
24
comments
8
min read
LW
link
Back to top
Next