Fabien Roger

Karma: 5,262

Modifying LLM Beliefs with Synthetic Document Finetuning

RowanWang, Johannes Treutlein, Avery, Ethan Perez, Fabien Roger and Sam Marks

Apr 24, 2025, 9:15 PM

70 points

12 comments2 min readLW link

(alignment.anthropic.com)

Reasoning models don’t always say what they think

Joe Benton, Ethan Perez, Vlad Mikulik and Fabien Roger

Apr 9, 2025, 7:48 PM

28 points

4 comments1 min readLW link

(www.anthropic.com)

Alignment Faking Revisited: Improved Classifiers and Open Source Extensions

John Hughes, abhayesian, Akbir Khan and Fabien Roger

Apr 8, 2025, 5:32 PM

146 points

20 comments12 min readLW link

Automated Researchers Can Subtly Sandbag

gasteigerjo, Akbir Khan, Sam Bowman, Vlad Mikulik, Ethan Perez and Fabien Roger

Mar 26, 2025, 7:13 PM

44 points

0 comments4 min readLW link

(alignment.anthropic.com)

Auditing language models for hidden objectives

Sam Marks, Johannes Treutlein, dmz, Sam Bowman, Hoagy, Carson Denison, Kei, 7vik, Akbir Khan, Austin Meek, Euan Ong, Christopher Olah, Fabien Roger, jeanne_, Meg, Drake Thomas, Adam Jermyn, Monte M and evhub

Mar 13, 2025, 7:18 PM

141 points

15 comments13 min readLW link

Do reasoning models use their scratchpad like we do? Evidence from distilling paraphrases

Fabien RogerMar 11, 2025, 11:52 AM

121 points

23 comments11 min readLW link

(alignment.anthropic.com)

Fuzzing LLMs sometimes makes them reveal their secrets

Fabien RogerFeb 26, 2025, 4:48 PM

61 points

13 comments9 min readLW link

How to replicate and extend our alignment faking demo

Fabien RogerDec 19, 2024, 9:44 PM

114 points

5 comments2 min readLW link

(alignment.anthropic.com)

Alignment Faking in Large Language Models

ryan_greenblatt, evhub, Carson Denison, Benjamin Wright, Fabien Roger, Monte M, Sam Marks, Johannes Treutlein, Sam Bowman and Buck

Dec 18, 2024, 5:19 PM

483 points

75 comments10 min readLW link

A toy evaluation of inference code tampering

Fabien RogerDec 9, 2024, 5:43 PM

52 points

0 comments9 min readLW link

(alignment.anthropic.com)

The case for unlearning that removes information from LLM weights

Fabien RogerOct 14, 2024, 2:08 PM

96 points

18 comments6 min readLW link

[Question] Is cybercrime really costing trillions per year?

Fabien RogerSep 27, 2024, 8:44 AM

63 points

28 comments1 min readLW link

An issue with training schemers with supervised fine-tuning

Fabien RogerJun 27, 2024, 3:37 PM

49 points

12 comments6 min readLW link

Best-of-n with misaligned reward models for Math reasoning

Fabien RogerJun 21, 2024, 10:53 PM

25 points

0 comments3 min readLW link

Memorizing weak examples can elicit strong behavior out of password-locked models

Fabien Roger and ryan_greenblatt

Jun 6, 2024, 11:54 PM

58 points

5 comments7 min readLW link

[Paper] Stress-testing capability elicitation with password-locked models

Fabien Roger and ryan_greenblatt

Jun 4, 2024, 2:52 PM

85 points

10 comments12 min readLW link

(arxiv.org)

Open consultancy: Letting untrusted AIs choose what answer to argue for

Fabien RogerMar 12, 2024, 8:38 PM

35 points

5 comments5 min readLW link

Fabien’s Shortform

Fabien RogerMar 5, 2024, 6:58 PM

6 points

114 comments1 min readLW link

Notes on control evaluations for safety cases

ryan_greenblatt, Buck and Fabien Roger

Feb 28, 2024, 4:15 PM

49 points

0 comments32 min readLW link

Protocol evaluations: good analogies vs control

Fabien RogerFeb 19, 2024, 6:00 PM

42 points

10 comments11 min readLW link

Keyboard shortcuts

Keys shown in yellow (e.g., ]) are accesskeys, and require a browser-specific modifier key (or keys).

Keys shown in grey (e.g., ?) do not require any modifier keys.

General
? Show keyboard shortcuts
Esc Hide keyboard shortcuts

Site navigation
h Go to Home (a.k.a. “Frontpage”) view
f Go to Featured (a.k.a. “Curated”) view
a Go to All (a.k.a. “Community”) view
m Go to Meta view
v Go to Tags view
c Go to Recent Comments view
r Go to Archive view
q Go to Sequences view
t Go to About page
u Go to User or Login page
o Go to Inbox page

Page navigation
, Jump up to top of page
. Jump down to bottom of page
/ Jump to top of comments section
s Search

Page actions
n New post or comment
e Edit current post

Post/comment list views
. Focus next entry in list
, Focus previous entry in list
; Cycle between links in focused entry
Enter Go to currently focused entry
Esc Unfocus currently focused entry
] Go to next page
[ Go to previous page
\ Go to first page
e Edit currently focused post

Editor
k Bold text
i Italic text
l Insert hyperlink
q Blockquote text

Appearance
= Increase text size
- Decrease text size
0 Reset to default text size
′ Cycle through content width settings
1 Switch to default theme [A]
2 Switch to dark theme [B]
3 Switch to grey theme [C]
4 Switch to ultramodern theme [D]
5 Switch to simple theme [E]
6 Switch to brutalist theme [F]
7 Switch to ReadTheSequences theme [G]
8 Switch to classic Less Wrong theme [H]
9 Switch to modern Less Wrong theme [I]
; Open theme tweaker
Enter Save changes and close theme tweaker
Esc Close theme tweaker (without saving)

Slide shows
l Start/resume slideshow
Esc Exit slideshow
→↓ Next slide
←↑ Previous slide
Space Reset slide zoom

Miscellaneous
x Switch to next view on user page
z Switch to previous view on user page
` Toggle compact comment list view
g Toggle anti-kibitzer