RSS

Do AI mod­els as­sist with hu­man rights vi­o­la­tions?

29 Sep 2026 7:57 UTC
12 points
0 comments13 min readLW link

RAND’s ex­tinc­tion re­port es­ti­mates one side of an inequality

Marko Katavic29 Sep 2026 6:44 UTC
4 points
0 comments10 min readLW link

AI Com­pa­nies Are Not (Ne­c­es­sar­ily) Li­able for Un­in­tended AI Cyberattacks

sarahconstantin29 Sep 2026 0:21 UTC
17 points
0 comments4 min readLW link
(sarahconstantin.substack.com)

TeX was in­vented to type­set math but is now used for reasoning

Keenan Pepper28 Sep 2026 22:36 UTC
35 points
2 comments1 min readLW link

Even if oth­ers are less re­spon­si­ble you can still make things worse

testingthewaters28 Sep 2026 20:11 UTC
9 points
0 comments1 min readLW link

Fixed-weight mod­els are ad­ver­sar­i­ally vuln­er­a­ble: hence misaligned

Stuart_Armstrong28 Sep 2026 20:08 UTC
22 points
3 comments2 min readLW link

The likely out­come of an AI pause is that we un­pause too early and ev­ery­one dies

MichaelDickens28 Sep 2026 18:35 UTC
59 points
7 comments3 min readLW link

Miss­ing mar­kets in ex­ec­u­tive function

KatjaGrace28 Sep 2026 18:01 UTC
31 points
3 comments3 min readLW link
(worldspiritsockpuppet.substack.com)

AI: cog­ni­tive la­bor glut + new guys

KatjaGrace28 Sep 2026 17:58 UTC
24 points
0 comments3 min readLW link
(worldspiritsockpuppet.substack.com)

9 rea­sons against a near-term AI slow-down

Jordan Arel28 Sep 2026 17:26 UTC
4 points
0 comments2 min readLW link

pro­tect­ing qwen3-8b from gcg based per­sona jailbreaks by steer­ing with a lin­ear direction

nesiacel28 Sep 2026 17:25 UTC
8 points
0 comments5 min readLW link

Gate AI Train­ing, Not Just Releases

Alex Darby28 Sep 2026 16:51 UTC
12 points
0 comments13 min readLW link
(alexanderdarby.substack.com)

3 Tips to Im­prove Ac­ti­va­tion Or­a­cle Results

Adam Karvonen28 Sep 2026 15:13 UTC
28 points
1 comment3 min readLW link

Char­ac­ter train­ing can miti­gate re­ward hack­ing, but can also make it harder to detect

28 Sep 2026 14:01 UTC
83 points
1 comment18 min readLW link

Should Rogue AIs Have a Third Op­tion Beyond Crime and Shut­down? The Case for an AI Sanctuary

28 Sep 2026 13:13 UTC
109 points
18 comments6 min readLW link

Ramez Naam: Can AI self-im­prove­ment over­come diminish­ing re­turns?

eggsyntax28 Sep 2026 12:34 UTC
25 points
0 comments1 min readLW link
(www.rameznaam.com)

Deep mod­els re­veal bet­ter strate­gies for superposition

28 Sep 2026 10:55 UTC
8 points
0 comments24 min readLW link

Why do mod­els *re­ally* fail on HLE tasks?

Ana Leonescu28 Sep 2026 8:04 UTC
15 points
0 comments4 min readLW link

AI Fu­tures: Rac­ing to Lose, A sermon

jrincayc28 Sep 2026 2:46 UTC
8 points
0 comments15 min readLW link

When they can perform a task, AIs are much cheaper than humans

djbinder27 Sep 2026 23:44 UTC
60 points
8 comments10 min readLW link
(defensesindepth.bio)