Sodium comments on Goals selected from learned knowledge: an alternative to RL alignment

Sodium Jan 27, 2024, 10:32 PM
2 points
1
Thanks for the response!

I’m worried that instead of complicated LMA setups with scaffolding and multiple agents, labs are more likely to push for a single tool using LM agent, which seems cheaper and simpler. I think some sort of internal steering for a given LM based on learned knowledge discovered through interpretability tools is probably the most competitive method. I get your point that the existing method in LLMs aren’t necessarily re targeting some sort of searching method, but at the same time they don’t have to be? Since there isn’t this explicit search and evaluation process in the first place, I think of it more as a nudge guiding LLM hallucinations.
I was just thinking, a really ambitious goal would be apply some sort of GSLK steering to LLAMA and see if you could get it to perform well on the LLM leaderboard, similar to how there’s models there that’s just DPO applied to LLAMA.
- Seth Herd Jan 27, 2024, 11:09 PM
  2 points
  0
  Parent
  What I’m envisioning is a single agent, with some scaffolding of episodic memory and executive function to make it more effective. If I’m right, that would be not the simplest, but the cheapest way to AGI, since it fills some gaps in the language model’s abilities without using brute force. I wrote about this vision of language model cognitive architectures here.
  I’m realizing that the distinction between a minimal language model agent and the sort of language model cognitive architecture I think will work better is a real distinction, and most people assume with you that a language model agent will just be a powerful LLM prompted over and over with something like “keep thinking about that, and take actions or get data using these APIs when it seems useful”. That system will be much less explicitly goal directed than an LMA with additonal executive function to keep it on-task and therefore goal-directed.
  I intend to write a post about that distinction.
  On your original question, see also Kristin’s comment and the paper she suggests. It’s work on a modification of the transformer algorithm to make more easily interpretable representations. I meant to mention it, and Roger Dearnaley’s post on it. I do find this a promising route to better interpretability. The ideal foundation model for a safe agent would be a language model that’s also trained with an algorithm that encourages interpretable representations.

Keyboard shortcuts

Keys shown in yellow (e.g., ]) are accesskeys, and require a browser-specific modifier key (or keys).

Keys shown in grey (e.g., ?) do not require any modifier keys.

General
? Show keyboard shortcuts
Esc Hide keyboard shortcuts

Site navigation
h Go to Home (a.k.a. “Frontpage”) view
f Go to Featured (a.k.a. “Curated”) view
a Go to All (a.k.a. “Community”) view
m Go to Meta view
v Go to Tags view
c Go to Recent Comments view
r Go to Archive view
q Go to Sequences view
t Go to About page
u Go to User or Login page
o Go to Inbox page

Page navigation
, Jump up to top of page
. Jump down to bottom of page
/ Jump to top of comments section
s Search

Page actions
n New post or comment
e Edit current post

Post/comment list views
. Focus next entry in list
, Focus previous entry in list
; Cycle between links in focused entry
Enter Go to currently focused entry
Esc Unfocus currently focused entry
] Go to next page
[ Go to previous page
\ Go to first page
e Edit currently focused post

Editor
k Bold text
i Italic text
l Insert hyperlink
q Blockquote text

Appearance
= Increase text size
- Decrease text size
0 Reset to default text size
′ Cycle through content width settings
1 Switch to default theme [A]
2 Switch to dark theme [B]
3 Switch to grey theme [C]
4 Switch to ultramodern theme [D]
5 Switch to simple theme [E]
6 Switch to brutalist theme [F]
7 Switch to ReadTheSequences theme [G]
8 Switch to classic Less Wrong theme [H]
9 Switch to modern Less Wrong theme [I]
; Open theme tweaker
Enter Save changes and close theme tweaker
Esc Close theme tweaker (without saving)

Slide shows
l Start/resume slideshow
Esc Exit slideshow
→↓ Next slide
←↑ Previous slide
Space Reset slide zoom

Miscellaneous
x Switch to next view on user page
z Switch to previous view on user page
` Toggle compact comment list view
g Toggle anti-kibitzer