Dakara comments on If we solve alignment, do we die anyway?

Dakara 20 Nov 2024 15:29 UTC
1 point
0
What would instrumental convergence mean in this case? I am not sure of what that means in this case.
- Noosphere89 20 Nov 2024 15:46 UTC
  2 points
  0
  Parent
  In this case, it would mean the convergence to preserve your current values.
  - Dakara 20 Nov 2024 15:56 UTC
    1 point
    0
    Parent
    Reading from LessWrong wiki, it says “Instrumental convergence or convergent instrumental values is the theorized tendency for most sufficiently intelligent agents to pursue potentially unbounded instrumental goals such as self-preservation and resource acquisition”
    
    It seems like it preserves exactly the goals we wouldn’t really need it to preserve (like resource acquisition). I am not sure how it would help us with preserving goals like ensuring humanity’s prosperity, which seem to be non-fundamental.
    - Noosphere89 20 Nov 2024 16:05 UTC
      3 points
      0
      Parent
      Yes, I admittedly want to point to something along the lines of preserving your current values being a plausibly major drive of AIs.
      - Dakara 20 Nov 2024 16:09 UTC
        1 point
        0
        Parent
        Ah, so you are basically saying that preserving current values is like a meta instrumental value for AGIs similar to self-preservation that is just kind of always there? I am not sure if I would agree with that (if I am correctly interpreting you) since, it seems like some philosophers are quite open to changing their current values.
        Noosphere89 20 Nov 2024 16:31 UTC
        2 points
        0
        Parent
        Not always, but I’d say often.
        I’d also say that at least some of the justification for changing values in philosophers/humans is because they believe the new values are closer to the moral reality/truth, which is an instrumental incentive.
        To be clear, I’m not going to state confidently that this will happen (maybe something like instruction following ala @Seth Herd is used instead, such that the pointer is to the human giving the instructions, rather than having values instead), but this is at least reasonably plausible IMO.
        Dakara 10 Dec 2024 13:48 UTC
        1 point
        0
        Parent
        I have found another possible concern of mine.
        
        Consider gravity on Earth, it seems to work every year. However, this fact alone is consistent with theories that gravity will stop working in 2025, 2026, 2027, 2028, 2029, 2030, etc. There are infinite such theories and only one theory that gravity will work as an absolute rule.
        
        We might infer from the simplest explaination that gravity holds as an absolute rule. However, the case is different with alignment. To ensure AI alignment, our evidence must rule out whether an AI is following a misaligned rule compared to an aligned rule based on time-and situation-limited data.
        
        While it may be safe, for all practical purposes, to assume that simpler explanations tend to be correct when it comes to nature, we cannot safely assume this for LLMs—for the reason that the learning algorithms that are programmed into them can have complex unintended consequences for how the LLM will behave in the future, given the changing conditions an LLM finds itself in.
        
        Doesn’t this mean that it is not possible to achieve alignment?
        Dakara 17 Dec 2024 18:28 UTC
        1 point
        0
        Parent
        I have posted this text as a standalone question here
        Dakara 20 Nov 2024 17:22 UTC
        1 point
        0
        Parent
        Fair enough. Would you expect that AI would also try to move its values to the moral reality? (something that’s probably good for us, cause I wouldn’t expect human extinction to be a morally good thing)
        Noosphere89 20 Nov 2024 17:27 UTC
        3 points
        0
        Parent
        The problem with that plan is that there are too many valid moral realities, so which one you do get is once again a consequence of alignment efforts.
        
        To be clear, I’m not stating that it’s hard to get the AI to value what we value, but it’s not so brain-dead easy that we can make the AI find moral reality and then all will be well.
        Dakara 20 Nov 2024 17:43 UTC
        3 points
        0
        Parent
        Noosphere, I am really, really thankful for your responses. You completely answered almost all (I am still not convinced about that strategy of avoiding value drift. I am probably going to post that one as a question to see if maybe other people have different strategies on preventing value drift) of the concerns that I had about alignment.
        
        This discussion, significantly increased my knowledge. If I could triple upvote your answers, I would. Thank you! Thank you a lot!
        Dakara 20 Nov 2024 22:41 UTC
        1 point
        0
        Parent
        P.S. Here is the link to the question that I posted.
        Dakara 22 Nov 2024 7:53 UTC
        1 point
        0
        Parent
        Another concern that I could see with the plan. Step 1 is to create safe and alignment AI, but there are some results which suggest that even current AIs may not be as safe as we want them to be. For example, according to this article, current AI (specifically o1) can help novices build CBRN weapons and significantly increase threat to the world. Do you think this is concerning or do you think that this threat will not materialize?
        Noosphere89 22 Nov 2024 17:39 UTC
        2 points
        0
        Parent
        The threat model is plausible enough that some political actions should be done, like banning open-source/open-weight models, and putting in basic Know Your Customer checks.
        Dakara 22 Nov 2024 17:41 UTC
        1 point
        0
        Parent
        Isn’t it a bit too late for that? If o1 gets publicly released, then according to that article, we would have an expert-level consultant in bioweapons available for everyone. Or do you think that o1 won’t be released?
        Expand this thread
        Noosphere89 22 Nov 2024 17:46 UTC
        3 points
        2
        Parent
        I don’t buy that o1 has actually given people expert-level bioweapons, so my actions here are more so about preparing for future AI that is very competent at bioweapon building.
        
        Also, even with the current level of jailbreak resistance/adversarial example resistance, assuming no open-weights/open sourcing of AI is achieved, we can still make AIs that are practically hard to misuse by the general public.
        
        See here for more:
        
        https://www.lesswrong.com/posts/KENtuXySHJgxsH2Qk/managing-catastrophic-misuse-without-robust-ais
        Dakara 22 Nov 2024 17:31 UTC
        1 point
        0
        Parent
        After some thought, I think this is a potentially really large issue which I don’t know how we can even begin to solve. We can have aligned AI, being aligned with someone who wants to create bioweapons. Is there anything being done (or anything that can be done) to prevent that?
        Noosphere89 22 Nov 2024 17:40 UTC
        3 points
        1
        Parent
        The answers to this question is actually 2 things:
        
        This is why I expect we will eventually have to fight to ban open-source, and we will have to get the political will to ban both open-source and open-weights AI.
        
        This is where the unlearning field comes in. If we could make the AI unlearn knowledge, an example being nuclear weapons, we could possibly distribute AI safely without causing novices to create dangerous stuff.
        
        More here:
        
        https://www.lesswrong.com/posts/mFAvspg4sXkrfZ7FA/deep-forgetting-and-unlearning-for-safely-scoped-llms
        
        https://www.lesswrong.com/posts/9AbYkAy8s9LvB7dT5/the-case-for-unlearning-that-removes-information-from-llm
        
        But the solutions are intentionally going to make AI safe without relying on alignment.