flâneur — a map of the web's best reading

Bing Chat is blatantly, aggressively misaligned - LessWrong

lesswrong.com · 4,333 words · saved by 1 readers

Comment by gwern - I've been thinking how Sydney can be so different from ChatGPT, and how RLHF could have resulted in such a different outcome, and here is a hypothesis no one seems to have brought up: "Bing Sydney is not a RLHF trained GPT-3 model at all! but a GPT-4 model developed in a hurry which has been finetuned on some sample dialogues and possibly some pre-existing dialogue datasets or instruction-tuning [https://gwern.net/doc/ai/nn/transformer/gpt/instruction-tuning/index], and this plus the wild card of being able to inject random novel web searches into the prompt are why it acts like it does". This seems like it parsimoniously explains everything thus far. So, some background: 1. The relationship between OA/MS is close but far from completely cooperative, similar to how DeepMind won't share [https://news.ycombinator.com/item?id=34804446] anything with Google Brain. Both parties are sophisticated and understand that they are allies - for now... They share as little as possible. When MS plugs in OA stuff to its services, it doesn't appear to be calling the OA API but running it itself. (That would be dangerous and complex from an infrastructure point of view, anyway.) MS 'licensed [https://news.microsoft.com/source/features/ai/new-azure-openai-service/] the GPT-3 source code [https://blogs.microsoft.com/blog/2020/09/22/microsoft-teams-up-with-openai-to-exclusively-license-gpt-3-language-model/]' for Azure use but AFAIK they did not get the all-important checkpoints or datasets (cf. their investments in ZeRO). So, what is Bing Sydney? It will not simply be unlimited access to the ChatGPT checkpoints, training datasets, or debugged RLHF code. It will be something much more limited, perhaps just a checkpoint. 2. This is not ChatGPT. MS has explicitly stated it is more powerful than ChatGPT, but refused to say anything more straightforward like "it's a more trained GPT-3" etc. If it's not a ChatGPT, then

x Conversations with AIs Language Models (LLMs) LLM Personas Microsoft Bing / Sydney AI Frontpage 395 Bing Chat is blatantly, aggressively misaligned by evhub 15th Feb 2023 2 min read 181 395 I haven't seen this discussed here yet, but the examples are quite striking, definitely worse than the ChatGPT jailbreaks I saw. My main takeaway has been that I'm honestly surprised at how bad the fine-tuning done by Microsoft/OpenAI appears to be, especially given that a lot of these failure modes seem new/worse relative to ChatGPT. I don't know why that might be the case, but the scary hypothesis here

Explore this link on the map →

related reading