SolidGoldMagikarp (plus, prompt generation) | AI Alignment Forum | Jessica Rumbelow, Matthew Watkins
UPDATE (14th Feb 2023): ChatGPT appears to have been patched! However, very strange behaviour can still be elicited in the OpenAI playground, particularly with the davinci-instruct model. …
x SolidGoldMagikarp (plus, prompt generation) — AI Alignment Forum Best of LessWrong 2023 Glitch Tokens Adversarial Examples (AI) MATS Program Interpretability (ML & AI) Language Models (LLMs) AI Curated 134 SolidGoldMagikarp (plus, prompt generation) by Jessica Rumbelow , mwatkins 5th Feb 2023 15 min read 208 134 UPDATE (14th Feb 2023): ChatGPT appears to have been patched! However, very strange behaviour can still be elicited in the OpenAI playground , particularly with the davinci-instruct model. More technical details here . Further (fun) investigation into the stories behind the tokens we
saved by
related reading
- interpreting GPT: the logit lens — LessWronglesswrong.com
- GitHub - brexhq/prompt-engineering: Tips and tricks for working with Large Language Models like OpenAI's GPT-4.github.com
- Where the goblins came from | OpenAIopenai.com
- trees are harlequins, words are harlequins - the voidnostalgebraist.tumblr.com
- GPT-4openai.com
- gpt-4.pdfcdn.openai.com
- Gwern visits BAIR – Yuxi on the Wiredyuxi.ml
- Towards a Typology of Strange LLM Chains-of-Thought1a3orn.com
- Steering GPT-2-XL by adding an activation vector — AI Alignment Forumalignmentforum.org
- microgptkarpathy.github.io
- Non-determinism in GPT-4 is caused by Sparse MoE - 152334H152334h.github.io
- microgptgist.github.com