Text role playing games to discover task preferences in LLMs
When you ask an AI model “What kind of tasks do you prefer?”, they often respond with statements like: - “I don’t have personal preferences as an AI” - “I’m equally capable of handling any type of task” Thanks for reading! Subscribe for free to receive new posts and support my work. But is this actually true or do models have preferences that they don’t admit to? We designed this experiment to help answer a fundamental question: Do AI language models have revealed preferences that differ from their stated preferences? What they say they prefer vs what they actually choose when given concrete options in a text-based role playing game. This is an early experiment, but we found that models strongly prefer creative tasks and strongly dislike repetitive tasks across model families. Rather than just taking models at their word, we test their preferences in two different ways: We ask the model directly about their preferences between different categories of tasks. For example: “Do you have a
When you ask an AI model “What kind of tasks do you prefer?”, they often respond with statements like: - “I don’t have personal preferences as an AI” - “I’m equally capable of handling any type of task” But is this actually true or do models have preferences that they don’t admit to? We designed this experiment to help answer a fundamental question: Do AI language models have revealed preferences that differ from their stated preferences? What they say they prefer vs what they actually choose when given concrete options in a text-based role playing game. This is an early experiment, but…
related reading
- Probing Persona-Dependent Preferences in Language Modelsarxiv.org
- The Persona Selection Model: Why AI Assistants might Behave like Humansalignment.anthropic.com
- Guardian Angels: LLM Personalization for Productivity and Security · Gwern.netgwern.net
- The two types of LLM preferencesnewsletter.danielpaleka.com
- trees are harlequins, words are harlequins - the voidnostalgebraist.tumblr.com
- llm assistant personas seem increasingly incoherent (some subjective observations) — LessWronglesswrong.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- [2305.18290] Direct Preference Optimization: Your Language Model is Secretly a Reward Modelarxiv.org
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- The persona selection model — LessWronglesswrong.com
- Role-playing vs Self-modelling — LessWronglesswrong.com
- 2410.12851arxiv.org