flâneur

Evaluating honesty and lie detection techniques on a diverse suite of dishonest models

alignment.anthropic.com · saved by 1 readers

We use a suite of testbed settings where models lie—i.e. generate statements they believe to be false—to evaluate honesty and lie detection techniques. The best techniques we studied involved fine-tuning on generic anti-deception data and using prompts that encourage honesty. Suppose we had a “truth serum for AIs”: a technique that reliably transforms a language model  into an honest model  that generates text which is truthful to the best of its own knowledge. How useful would this discovery be for AI safety? We believe it would be a major boon. Most obviously, we could deploy  in place of . Or, if our “truth serum” caused side-effects that limited ’s commercial value (like capabilities degradation or refusal to engage in harmless fictional roleplay),  could still be used by AI developers as a tool for ensuring ’s safety. For example, we could use to audit  for alignment pre-deployment. More ambitiously (and speculatively), while training , we could leverage for oversight by incorpo

saved by