Evaluating honesty and lie detection techniques on a diverse suite of dishonest models
We use a suite of testbed settings where models lie—i.e. generate statements they believe to be false—to evaluate honesty and lie detection techniques. The best techniques we studied involved fine-tuning on generic anti-deception data and using prompts that encourage honesty. Suppose we had a “truth serum for AIs”: a technique that reliably transforms a language model into an honest model that generates text which is truthful to the best of its own knowledge. How useful would this discovery be for AI safety? We believe it would be a major boon. Most obviously, we could deploy in place of . Or, if our “truth serum” caused side-effects that limited ’s commercial value (like capabilities degradation or refusal to engage in harmless fictional roleplay), could still be used by AI developers as a tool for ensuring ’s safety. For example, we could use to audit for alignment pre-deployment. More ambitiously (and speculatively), while training , we could leverage for oversight by incorpo