[2401.12070] Spotting LLMs With Binoculars: Zero-Shot Detection of Machine-Generated Text
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs. arXiv Operational Status Get status notifications via email or slack
View PDF HTML (experimental) Abstract:Detecting text generated by modern large language models is thought to be hard, as both LLMs and humans can exhibit a wide range of complex behaviors. However, we find that a score based on contrasting two closely related language models is highly accurate at separating human-generated and machine-generated text. Based on this mechanism, we propose a novel LLM detector that only requires simple calculations using a pair of pre-trained LLMs. The method, called Binoculars, achieves state-of-the-art accuracy without any training data. It is capable of…
saved by
related reading
- Ghostbuster: Detecting Text Ghostwritten by Large Language Modelsarxiv.org
- Pangram 4 Technical Reportpangram-public.s3.us-east-1.amazonaws.com
- GenAI Handbookgenai-handbook.github.io
- How does Pangram work?pangram.substack.com
- Cognitive Bias Detection Using Advanced Prompt Engineeringarxiv.org
- Pangram 4 Technical Overviewpangram.com
- How does Pangram work?substack.com
- Seeing in Pangram Space | Pangrampangram.com
- Prefill awareness: can LLMs tell when “their” message history has been tampered with? — LessWronglesswrong.com
- The Mark of the Machinenewyorker.com
- Why Perplexity and Burstiness Fail to Detect AI | Pangram Labspangram.com
- (Im)possibility of Automated Hallucination Detection in Large Language Modelsarxiv.org