✳flâneur — a map of the web's best reading
Several frontier models are substantially prefill aware — LessWrong
lesswrong.com · 1,765 words · saved by 1 readers
This blog post discusses work in a recently-published paper. However, this blogpost was primarily written by Parv Mahajan and Andy Wang, and several…
x Several frontier models are substantially prefill aware — LessWrong AI Frontpage 59 Several frontier models are substantially prefill aware by yeedrag , Parv Mahajan , David Africa , alexsouly , Jordan Taylor , RobertKirk 17th Jun 2026 6 min read 2 59 This blog post discusses work in a recently-published paper. However, this blogpost was primarily written by Parv Mahajan and Andy Wang, and several of the more speculative takes may not represent the all-things-considered view of the entire team. Link to paper: https://arxiv.org/abs/2606.12747 TL;DR: We provide more conceptual grounding and ex
Explore this link on the map →saved by
related reading
- [2603.02202] Frontier Models Can Take Actions at Low Probabilitiesarxiv.org
- How confessions can keep language models honest | OpenAIopenai.com
- Prefill awareness: can LLMs tell when “their” message history has been tampered with? — LessWronglesswrong.com
- AI in 2025: gestalt — LessWronglesswrong.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Predicting LLM Safety Before Release by Simulating Deploymentcdn.openai.com
- Predicting model behavior before release by simulating deployment | OpenAIopenai.com
- Claude Sonnet 3.7 (often) knows when it’s in alignment evaluations — LessWronglesswrong.com
- Realistic Evaluations Will Not Prevent Evaluation Awareness — LessWronglesswrong.com
- Peer-Preservation in Frontier Modelsrdi.berkeley.edu
- Petri 2.0: New Scenarios, New Model Comparisons, and Improved Eval-Awareness Mitigationsalignment.anthropic.com
- Models May Behave Worse When Eval Aware — LessWronglesswrong.com