[2510.26418] Chain-of-Thought Hijacking
Abstract:Large Reasoning Models (LRMs) improve task performance through extended inference-time reasoning. Although previous studies suggest that longer reasoning should lead to more robust safety behavior, we find evidence to the contrary: over-extended reasoning can instead be exploited to systematically weaken refusal behavior. We propose Chain-of-Thought Hijacking, a simple yet effective black-box jailbreak attack that induces LRMs to engage in prolonged benign puzzle-solving reasoning, often lasting more than five minutes, before eliciting harmful compliance. Across HarmBench, CoT Hijacking achieves attack success rates of 99%, 94%, 100%, and 94% on Gemini 2.5 Pro, ChatGPT o4 Mini, Grok 3 Mini, and Claude 4 Sonnet, respectively. To understand why this attack succeeds, we conduct activation probing, attention-pattern analysis, and causal interventions on open-source reasoning models. Our results indicate that refusal behavior depends on a low-dimensional safety signal whose expression weakens as reasoning traces grow longer. In particular, extended benign reasoning shifts attention away from harmful intentions and attenuates refusal-related activations, producing what we call refusal dilution. These findings demonstrate that excessively prolonged reasoning can introduce a systematic jailbreak attack surface. We release our evaluation materials to support reproducibility and further research.
View PDF HTML (experimental) Abstract:Large Reasoning Models (LRMs) improve task performance through extended inference-time reasoning. Although previous studies suggest that longer reasoning should lead to more robust safety behavior, we find evidence to the contrary: over-extended reasoning can instead be exploited to systematically weaken refusal behavior. We propose Chain-of-Thought Hijacking, a simple yet effective black-box jailbreak attack that induces LRMs to engage in prolonged benign puzzle-solving reasoning, often lasting more than five minutes, before eliciting harmful…
saved by
related reading
- [2608.09867] Stealing Reasoning Traces from Proprietary LLM APIsarxiv.org
- the-illusion-of-thinking.pdfml-site.cdn-apple.com
- As Rocks May Think | Eric Jangevjang.com
- DeepSeek-R1arxiv.org
- [2603.07267] How to Steal Reasoning Without Reasoning Tracesarxiv.org
- Thought Branches: Interpreting LLM Reasoning Requires Resamplingarxiv.org
- Reasoning models don't always say what they think \ Anthropicanthropic.com
- Refusal in LLMs is mediated by a single direction — LessWronglesswrong.com
- [2201.11903] Chain of Thought Prompting Elicits Reasoning in Large Language Modelsarxiv.org
- Stolen Thoughtsstolen-thoughts.com
- Stealing Reasoning Traces from Proprietary LLM APIsarxiv.org
- Can Reasoning Models Obfuscate Reasoning? Stress-Testing Chain-of-Thought Monitorabilityarxiv.org