Yeedrag
7 followers · 10 following · 302 views
on the atlas — 2
- Can Reasoning Models Obfuscate Reasoning? Stress-Testing Chain-of-Thought Monitorability1 savers
- Reality has a surprising amount of detail49 savers
highlights — 2
In the absence of obfuscation pressure, CoT monitoring proves highly effective: CoT -only monitors achieve near-perfect detection ( ∼ 96–98%) compared to action -only monitors (8–20%), demon- strating that malicious intent is transparent in reasoning traces. Notably, even a weak CoT monitor (GPT-4o mini) achieves ∼ 96% detection—outperforming the strongest action and full (sees both actions and CoT) monitors (GPT-5).
Can Reasoning Models Obfuscate Reasoning? Stress-Testing Chain-of-Thought MonitorabilityIf you wish to not get stuck, seek to perceive what you have not yet perceived.
Reality has a surprising amount of detail