[2606.20560] How Transparent is DiffusionGemma?
Abstract:LLM reasoning transparency is a critical affordance for understanding model decisions, mitigating misuse and misalignment, and debugging surprising model behaviors. However, DiffusionGemma performs a larger fraction of its computation in a continuous latent space; does this make its reasoning less transparent? We study this question by decomposing transparency into two components: variable transparency, whether we understand intermediate snapshots of a model's computational state; and algorithmic transparency, whether we can use these snapshots to reconstruct the process by which the model arrived at its outputs. Naively, DiffusionGemma has poor variable transparency: its opaque serial depth, the amount of serial computation that occurs in between interpretable model states, seems at first 28.6X higher than the corresponding autoregressive Gemma 4 model. However, we show that we can map the information flowing between denoising steps through an interpretable token bottleneck with no decrease in downstream performance. Treating these intermediate states as interpretable reduces the opaque serial depth to just 1.1X that of Gemma 4. Algorithmic transparency is harder for diffusion models than for autoregressive models because all token predictions in the canvas can change at every denoising step, giving the model the power to implement complicated distributed algorithms during the denoising process. To begin bridging this gap, we conduct a suite of interpretability case studies, uncovering initial evidence of novel diffusion-specific phenomena such as non-chronological reasoning, token and sequence smearing, and intermediate-context reasoning. Finally, we test monitorability, a key application of transparency that measures whether model outputs are useful for downstream tasks. We find that DiffusionGemma is similarly monitorable to Gemma 4.
[2606.20560] How Transparent is DiffusionGemma? Skip to main content arXiv is now an independent nonprofit! Learn more × Search arXiv Press Enter to search · Advanced search --> Computer Science > Machine Learning arXiv:2606.20560 (cs) [Submitted on 18 Jun 2026] Title: How Transparent is DiffusionGemma? Authors: Joshua Engels , Callum McDougall , Bilal Chughtai , Janos Kramar , Senthoran Rajamanoharan , Cindy Wu , Arthur Conmy , Asic Q Chen , Jean Tarbouriech , Min Ma , Brendan O'Donoghue , João Gabriel Lopes de Oliveira , Rohin Shah , Neel Nanda View a PDF of the paper titled How
Explore this link on the map →related reading
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- Transformer Circuits Threadtransformer-circuits.pub
- A Visual Guide to DiffusionGemma - by Maarten Grootendorstnewsletter.maartengrootendorst.com
- What are Diffusion Models? | Lil'Loglilianweng.github.io
- Large Language Diffusion Modelsarxiv.org
- Transformer Explainer: LLM Transformer Model Visually Explainedpoloclub.github.io
- Announcing ReasoningLens — Visualizing and Diagnosing LLM Reasoning at a Glancehuggingface.co
- 2409.02908arxiv.org
- 2506.17298arxiv.org
- Neuronpedianeuronpedia.org
- Esoteric Language Modelsarxiv.org
- VaultGemma: The world's most capable differentially private LLMresearch.google