How Transparent is DiffusionGemma?
- Published
- Source
- arXiv
- Paper number
- 458
- Field
- Machine Learning
- arXiv ID
- 2606.20560
Key points
- The paper decomposes transparency into variable transparency, which is how well intermediate states can be understood, and algorithmic transparency, which is how well the process can be reconstructed.
- Opaque serial depth drops from 28.6 times the initial Gemma 4 level to 1.1 times when mapped to a token bottleneck.
- Compressing self-conditioning between denoising steps into O(c) tokens causes only a small downstream performance drop.
- On monitoring evaluation, DiffusionGemma performs about the same as Gemma 4 in terms of CoT monitoring quality.
- The paper finds diffusion-specific phenomena such as non-sequential reasoning, token and sequence smearing, and intermediate-context reasoning.
- It provides a framework and benchmark for auditing the transparency of future latent reasoning models.
Paper links
External research summaries. These are not HDATF publications or measured product results.