Gemma 4 Technical Report
- Published
- Source
- arXiv
- Paper number
- 570
- Field
- LLMs / NLP
- arXiv ID
- 2607.02770
Key points
- It consists of dense models from 2.3B to 31B and a 26B-A4B MoE, and every model includes a thinking mode.
- The 12B model removes the vision and audio encoders and uses an encoder-free architecture that projects raw patches directly into the LLM embedding space.
- KV-cache sharing, keys-as-values, and p-RoPE reduce the global KV cache by 37.5 percent.
- With QAT quantization, the 31B model shrinks from 64 GB to 19.2 GB, and the audio encoder shrinks from 390 MB to 87 MB, which is a 78 percent reduction.
- It is the leading dense open model on Arena Elo at 1451 for the 31B model and reaches 96.4 percent accuracy on RULER 128k.
Paper links
External research summaries. These are not HDATF publications or measured product results.