Gemma 4 Technical Report

Published
Source
arXiv
Paper number
570
Field
LLMs / NLP
arXiv ID
2607.02770

Key points

  • It consists of dense models from 2.3B to 31B and a 26B-A4B MoE, and every model includes a thinking mode.
  • The 12B model removes the vision and audio encoders and uses an encoder-free architecture that projects raw patches directly into the LLM embedding space.
  • KV-cache sharing, keys-as-values, and p-RoPE reduce the global KV cache by 37.5 percent.
  • With QAT quantization, the 31B model shrinks from 64 GB to 19.2 GB, and the audio encoder shrinks from 390 MB to 87 MB, which is a 78 percent reduction.
  • It is the leading dense open model on Arena Elo at 1451 for the 31B model and reaches 96.4 percent accuracy on RULER 128k.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)