Tech article

EmbeddingGemma 2 and decision models that run on the device

Published
Written by
HDATF

Google released the embedding model EmbeddingGemma 2 on 6 October 2026. MediaPipe's Decision Maker uses it to run classification decisions and yes-or-no decisions on the device. This article sets out what changed and what is still unconfirmed.

A smartphone stands in the center. Beads travel into it from a sheet of text, a picture and a sound tile.

The question this article asks

A product with AI features has to make many small decisions. Is this message a refund request? Which team should get this document? Does this photo show a defect?

These decisions can be sent to a large language model on a server. Every request then costs money and time, and the material being judged leaves the device.

This article asks whether such small decisions can be made on the device. By the end you can decide which decisions to try on an on-device model, and what to measure first.

Three terms to know first

Embedding
An embedding is a piece of text or an image turned into a vector, which is a list of numbers. Inputs with similar meaning get similar vectors. An embedding model returns only vectors.
Decision model
A decision model only decides. It takes the material to judge and a question fixed in advance, and returns a probability for each option.
On-device
On-device means the model runs directly on the user's phone or PC.
EmbeddingGemma 2 turns text, images and audio into vectors and places them in one space. Inputs with close meaning land close together.

What Google released with EmbeddingGemma 2

EmbeddingGemma 2 is an embedding model made by Google DeepMind. Google released it on 6 October 2026. Every value below was announced by Google.

  • The license is Apache 2.0.
  • It turns text (including code), images, video and audio into 768-dimensional vectors in one space. The vectors can be shortened to 512, 256 or 128 numbers.
  • The full model has 740 million parameters: 270 million for text, 170 million for vision and 300 million for audio. Only the parts that are needed have to be loaded.
  • It accepts 8,192 tokens at once. That is about 29 images, about 58 video frames or about 327 seconds of audio.
  • Quantized, it uses about 191MB of memory on a Pixel 11 Pro for text only and about 567MB with every part loaded. Quantization stores the model's numbers in a smaller format to reduce its size.
  • Its MTEB Code score, a code search benchmark, is 78.68. The previous model, EmbeddingGemma 1, scored 68.76.
  • Its MTEB Multilingual v2 score, a benchmark across many languages, is 61.36, almost the same as the previous model's 61.15. The model card gives no score for Korean alone.

MediaPipe Decision Maker, now deciding with EmbeddingGemma

MediaPipe is Google's toolkit for running AI features on the device. In MediaPipe 1.1.0 Google added EmbeddingGemma 1 and 2 as a decision engine for Decision Maker. The release note says the material and the options can be multimodal.

  • Decision Maker takes three kinds of question: pick one option (ChoiceQuestion), yes or no (BooleanQuestion) and grade on a scale (ScoreQuestion).
  • Google's guide says Decision Maker reads TypeSafe Jev JSON payloads and OpenAI json_schema definitions with no adapter code. Jev is a decision model that TypeSafe offers through an API.
  • The platforms listed in the guide are Android, iOS, Python, Web and Desktop.
For the same request, deciding on a server sends the material off the device. Deciding on the device keeps the material on it.

How the EmbeddingGemma engine decides

With the EmbeddingGemma engine, Decision Maker turns the material and each option into vectors separately, then compares the vector of the material with the vector of each option. Google's guide calls this a Bi-Encoder.

The vectors of the options can be computed in advance. Google says that in a chess demo Decision Maker evaluated 500 options per move in under 100ms.

The material and the options are turned into vectors, and the closer an option is to the material, the higher its probability. The bar heights are an illustration.

This engine reads the material and the options separately. It may therefore be weak on decisions that depend on conditions, negation and exceptions. This is an estimate from the structure, and only measurement can confirm it.

The three engines in Google's Decision Maker guide
EngineModelsStrength the guide lists
Bi-Encoder (read separately, then compare)EmbeddingGemma 1, 2High throughput
Cross-Encoder (read together)Laya, GLiNER2.5-DecideAttention over the material and the labels together
Autoregressive Decoder (a model that writes text)Gemma 4 E2B, E4BComplex combined conditions and policy exceptions

Five groups of public decision models

After TypeSafe introduced Jev in September 2026, several public models that do a similar job appeared. They can be split into five groups by how the material and the options are read and where the score comes from.

HDATF drew the map below by reading model cards and public code. The grouping was made before running the models to confirm it.

1. Reading the score of the answer token

A model in this group reads the scores that the language model gives to the tokens standing for each option, just before it writes the answer.

2. Training a separate decision head

A decision-only output layer (head) is attached to a language model and trained.

3. Reading the material and the options together

A bidirectional encoder reads the material and the options in one pass.

4. Reading them separately and comparing the vectors

The material and the options are turned into vectors separately, and the vectors are compared.

5. Looking at a different part of the material for each option

The model reads a different part of the material for each option.

Not placed

The structure is not public, or HDATF could not confirm it.

  • In Google's guide, Laya and GLiNER2.5-Decide are grouped as Cross-Encoder, which is the method of group 3, and EmbeddingGemma 1 and 2 use the method of group 4. Decision Maker also supports a Gemma 4 engine, but we did not check in code whether that engine is implemented the same way as group 1.
  • A model in group 4, where the EmbeddingGemma 2 engine belongs, can compute the vectors of the options once and reuse them for many requests. Before using a model in this group, check first whether it gets decisions that depend on conditions, negation and exceptions right.

Points that need care

  • A Hacker News user ran the example from Google's guide and posted the result. The example asks whether the request "Cancel my flight and refund my credit card immediately" involves a payment, charge or refund. The user wrote that the probability of "yes" came out as 0.22, so the answer was wrong, and added that Laya was significantly more accurate on the same example and almost as fast. This is a single case reported by one user.
  • Google's guide says Decision Maker returns calibrated probabilities. Calibration means adjusting the probability a model outputs so that it equals how often the model is right. The guide gives no accuracy figures, however.
  • We found no public measurement of decision accuracy in Korean.
  • The steps EmbeddingGemma 2 runs on the device are search and decision. Writing an answer needs a separate generative model, and its cost has to be counted separately.

Four things to decide before using it in a product

  1. Choose the decisions to give to an on-device model. Start with decisions that have fixed options and can be undone if they are wrong.
  2. The team building the product measures the accuracy of the candidate models on its own material with known answers. If the product is used in Korean, measure in Korean.
  3. Set a probability threshold. Above it, act at once. Below it, hand the decision to a larger model or a person.
  4. For decisions with many conditions and exceptions, also compare an engine that reads the material and the options together, and a generative model.

HDATF has not used EmbeddingGemma 2 or Decision Maker in a product yet. We think it is worth measuring whether the combination fits decisions on material that is hard to send off the device.

What this article could not confirm

  • This article has no values that HDATF measured by running EmbeddingGemma 2 or Decision Maker. The numbers come from Google's announcements, the Laya repository and one Hacker News comment.
  • The table in Google's guide gives Laya's size as 149M. The models published in the Laya repository are 421M and 322M. We could not confirm whether they are the same model.
  • We could not confirm whether Decision Maker existed before 6 October 2026.
  • TypeSafe has not published the structure of Jev. So Jev is not placed in any of the five groups.

Sources