Retail / Online grocery

Instacart: Interpreting search intent and product attributes

Company
Instacart
Country
United States
Adoption stage
In operation
Source published
Date basis
The date the source was published. It can differ from the date adoption started.
How the source was checked
Read the full source text

The work problem

Rough, unpolished wording arrives in the Instacart search box. In the past a different model was built for each task. Query classification was handled by a FastText model, and query rewriting by a separate system that mined user session logs. The rewriting system covered only 50 percent of search traffic and could not produce a usable alternative for rare search terms.

Technology and data

The company rebuilt query understanding around LLMs. Frequently seen search terms are handled offline. This is a RAG approach that puts the top brands and categories from conversion history, and brand name similarity from the product catalogue, into the prompt. The results go into a cache and at the same time serve as training material for the real-time model. Rare search terms not in the cache are handled by a real-time model, Llama-3-8B further trained with LoRA. The output is checked by a post-processor. It computes embedding similarity between the query and the predicted category, discards anything below the threshold, and cross-checks against the catalogue.

Results

Instacart said rewriting coverage exceeds 95 percent and precision for all three types is above 90 percent. The further trained 8B model reached 96.4 percent precision against the baseline model's 95.4 percent, while F1 was 95.7 percent, close to 95.8 percent. Latency was close to 700ms on an A100, but merging the LoRA weights into the base model and switching to H100 brought it to the 300ms target. The company said this system is already running and handles several million new search terms a week. It explained that dissatisfaction with results for tail search terms fell by 50 percent and average scroll depth fell by 6 percent.

Limits and open questions

The metrics disclosed are limited to search quality and latency. Figures showing that this led to a change in orders or revenue are not in the source.

Sources

Compiled from public sources. These are not results from ATF Works customers.

Read original (opens in a new tab)