Delivery and mobility / Ride-hailing and delivery
Uber: Improving food delivery search
- Company
- Uber
- Country
- United States
- Adoption stage
- In operation
- Source published
- Date basis
- The date the source was published. It can differ from the date adoption started.
- How the source was checked
- Read the full source text
The work problem
On Uber Eats, search is the main path to an order. Matching on characters alone breaks down on synonyms, misspellings, abbreviations, queries that mix languages, and words with several meanings. The source wrote that lexical methods see strings, not meaning. A structure was therefore needed that turns queries and documents into vectors and retrieves by meaning.
Technology and data
The design is a two tower structure that encodes queries and documents separately. Query embeddings are produced in real time by an online service, while document embeddings are prepared in advance by scheduled batch jobs. The backbone is a Qwen model fine tuned on internal Uber Eats data, and one model covers every vertical and market. Training uses an MRL based infoNCE loss, so a single model emits embeddings at several sizes. The index is an HNSW graph in Apache Lucene Plus holding both float32 and int8 vectors. Conditions such as area, city, document type and fulfillment type filter first, and approximate nearest neighbour search runs after that. Retraining and reindexing happen every two weeks, swapping blue and green at the column level inside one index rather than keeping separate indexes. Before deployment a run must pass a document count comparison, a byte for byte match on columns that should not change, and a recall comparison from replaying real queries. In operation, sampled requests check that the query model matches the model identifier on the index column, and the model deployment rolls back automatically if mismatches persist.
Results
The source said that lowering the shard level candidate count from 1,200 to about 200 cut latency by 34% and saved 17% of CPU with almost no effect on recall. int7 scalar quantization cut latency by more than 50% against fp32 while holding recall above 0.95. Cutting embeddings down to 256 dimensions kept quality loss under 0.3% for English and Spanish and reduced storage by about 50%. The company said the structure powers restaurant, grocery and retail search together, and that scheduled refreshes run without disrupting live traffic.
Limits and open questions
The source carries no results from an experiment that split real users. It does not report business measures such as order conversion or click changes. The figures given are offline evaluations and infrastructure performance measurements. The source also does not say when adoption began.
Sources
- Evolution and Scale of Uber's Delivery Search Platformuber.com, Accessed
Compiled from public sources. These are not results from ATF Works customers.