Retail / Online grocery

Instacart: Producing images of products and recommendation sets

Company
Instacart
Country
United States
Adoption stage
In operation
Source published
Date basis
The date the source was published. It can differ from the date adoption started.
How the source was checked
Read the full source text

The work problem

In online grocery, shoppers cannot pick up a product and look at it as they would in a store. Inside Instacart, image generation ran separately in each team. Every team used different models, different prompting styles and different evaluation criteria. Each team had to learn from scratch which prompts work for food photography, which model looks most realistic and how to measure quality.

Technology and data

Instacart built a single image generation platform called PIXEL. It gives access to several models and builds the required parameters and settings on the user's behalf. It carries default prompts for generating and for evaluating images, and teams can change them. It has a screen usable by people without technical knowledge. Automatic evaluation runs in this order: an LLM writes a prompt and the image is generated, an LLM then writes evaluation questions to suit the project's needs, and those questions and the image are passed to a vision language model for judgement. The evaluation questions check composition, consistency, style and overall appeal. Final quality is judged by human reviewers.

Results

Instacart said teams using PIXEL cut the time to create a new image by a factor of 10. The company explained that using a vision language model as a feedback loop lifted the human reviewer approval rate for images from 20 percent to 85 percent. It said that after adding images of meat cuts, browsing time and time to add to cart for those items fell by more than 25 percent, and cart conversion for personalised recommendation sets rose by 15 percent.

Limits and open questions

The figures come from a specific image production procedure and from some items such as meat cuts. The source does not disclose the total volume of images generated or the amount of cost saved.

Sources

Compiled from public sources. These are not results from ATF Works customers.

Read original (opens in a new tab)