Retail / E-commerce

Shopify: Replacing the product classification model

Company
Shopify
Country
Canada
Adoption stage
In operation
Source published
Date basis
The date the source was published. It can differ from the date adoption started.
How the source was checked
Read the full source text

The work problem

According to a Shopify engineering blog post, Shopify handles products from millions of businesses, from handcrafted jewelry to industrial equipment, and understanding product categories and attributes is what lets it improve search, discovery and recommendations. The same post explained that classification began in 2018 with TF-IDF and logistic regression and struggled as products grew more complex and diverse, and that even after a 2020 approach combining images and text, category classification alone fell short of understanding products. The same post said that by early 2023 Shopify needed more granular product understanding, a consistent taxonomy across the platform, attribute extraction that depends on the category, metadata such as simplified product descriptions and content tags, and content safety features.

Technology and data

The same post explained that Shopify combined the Shopify Product Taxonomy, which spans more than 26 business verticals, more than 10,000 product categories and more than 1,000 attributes, with Vision Language Models. According to the same post, the model moved from LLaVA 1.5 7B to LLaMA 3.2 11B and now Qwen2VL 7B, and each change was compared with the existing pipeline on performance metrics and compute cost. The same post said a Dataflow pipeline calls the model service twice, first producing a category and a simplified description, then predicting attributes with that category as context. According to the same post, the service runs with Nvidia Dynamo on a Kubernetes cluster with NVIDIA GPUs, raises throughput with FP8 quantization, in-flight batching and a KV cache, handles category and attribute predictions like a transaction so that both must succeed, and retries automatically when part of a prediction fails. According to the same post, the training data uses labels that several LLMs assign to each product independently, and a dedicated judge model settles cases where the labels disagree. The same post explained that people pick out complex edge cases and new product types for review, and regular quality audits keep the labeling standards in place.

Results

According to the same post, merchants accepted 85% of predicted categories, and the system makes more than 30 million predictions a day and has processed billions of historical products. The same post said hierarchical precision and recall doubled compared with the earlier neural network approach. The same post explained that accurate categories improve search relevance and tax calculations, automatic attribute tagging reduces manual work for merchants, and the attribute system now spans all product categories and is also used for automated content screening.

Limits and open questions

The same post does not say over what period or on what sample the 85% acceptance rate and the doubling were measured. It also does not give the date of the switch to the current Qwen2VL 7B, and changes in business metrics such as sales were not confirmed in this research.

Sources

Compiled from public sources. These are not results from ATF Works customers.

Read original (opens in a new tab)