Finance / Crypto exchange
Coinbase: Additional fraud risk review for crypto purchases in partner apps
- Company
- Coinbase
- Country
- United States
- Adoption stage
- Limited operation
- Source published
- Date basis
- The date the source was published. It can differ from the date adoption started.
- How the source was checked
- Read the full source text
The work problem
According to a Coinbase engineering blog post, Coinbase Onramp lets users buy crypto inside partner apps and also supports guest checkout for people without a Coinbase account. The same post explained that the existing machine learning models and rules engine screen guest checkouts quickly and cheaply, but find it hard to recognize attacks whose signals are spread across several transactions. The post said that predefined features such as transaction counts or spending totals cannot capture combinations of timing, payment method and destination, and that confirmed fraud outcomes arrive late, which makes it hard to evaluate and adjust for new patterns.
Technology and data
According to the same post, when the Onramp service sends a transaction to the Risk Service, the existing models and rules check it first, and an LLM risk review agent then reviews only selected transactions once more. The post explained that the models and the agent run in the same Ray Serve deployment, and that the agent can block transactions the existing checks allowed but cannot release transactions an earlier stage blocked. The post said the agent receives transaction details, a summary of recent transaction history, fraud review guidelines written in natural language and calibration examples, and outputs only a risk tier, while application code applies the decision rules and handles exceptions and the model does not execute any payment action. According to the post, the first deployment used Claude Opus, and median LLM call latency (P50) was about 1.5 seconds for an input of about 7,000 tokens. The post added that changing the guidelines does not require retraining the model, but every change must be validated against real transaction outcomes.
Results
According to the same post, an online experiment compared a control group using only the existing models and rules with a treatment group that added agent review on 50% of traffic, and the treatment group had 30% fewer fraud transactions and a 22% lower fraud amount. The post said the treatment group had a lower fraud count and fraud rate, while total transactions, users and gross revenue were higher. The post explained that the fraud savings estimated from the experiment were about 3 times the inference cost calculated at the model provider's public pricing. According to the series plan in the post, later parts cover an evaluation in which newer model versions scored lower recall and F1 on the company's fraud benchmark, and a result in which post-training Qwen3.5-9B cut latency by 55%.
Limits and open questions
The post did not disclose the experiment period, the share of transactions sent to agent review, or absolute fraud counts and amounts, and it showed the increases in transactions and revenue only in charts without figures in the text. The 30% and 22% figures are results of an online experiment Coinbase measured itself, and the post does not say whether the agent was extended to all transactions after the experiment.
Sources
Compiled from public sources. These are not results from ATF Works customers.