Finance / Crypto exchange
Coinbase: Feeding the order of user actions into fraud detection
- Company
- Coinbase
- Country
- United States
- Adoption stage
- In operation
- Source published
- Date basis
- The date the source was published. It can differ from the date adoption started.
- How the source was checked
- Read the full source text
The work problem
According to a Coinbase engineering blog post, Coinbase uses machine learning for many tasks, from notifications to fraud detection, and its ML Platform manages more than 5,000 hand-crafted features. The same post explained that these features are long-running aggregations built by domain experts and refreshed daily in batch or in near real time through streaming. The same post said that a feature such as the number of digital asset purchases in the last 30 days leaves out which assets were bought, how much and when. The same post explained that building more features to capture that information takes a long time and the performance gain is uncertain.
Technology and data
According to the same post, Coinbase built a sequence feature framework on top of Tecton and Databricks Spark that feeds user actions such as sign-ins, buys, sells and crypto sends into models in time order in place of hand-crafted features. The same post listed shared requirements of fewer than 100 event types per sequence, up to 1,000 recent events at read time, freshness within 1 to 2 seconds and online read latency under 100ms at p99. According to the same post, the framework is designed so that machine learning engineers declare only the Kafka topic, the event schema and the sequence conditions, and it then creates the data sources and Tecton feature views automatically with observability built in. The same post explained that hundreds of Kafka topics hold thousands of event types, each with its own schema, so each event used for machine learning is registered explicitly in the ML Event Registry, and that streaming reads come from Kafka while batch jobs and backfills read from Delta Lake, all converted into a shared schema of user_id, event_name, timestamp and metadata. According to the same post, a continuous mode with a processing interval of 0 seconds writes each event to DynamoDB as soon as it arrives, and a user's events are assembled into the full sequence at read time. Engineers decide which events and time windows go in, and the same post explained that deep learning models such as Transformers and LSTMs can learn directly from these sequences.
Results
According to the same post, one of the most demanding streaming jobs receives about 2,000 events per second with an average end-to-end latency under 500ms. The same post said that after moving to Databricks Single Node clusters, which run the Spark driver and worker on one machine, freshness for all streaming jobs stayed the same or improved while CPUs and compute cost fell by 20 to 40%. The same post said early results improved fraud detection and recommendation models and delivered tens of millions of dollars in impact on key business metrics over the past year. According to the same post, several sequence features rank in the top 10 of global feature importance in many critical models.
Limits and open questions
The same post does not break the tens of millions of dollars in impact down by model or metric, and it does not separately state the reduction in fraud losses. A search tool excerpt showed a customer story about a later move to Databricks real-time processing, but that source was not opened in this research.
Sources
- How Coinbase Builds Sequence Features for Machine Learningcoinbase.com, Accessed
Compiled from public sources. These are not results from ATF Works customers.