Part 1/3 Owning Intelligence Series: How Coinbase Built, Evaluated, and Post-Trained a Fraud Agent
TL;DR: How Coinbase built, evaluated, and post-trained an Onramp fraud agent to achieve 9.6% higher performance (F1 score) and 55% lower end to end latency than frontier models.

In This Series
Part 1 — Build: We layer an LLM risk agent onto Onramp’s existing ML model and rules. Online A/B experiment shows 30% fewer fraudulent transactions and 22% less fraud value.
Part 2 — Evaluate: We compared Opus, Sonnet, and GPT releases on our fraud benchmark. Each newer version scored lower on recall and F1, reinforcing the need for domain-specific evaluation.
Part 3 — Own: Using SkyRL on Anyscale, we post-trained Qwen3.5-9B on our proprietary Onramp fraud dataset. The resulting model outperformed Opus 4.5 on all four fraud-detection metrics and delivered 55% lower end-to-end online serving latency.
Part 1 — Build: Reducing Onramp Fraud with an LLM Risk Agent
TL;DR: Coinbase Onramp lets users buy crypto within partner apps, including through guest checkout without a Coinbase account. To detect fraud with limited account history, we built and deployed an LLM-powered risk agent that reviews recent transaction sequences alongside our existing ML model and rules. In an online experiment, the agent-enabled flow recorded 30% fewer fraudulent transactions and 22% less fraud value than the existing stack alone.
The Challenge: Fraud Patterns Across Transactions
Our existing machine-learning model and rules engine provide fast, low-cost screening for guest-checkout payments. The harder problem is recognizing attacks whose signals emerge across several transactions rather than within a single purchase.
Some fraud is visible within a single transaction. Other patterns emerge only when related activity is considered together. What looks routine in isolation may warrant closer review in a broader context. This distinction motivated us to explore an additional layer of contextual reasoning alongside our existing models and rules.
Even without an established account history, a guest’s recent activity can provide useful context. Turning that context into a timely risk decision presents two challenges:
Predefined features capture only part of the sequence. Traditional models can use historical aggregates, such as transaction counts or spending totals. Those features must be designed in advance and may not preserve the combinations of timing, payment methods, and destinations that distinguish an emerging attack.
Confirmed outcomes arrive later. Fraud labels are often delayed, which complicates timely evaluation and adaptation to emerging patterns.
We wanted a way to review recent behavior without engineering a new feature for every pattern. An LLM can compare the current transaction with the supplied history and apply fraud-review guidance expressed in natural language. Updating that guidance does not require retraining the model, although every change still needs validation against transaction outcomes.
The Solution: An Additional Layer of Risk Review
Onramp Service sends transactions to Risk Service, where the existing model and rules provide the first line of defense. We placed the LLM agent downstream of those checks to provide an additional contextual review for selected transactions. The model and agent run within the same Ray Serve deployment.

Figure 1: Ray Serve orchestrates the ML model, rule engine, and LLM agent for transaction risk review. The agent adds contextual review after the existing checks: it can block a transaction that would otherwise be allowed, but cannot override an upstream block.
1. Route Transactions Selectively
We invoke the LLM selectively to balance fraud detection, latency, and cost. The model and rules handle broad screening, while the agent provides an additional layer of contextual review.
2. Give the Agent Structured Context
Each review combines structured transaction context, summaries of relevant historical activity, and domain-specific review guidance. Calibration examples help maintain consistent assessments without requiring the model to retain a conversation across transactions.
The history preserves the sequence of events. The summaries make bursts and changes explicit, without requiring the LLM to calculate them. Together, they let the agent assess not just whether a transaction looks unusual, but how it differs from the activity that preceded it.
The serving layer assembles this context from the current transaction and streaming history supplied with the request. It removes the current transaction from that history to avoid comparing the event against itself, and formats the two separately. Review instructions and learned guidance are loaded at initialization, so each request supplies fresh evidence without requiring the model to maintain a conversation across transactions.
3. Keep the Decision Path Simple and Bounded
The agent is a constrained reviewer, not an open-ended tool-calling workflow. Application code prepares the context, the model returns a risk score, and application code parses that score and applies the decision policy. The model is not responsible for executing payment actions.
We used Claude Opus for the initial deployment. The serving layer keeps model configuration separate from the decision policy, allowing us to evaluate alternative models without redesigning the risk flow. Part 2 examines those comparisons.
To reduce latency, we intentionally limit the agent’s output to a risk-level classification, such as RISK_LEVEL: 5, without a written explanation. With an input prompt of roughly 7,000 tokens, the LLM call had a median (P50) latency of approximately 1.5 seconds.
The agent returns a structured risk assessment. Application code applies the decision policy and handles exceptions within the existing risk-control framework. The agent remains a bounded component of the broader risk stack.
4. Build for Reliability and Observability
Within a batch, eligible reviews run asynchronously rather than waiting for each LLM call in sequence. The orchestrator collects individual call failures without discarding other results in the batch. The client also supports limited retries and a configurable fallback model for provider capacity errors.
The serving response includes the agent’s decision and risk level, the model used, elapsed review time, retry count, and reported token usage. These fields help separate detection behavior from serving behavior: a change in outcomes can be investigated alongside model choice, latency, and provider failures rather than treated as a single unexplained shift in the final decision.
Online Experiment Result
In the online experiment, the control group used the existing ML model and rules engine. The treatment group, which received 50% of traffic, used the same risk stack with an additional layer of selective LLM-agent review. Compared with control, the agent-enabled treatment group showed the following changes in fraud, volume, and revenue:

Figure 2: Relative changes in fraud, transaction volume, user count, and revenue in the online experiment. All metrics share the same scale.
The treatment group recorded lower fraud counts and rates alongside higher total transactions, user count, and total revenue. The fraud-rate metric accounts for differences in transaction volume; the growth metrics shown are aggregate totals, not per-user or per-transaction gains.
These findings support the value of contextual review after the existing checks.
Cost and Operational Trade-offs
Based on the online experiment results, projected savings from fraud prevention were roughly three times the estimated inference cost, calculated using input and output token usage at the model provider’s published rates.
The design makes those economics possible. We reserve the more expensive review for a subset of activity while retaining the speed and low cost of the traditional stack for broad screening.
From Building the Agent to Owning the Evaluation
The online experiment showed that an LLM can strengthen fraud detection without replacing the existing risk stack. By reviewing transactions in context, a narrowly scoped agent caught patterns that passed the first layer of checks. The value came from giving the model a complementary role and keeping its authority explicit.
But success with one model raises another question: does the improvement hold when the model changes? We control the prompt, transaction context, and decision policy, but not the weights of a closed-source model. Changing model versions can shift risk judgments even when the rest of the system stays the same. Stronger performance on general-purpose benchmarks is not enough to tell us whether that shift helps or hurts fraud detection.
The online experiment established that contextual LLM review could strengthen our existing risk stack. It did not tell us which model should power the agent or whether a newer release would improve its decisions. In Part 2, we build the evaluation needed to answer those questions.



