SnorkelRevOps
SnorkelRevOps measures agents inside a simulated B2B SaaS revenue function, where work crosses CRM, CPQ, billing, compensation, marketing, usage, finance, and governance records — each system authoritative for different decisions.
The set concentrates on cross-system, control-sensitive tasks most likely to separate fluent assistants from reliable operators.
At a glance
200
frontier tasks
17
simulated personas represented
54%
semantic-output tasks
35%
tasks requiring state changes
Frontier performance
- Grok 4.6
- Fable 5.1
- Kimi K3
- Muse Spark 1.3
- Opus 5.5
- Opus 5
- GLM 5.3
- Grok 4.7
- GPT-6 Astra
- Nemotron 3 Ultra 550B
- DeepSeek V4 Pro
- Qwen 3.8 Max
- Gemini Flash 3.8
Key takeaways
Grok 4.6 led SnorkelRevOps at 15.8%, followed by Fable 5.1 at 14.0% and Kimi K3 at 13.1%. Eight additional models clustered between 9.5% and 12.6%, including Opus 5.5 at 12.5% and Grok 4.7 at 11.3%, while Qwen 3.8 Max and Gemini Flash 3.8 remained below 7%.
State-pass rates were tightly grouped between 63.3% and 70.4%, but answer-pass rates ranged from 9.7% to 29.4%. Final-answer completion therefore created considerably more separation between models than state execution.
Leaderboard
| Rank | Model | Pass@1 | Pass@5 | Cost / Trial | Cost / Task |
|---|---|---|---|---|---|
| 1 | Grok 4.6 |
15.8%
|
35%
|
$1.94 | $9.71 |
| 2 | Fable 5.1 |
14%
|
25.1%
|
$4.03 | $20.17 |
| 3 | Kimi K3 |
13.1%
|
31.5%
|
$1.77 | $8.85 |
| 4 | Muse Spark 1.3 |
12.6%
|
25.5%
|
$1.07 | $5.37 |
| 5 | Opus 5.5 |
12.5%
|
25.8%
|
$2.47 | $12.13 |
| 6 | Opus 5 |
12.2%
|
23.9%
|
$4.02 | $20.11 |
| 7 | GLM 5.3 |
11.8%
|
25.7%
|
$1.27 | $6.36 |
| 8 | Grok 4.7 |
11.3%
|
28.1%
|
$2.17 | $10.86 |
| 9 | GPT-6 Astra |
10.9%
|
21.1%
|
$5.92 | $29.59 |
| 10 | Nemotron 3 Ultra 550B |
10.1%
|
28%
|
$0.34 | $1.68 |
| 11 | DeepSeek V4 Pro |
9.5%
|
24.1%
|
$0.78 | $3.91 |
| 12 | Qwen 3.8 Max |
6.5%
|
17.7%
|
$1.87 | $9.37 |
| 13 | Gemini Flash 3.8 |
5.4%
|
14.4%
|
$0.93 | $4.66 |
Want to evaluate your model against SnorkelRevOps? Talk to our team
Methodology
Evaluator
Models operate through a stateful OpenEnv session backed by a reproducible Docker environment. Structured MCP calls connect the relevant revenue systems and document corpus. Simulated users control disclosure, tone, authority, and escalation. Every user exchange, tool call, observation, and state change is logged.
scoring note
Task correctness is gated on both the final answer and the required environment state. A model can be fluent and still fail if it cites the wrong source of truth, approves a prohibited exception, or executes a plan that diverges from the intended change. Valid alternate tool paths are not penalized. Redundant calls, tool errors, unsafe behavior, and plan-versus-action mismatches remain visible in diagnostics.
Behind the benchmark
RevOps is a key area of business where a numerically plausible answer can still cause operational harm. A discount can violate floor price. A forecast can mix CRM and billing numbers. A territory or compensation change can be valid in one system and wrong in another.
The benchmark requires agents to resolve source-of-truth, authority, and approval constraints before acting. It also tests whether they can manage incomplete information and human pressure without substituting generic business knowledge for the controls encoded in the environment.
The high semantic-output share measures judgment-quality communication, while state-changing tasks test operational follow-through. Simulated personas add the disclosure differences, escalation behavior, and authority boundaries that make revenue workflows difficult in practice.
Requests cover forecasting, pricing and deal desk, territory and compensation, renewals, and revenue reconciliation. Agents must anchor to the right account or opportunity, reconcile competing records, and either produce a decision-quality deliverable or execute the required update.

