Enterprise Environments

SnorkelRevOps

SnorkelRevOps measures agents inside a simulated B2B SaaS revenue function, where work crosses CRM, CPQ, billing, compensation, marketing, usage, finance, and governance records — each system authoritative for different decisions.

The set concentrates on cross-system, control-sensitive tasks most likely to separate fluent assistants from reliable operators.

At a glance

200

frontier tasks

17

simulated personas represented

54%

semantic-output tasks

35%

tasks requiring state changes

Frontier performance

  • Grok 4.6
  • Fable 5.1
  • Kimi K3
  • Muse Spark 1.3
  • Opus 5.5
  • Opus 5
  • GLM 5.3
  • Grok 4.7
  • GPT-6 Astra
  • Nemotron 3 Ultra 550B
  • DeepSeek V4 Pro
  • Qwen 3.8 Max
  • Gemini Flash 3.8
Grok 4.6 ×
Fable 5.1 ×
Kimi K3 ×
Muse Spark 1.3 ×
Opus 5.5 ×
Opus 5 ×
GLM 5.3 ×
Grok 4.7 ×
GPT-6 Astra ×
Nemotron 3 Ultra 550B ×
DeepSeek V4 Pro ×
Qwen 3.8 Max ×
Gemini Flash 3.8 ×
Loading chart data...

Key takeaways

Grok 4.6 led SnorkelRevOps at 15.8%, followed by Fable 5.1 at 14.0% and Kimi K3 at 13.1%. Eight additional models clustered between 9.5% and 12.6%, including Opus 5.5 at 12.5% and Grok 4.7 at 11.3%, while Qwen 3.8 Max and Gemini Flash 3.8 remained below 7%.

State-pass rates were tightly grouped between 63.3% and 70.4%, but answer-pass rates ranged from 9.7% to 29.4%. Final-answer completion therefore created considerably more separation between models than state execution.

Leaderboard

Rank Model Pass@1 Pass@5 Cost / Trial Cost / Task
1 Grok 4.6
15.8%
35%
$1.94 $9.71
2 Fable 5.1
14%
25.1%
$4.03 $20.17
3 Kimi K3
13.1%
31.5%
$1.77 $8.85
4 Muse Spark 1.3
12.6%
25.5%
$1.07 $5.37
5 Opus 5.5
12.5%
25.8%
$2.47 $12.13
6 Opus 5
12.2%
23.9%
$4.02 $20.11
7 GLM 5.3
11.8%
25.7%
$1.27 $6.36
8 Grok 4.7
11.3%
28.1%
$2.17 $10.86
9 GPT-6 Astra
10.9%
21.1%
$5.92 $29.59
10 Nemotron 3 Ultra 550B
10.1%
28%
$0.34 $1.68
11 DeepSeek V4 Pro
9.5%
24.1%
$0.78 $3.91
12 Qwen 3.8 Max
6.5%
17.7%
$1.87 $9.37
13 Gemini Flash 3.8
5.4%
14.4%
$0.93 $4.66
Image

Want to evaluate your model against SnorkelRevOps? Talk to our team

Methodology

Evaluator

Numeric and structured outputs are verified programmatically. Memos, forecasts, adjudications, and other semantic outputs are evaluated for meaning against scenario-specific rubrics. State hashes verify required updates. Additional judges assess policy-safe communication, tool use, user interaction, and efficiency.
timeout
A fixed run budget governs time and agent steps so multi-tool execution can be compared across models. Results should be reported with the timeout and step configuration used for the evaluation.
integration

Models operate through a stateful OpenEnv session backed by a reproducible Docker environment. Structured MCP calls connect the relevant revenue systems and document corpus. Simulated users control disclosure, tone, authority, and escalation. Every user exchange, tool call, observation, and state change is logged.

scoring note

Task correctness is gated on both the final answer and the required environment state. A model can be fluent and still fail if it cites the wrong source of truth, approves a prohibited exception, or executes a plan that diverges from the intended change. Valid alternate tool paths are not penalized. Redundant calls, tool errors, unsafe behavior, and plan-versus-action mismatches remain visible in diagnostics.

Behind the benchmark

RevOps is a key area of business where a numerically plausible answer can still cause operational harm. A discount can violate floor price. A forecast can mix CRM and billing numbers. A territory or compensation change can be valid in one system and wrong in another.

The benchmark requires agents to resolve source-of-truth, authority, and approval constraints before acting. It also tests whether they can manage incomplete information and human pressure without substituting generic business knowledge for the controls encoded in the environment.

The high semantic-output share measures judgment-quality communication, while state-changing tasks test operational follow-through. Simulated personas add the disclosure differences, escalation behavior, and authority boundaries that make revenue workflows difficult in practice.

Requests cover forecasting, pricing and deal desk, territory and compensation, renewals, and revenue reconciliation. Agents must anchor to the right account or opportunity, reconcile competing records, and either produce a decision-quality deliverable or execute the required update.

More benchmarks

Software Engineering

SWE-bench CLI

Evaluates validated multi-file code changes in open-source repositories.

By pass@1
1
Image
Opus 5.5
16.6%
2
Image
Fable 5.1
14.5%
3
Image
Opus 5
14.0%
Enterprise Environments

SnorkelFinance 2.0

Evaluates financial workflows requiring evidence gathering, analysis, and compliance.

By pass@1
1
Image
Grok 4.6
25.1%
2
Image
GLM 5.3
21.2%
3
Image
Grok 4.7
20.2%
Enterprise Environments

SnorkelUnderwrite 2.0

Evaluates auditable insurance underwriting decisions using incomplete, distributed evidence.

By pass@1
1
Image
Grok 4.7
30.5%
2
Image
GLM 5.3
30.4%
3
Image
Nemotron 3 Ultra 550B
30.0%
Agentic Coding

Agentic Coding 2.0

Evaluating whether coding agents can plan, execute, verify, and recover across complex terminal-native engineering tasks.

By pass@1
1
Image
GPT-6 Astra
47.6%
2
Image
Opus 5.5
40.4%
3
Image
Fable 5.1
39.6%
Enterprise Environments

SnorkelManufacturing

Evaluates manufacturing decisions using fragmented plant-floor, engineering, and supplier evidence.

By pass@1
1
Image
Grok 4.6
16.4%
2
Image
Grok 4.7
12.1%
3
Image
GLM 5.3
11.9%
Knowledge Work

WorkplaceAgents

Evaluates AI agents on economically valuable professional work and verifiable workplace deliverables.

By pass@1
1
Image
Grok 4.6
17.1%
2
Image
Fable 5.1
15.8%
3
Image
Grok 4.7
15.2%
Enterprise Environments

SnorkelLegal

Evaluates legal workflows requiring evidence, procedure, authority, and action.

By pass@1
1
Image
Grok 4.6
40.7%
2
Image
GLM 5.3
37.8%
3
Image
Muse Spark 1.3
36.5%
Software Engineering

LibraryDesignBench

Measures how well AI agents design libraries that other agents can use.

Pass Rate
1
Image
Opus 5.5
48.9%
2
Image
Fable 5.1
47.5%
3
Image
GPT-6 Astra (mini-SWE)
45.1%
Agentic Coding

Terminal-Bench 4.0

Evaluates terminal agents on continuously updated software-engineering tasks.

By Resolution Rate
1
Image
GPT-6 Astra
58.2%
2
Image
Fable 5.1
57.9%
3
Image
Opus 5
53.9%
Science & Research

Terminal-Bench-Science

Evaluates agents on scientific workflows derived from researchers’ own work.

By Resolution Rate
1
Image
GPT-6 Astra
68.1%
2
Image
Opus 5.5
63.3%
3
Image
Fable 5.1
40.0%
Agentic Coding

Terminal-Bench 3.0

Evaluates terminal agents on containerized tasks across seven domains.

By Resolution Rate
1
Image
Opus 5
42.7%
2
Image
GPT-5.6 Sol
34.6%
3
Image
Fable 5
34.1%
Software Engineering

Senior SWE-bench

Evaluates coding agents on senior-level software engineering tasks.

Tasteful Solve Rate
1
Image
Fable 5.1
34.7%
2
Image
Fable 5
34.7%
3
Image
Opus 5
34.7%
Computer Use

OSWorld 2.0

Evaluates computer-use agents on 108 long-horizon workflows.

By Binary Accuracy (500 steps)
1
Image
Opus 5 · max
44.33%
2
Image
Opus 5 · high
36.89%
3
Image
Opus 5 · xhigh
33.33%
Knowledge Work

Agents’ Last Exam

Evaluating AI agents on long-horizon, economically valuable professional workflows with verifiable outcomes.

By Pass Rate
1
Image
Claude Opus 5.5
38.2%
2
Image
GPT-6 Astra
34.2%
3
Image
GPT-6 Sol
32.2%
Software Engineering

SlopCode Bench

Measures coding-agent performance across evolving software requirements.

By Iso Solve
1
Image
GPT-5.5
28.06%
2
Image
GPT-5.3
26.02%
3
Image
GPT-5.4
23.47%
Capability/Efficiency

Continual Learning Bench

Measures improvement across sequential, stateful tasks.

By Agg. Reward
1
Image
Sonnet 4.6
+0.196
2
Image
GPT-5.4
+0.189
3
Image
Sonnet 4.6
+0.185
of
 Illution Back
Illution Front

For models that need to be right. Not just good enough.