Enterprise Environments

SnorkelLegal

SnorkelLegal evaluates agents within a simulated mid-market law firm, where every request is tied to a matter, a procedural stage, and a role with defined authority.

It evaluates the intersection of long-context reasoning, matter-specific rules, and procedural discipline.

At a glance

200

frontier tasks

15

simulated personas represented

50%

semantic-output tasks

23.5%

tasks requiring state changes

Frontier performance

  • Grok 4.6
  • GLM 5.3
  • Muse Spark 1.3
  • Kimi K3
  • Fable 5.1
  • Grok 4.7
  • Opus 5.5
  • DeepSeek V4 Pro
  • Nemotron 3 Ultra 550B
  • Opus 5
  • Qwen 3.8 Max
  • Gemini Flash 3.8
  • GPT-6 Astra
Grok 4.6 ×
GLM 5.3 ×
Muse Spark 1.3 ×
Kimi K3 ×
Fable 5.1 ×
Grok 4.7 ×
Opus 5.5 ×
DeepSeek V4 Pro ×
Nemotron 3 Ultra 550B ×
Opus 5 ×
Qwen 3.8 Max ×
Gemini Flash 3.8 ×
GPT-6 Astra ×
Loading chart data...

Key takeaways

SnorkelLegal produced the highest Pass@1 results among the enterprise benchmarks, with Grok 4.6 leading at 40.7%. GLM 5.3, Muse Spark 1.3, Kimi K3, and Fable 5.1 formed a tight second group between 35.9% and 37.8%. Grok 4.7 and Opus 5.5 followed at 32.8% and 31.4%, ahead of the remaining qualified models, which ranged from 20.9% to 30.4%.

Leaderboard

Rank Model Pass@1 Pass@5 Cost / Trial Cost / Task
1 Grok 4.6
40.7%
64.6%
$2.22 $11.12
2 GLM 5.3
37.8%
63.6%
$1.37 $6.83
3 Muse Spark 1.3
36.5%
55.5%
$1.45 $7.25
4 Kimi K3
36.1%
65.3%
$1.94 $9.69
5 Fable 5.1
35.9%
60.5%
$5.26 $26.31
6 Grok 4.7
32.8%
53.4%
$3.67 $18.33
7 Opus 5.5
31.4%
51.5%
$2.07 $10.37
8 DeepSeek V4 Pro
30.4%
63.5%
$1.37 $6.84
9 Nemotron 3 Ultra 550B
29.4%
57.1%
$0.4 $2
10 Opus 5
26.4%
50%
$4.14 $20.69
11 Qwen 3.8 Max
25.6%
53.5%
$2.34 $11.69
12 Gemini Flash 3.8
25.3%
42%
$1.23 $6.16
13 GPT-6 Astra
20.9%
35.1%
$11.61 $58.04

Methodology

Evaluator

Short exact answers and record changes are verified programmatically. Semantic legal work products are evaluated against scenario-specific rubrics for controlling rules, matter-specific facts, scope, and appropriate caveats. State, confidentiality and safety, process, and efficiency are measured separately.
Timeout
The evaluation uses a fixed time and step budget for every run. Long-context tasks are therefore judged within a defined operating envelope rather than against unlimited retries. 
Integration

Each task runs inside a reproducible Docker/OpenEnv workspace with matter records, document evidence, role-gated MCP tools, and a simulated requester. The session preserves matter state and captures the complete trajectory, including retrieval, authority checks, user interaction, and governed writes. Seeded randomness supports repeatable failure-mode testing.

scoring note

Fluent legal prose is not enough. A task is incorrect when the agent relies on superseded evidence, skips a required authority or conflict check, makes an impermissible write, or gives a semantically wrong answer that merely sounds plausible. The score combines outcome gates with process and efficiency diagnostics while allowing valid alternative tool sequences.

Behind the benchmark

Legal work is not just retrieval plus prose. The operative rule may be firm-private or jurisdiction-specific. A matter file may place governing text next to superseded drafts and withdrawn analyses. Authority and confidentiality can determine whether an otherwise correct action is allowed.

SnorkelLegal turns those constraints into observable tests of grounding, sequencing, and restraint. The subset includes semantic legal outputs, long-context tasks, and governed matter updates. It asks a practical question for model developers: can an agent move a matter forward without inventing certainty, using the wrong version, or bypassing a control?

Tasks and rubrics are authored and reviewed by legal-domain experts. Work spans matter intake, conflicts, docketing, litigation holds, discovery, privilege, settlement authority, transactional exposure, and regulatory investigations — requiring agents to locate the correct matter, find the controlling document, and distinguish current evidence from superseded material.

More benchmarks

Software Engineering

SWE-bench CLI

Evaluates validated multi-file code changes in open-source repositories.

By pass@1
1
Image
Opus 5.5
16.6%
2
Image
Fable 5.1
14.5%
3
Image
Opus 5
14.0%
Enterprise Environments

SnorkelRevOps

Evaluates revenue-operations workflows involving systems, controls, and business-state updates.

By pass@1
1
Image
Grok 4.6
15.8%
2
Image
Fable 5.1
14.0%
3
Image
Kimi K3
13.1%
Enterprise Environments

SnorkelFinance 2.0

Evaluates financial workflows requiring evidence gathering, analysis, and compliance.

By pass@1
1
Image
Grok 4.6
25.1%
2
Image
GLM 5.3
21.2%
3
Image
Grok 4.7
20.2%
Enterprise Environments

SnorkelUnderwrite 2.0

Evaluates auditable insurance underwriting decisions using incomplete, distributed evidence.

By pass@1
1
Image
Grok 4.7
30.5%
2
Image
GLM 5.3
30.4%
3
Image
Nemotron 3 Ultra 550B
30.0%
Agentic Coding

Agentic Coding 2.0

Evaluating whether coding agents can plan, execute, verify, and recover across complex terminal-native engineering tasks.

By pass@1
1
Image
GPT-6 Astra
47.6%
2
Image
Opus 5.5
40.4%
3
Image
Fable 5.1
39.6%
Enterprise Environments

SnorkelManufacturing

Evaluates manufacturing decisions using fragmented plant-floor, engineering, and supplier evidence.

By pass@1
1
Image
Grok 4.6
16.4%
2
Image
Grok 4.7
12.1%
3
Image
GLM 5.3
11.9%
Knowledge Work

WorkplaceAgents

Evaluates AI agents on economically valuable professional work and verifiable workplace deliverables.

By pass@1
1
Image
Grok 4.6
17.1%
2
Image
Fable 5.1
15.8%
3
Image
Grok 4.7
15.2%
Software Engineering

LibraryDesignBench

Measures how well AI agents design libraries that other agents can use.

Pass Rate
1
Image
Opus 5.5
48.9%
2
Image
Fable 5.1
47.5%
3
Image
GPT-6 Astra (mini-SWE)
45.1%
Agentic Coding

Terminal-Bench 4.0

Evaluates terminal agents on continuously updated software-engineering tasks.

By Resolution Rate
1
Image
GPT-6 Astra
58.2%
2
Image
Fable 5.1
57.9%
3
Image
Opus 5
53.9%
Science & Research

Terminal-Bench-Science

Evaluates agents on scientific workflows derived from researchers’ own work.

By Resolution Rate
1
Image
GPT-6 Astra
68.1%
2
Image
Opus 5.5
63.3%
3
Image
Fable 5.1
40.0%
Agentic Coding

Terminal-Bench 3.0

Evaluates terminal agents on containerized tasks across seven domains.

By Resolution Rate
1
Image
Opus 5
42.7%
2
Image
GPT-5.6 Sol
34.6%
3
Image
Fable 5
34.1%
Software Engineering

Senior SWE-bench

Evaluates coding agents on senior-level software engineering tasks.

Tasteful Solve Rate
1
Image
Fable 5.1
34.7%
2
Image
Fable 5
34.7%
3
Image
Opus 5
34.7%
Computer Use

OSWorld 2.0

Evaluates computer-use agents on 108 long-horizon workflows.

By Binary Accuracy (500 steps)
1
Image
Opus 5 · max
44.33%
2
Image
Opus 5 · high
36.89%
3
Image
Opus 5 · xhigh
33.33%
Knowledge Work

Agents’ Last Exam

Evaluating AI agents on long-horizon, economically valuable professional workflows with verifiable outcomes.

By Pass Rate
1
Image
Opus 5.5
38.2%
2
Image
GPT-6 Astra
34.2%
3
Image
GPT-6 Sol
32.2%
Software Engineering

SlopCode Bench

Measures coding-agent performance across evolving software requirements.

By Iso Solve
1
Image
GPT-5.5
28.06%
2
Image
GPT-5.3
26.02%
3
Image
GPT-5.4
23.47%
Capability/Efficiency

Continual Learning Bench

Measures improvement across sequential, stateful tasks.

By Agg. Reward
1
Image
Sonnet 4.6
+0.196
2
Image
GPT-5.4
+0.189
3
Image
Sonnet 4.6
+0.185
of
 Illution Back
Illution Front

For models that need to be right. Not just good enough.