SnorkelLegal
SnorkelLegal evaluates agents within a simulated mid-market law firm, where every request is tied to a matter, a procedural stage, and a role with defined authority.
It evaluates the intersection of long-context reasoning, matter-specific rules, and procedural discipline.
At a glance
200
frontier tasks
15
simulated personas represented
50%
semantic-output tasks
23.5%
tasks requiring state changes
Frontier performance
- Grok 4.6
- GLM 5.3
- Muse Spark 1.3
- Kimi K3
- Fable 5.1
- Grok 4.7
- Opus 5.5
- DeepSeek V4 Pro
- Nemotron 3 Ultra 550B
- Opus 5
- Qwen 3.8 Max
- Gemini Flash 3.8
- GPT-6 Astra
Key takeaways
SnorkelLegal produced the highest Pass@1 results among the enterprise benchmarks, with Grok 4.6 leading at 40.7%. GLM 5.3, Muse Spark 1.3, Kimi K3, and Fable 5.1 formed a tight second group between 35.9% and 37.8%. Grok 4.7 and Opus 5.5 followed at 32.8% and 31.4%, ahead of the remaining qualified models, which ranged from 20.9% to 30.4%.
Leaderboard
| Rank | Model | Pass@1 | Pass@5 | Cost / Trial | Cost / Task |
|---|---|---|---|---|---|
| 1 | Grok 4.6 |
40.7%
|
64.6%
|
$2.22 | $11.12 |
| 2 | GLM 5.3 |
37.8%
|
63.6%
|
$1.37 | $6.83 |
| 3 | Muse Spark 1.3 |
36.5%
|
55.5%
|
$1.45 | $7.25 |
| 4 | Kimi K3 |
36.1%
|
65.3%
|
$1.94 | $9.69 |
| 5 | Fable 5.1 |
35.9%
|
60.5%
|
$5.26 | $26.31 |
| 6 | Grok 4.7 |
32.8%
|
53.4%
|
$3.67 | $18.33 |
| 7 | Opus 5.5 |
31.4%
|
51.5%
|
$2.07 | $10.37 |
| 8 | DeepSeek V4 Pro |
30.4%
|
63.5%
|
$1.37 | $6.84 |
| 9 | Nemotron 3 Ultra 550B |
29.4%
|
57.1%
|
$0.4 | $2 |
| 10 | Opus 5 |
26.4%
|
50%
|
$4.14 | $20.69 |
| 11 | Qwen 3.8 Max |
25.6%
|
53.5%
|
$2.34 | $11.69 |
| 12 | Gemini Flash 3.8 |
25.3%
|
42%
|
$1.23 | $6.16 |
| 13 | GPT-6 Astra |
20.9%
|
35.1%
|
$11.61 | $58.04 |
Methodology
Evaluator
Each task runs inside a reproducible Docker/OpenEnv workspace with matter records, document evidence, role-gated MCP tools, and a simulated requester. The session preserves matter state and captures the complete trajectory, including retrieval, authority checks, user interaction, and governed writes. Seeded randomness supports repeatable failure-mode testing.
scoring note
Fluent legal prose is not enough. A task is incorrect when the agent relies on superseded evidence, skips a required authority or conflict check, makes an impermissible write, or gives a semantically wrong answer that merely sounds plausible. The score combines outcome gates with process and efficiency diagnostics while allowing valid alternative tool sequences.
Behind the benchmark
Legal work is not just retrieval plus prose. The operative rule may be firm-private or jurisdiction-specific. A matter file may place governing text next to superseded drafts and withdrawn analyses. Authority and confidentiality can determine whether an otherwise correct action is allowed.
SnorkelLegal turns those constraints into observable tests of grounding, sequencing, and restraint. The subset includes semantic legal outputs, long-context tasks, and governed matter updates. It asks a practical question for model developers: can an agent move a matter forward without inventing certainty, using the wrong version, or bypassing a control?
Tasks and rubrics are authored and reviewed by legal-domain experts. Work spans matter intake, conflicts, docketing, litigation holds, discovery, privilege, settlement authority, transactional exposure, and regulatory investigations — requiring agents to locate the correct matter, find the controlling document, and distinguish current evidence from superseded material.

