How it owrks Banner
Research-led data development

Datasets and environments that give frontier models domain expertise

Snorkel builds the human expert-authored datasets, evaluation environments, and benchmarks calibrated to push the limits of frontier model capability. Off-the-shelf or custom.

Where generic data runs out

Frontier model development stalls on data problems generic pipelines weren't built to solve, including distributional gaps in specialized domains, benchmark blind spots, and failure modes that only surface at scale. We build the data to solve them.

get started

Two ways to get the data you need

The data frontier models need most is rarely the data that already exists. Snorkel delivers it two ways: off-the-shelf for well-defined task areas, or custom-built for the gaps only you can see.

Snorkel Data Series icon
Snorkel Data Series

Curriculum-structured datasets for the task areas frontier models are pushing hardest, with rubrics, reviewer guidance, difficulty tiers, and eval slices built in.

Custom data development Icon
Custom data development

Bespoke datasets, evaluation environments, and benchmark expansions to target the exact failure surface you're trying to close.

SNORKEL DATA SERIES

Datasets and environments

Readily available datasets developed in close collaboration with leading frontier AI teams – curriculum-structured to build difficulty progressively across a task area, with the evaluation infrastructure to match.
Alignment for better code generation icon
Terminal Coding
Long-horizon agents in real containerized terminals — planning, execution, and error recovery.
Terminal-Bench 2.0 & 3.0
AI voice assistant training data for a tech industry giant icon
Software Engineering
Debugging, codebase understanding, and multi-file changes across 7+ languages.

SWE-Bench Pro, Senior SWE-bench

Icon Chair
Enterprise & Workplace Environments
Multi-turn, tool-rich professional workflows across industries, policies, and 100+ occupations.
τ²-bench, τ³-bench, GDPval
Image
Computer Use
GUI interaction and desktop workflow execution across real applications.
OSWorld 2.0, Agent’s Last Exam
Image
Scientific & Research Workflows
Research execution and technical reasoning across scientific domains.
Terminal-Bench-Science, PaperBench
The frontiers of multi-turn math reasoning  icon
STEM Knowledge & Reasoning
Expert scientific problems requiring multi-step reasoning.
Humanity’s Last Exam, FrontierMath

Featured data series

Terminal coding

Senior SWE-bench+

Real-world PR tasks at a senior engineer’s bar
— inferring intent, runtime debugging, and shipping code that fits the codebase. 

Led benchmark development

Enterprise agents

GDPval+

Economically-grounded professional
tasks spanning all 20 O*NET sectors
and 100+ occupations.

Terminal coding

Frontier-Bench+

Multi-step terminal engineering tasks, calibrated to the Frontier-Bench standard.

Core benchmark contributor and reviewer

Image
For the full Snorkel data catalog, talk to our team.
Custom data development image
CUSTOM DATA DEVELOPMENT

Closing gaps existing datasets can’t reach

Custom data development engagements start with the failure surface: what the model can't do, where it's brittle, and what the correct evaluation criteria are. From there, Snorkel builds the datasets, environments, and benchmark expansions needed to close it.

01
Task specification and rubric design
02
Bespoke dataset construction
03
RL environment development
04
Benchmark and eval expansion
05
Provenance and adjudication
PUBLISHED RESEARCH

Research-validated methodology

Our datasets are built using research-backed methods for benchmark design, evaluator calibration, and failure analysis.

For models that need to be right. Not just good enough.