RESOURCES

Blog

Ideas, updates, and practical guidance from the Snorkel team.

Image for Data 2.0 and the research era of AI data

Data 2.0 and the research era of AI data

Announcing Snorkel’s $350M fundraise to build the frontier lab for AI data.

September 22, 2026
All articles
Sort: Newest
MedPAIR: Measuring Whether Physicians and AI Agree on What Matters in Medical QA
NEW
MedPAIR: Measuring Whether Physicians and AI Agree on What Matters in Medical QA

Yuexing Hao presents MedPAIR, a dataset comparing which sentences physicians and LLMs find relevant in clinical questions. Humans and LLMs agree on only 50 to 60% of relevance labels.

Oct 02, 2026 •
Learn more about MedPAIR: Measuring Whether Physicians and AI Agree on What Matters in Medical QA
RL environments for LLM agents: Design, rewards, and validation
NEW
RL environments for LLM agents: Design, rewards, and validation

TLDR: An agent, by definition, can take actions and is more than just a language model. The model is only one part of the system, and is only one element that you are training. The environment determines what the agent can observe, what it can change, which actions are available, and what behavior receives a reward. For a coding agent,…

Sep 28, 2026 •
Learn more about RL environments for LLM agents: Design, rewards, and validation
Opus 5.5 vs Opus 5 vs Fable 5.1: Coding Benchmark Results
NEW
Opus 5.5 vs Opus 5 vs Fable 5.1: Coding Benchmark Results

We’ve now run three generations of frontier models through the same expert-created terminal-bench style task set: Fable 5.1 in September, Opus 5 in July, and, now Opus 5.5. For this analysis, the task set contained 200 trajectories, of 24 tasks, with every failure traced to a judge-confirmed root cause. A generational climb On the task set, pass@1 is 61.5% for…

Sep 23, 2026 •
Learn more about Opus 5.5 vs Opus 5 vs Fable 5.1: Coding Benchmark Results
Data 2.0 and the research era of AI data
Data 2.0 and the research era of AI data

Today, I’m excited to announce Snorkel’s $350M Series E financing at a $3.5B valuation, led by Insight and S32, with participation from Third Point, March, Blumberg, Allegis, Standard VC, Frontline, and existing investors Addition, Lightspeed, Greylock, GV, P7, Wells Fargo, Walden Catalyst Ventures, and Factory.

Sep 22, 2026 •
Learn more about Data 2.0 and the research era of AI data
Grok 4.7 on Senior SWE-Bench: Strong pass@3, Cheaper Cost per Trial
Grok 4.7 on Senior SWE-Bench: Strong pass@3, Cheaper Cost per Trial

Grok 4.7 was evaluated on Senior SWE-Bench, our benchmark for measuring whether coding agents work like senior engineers. It reaches a tasteful pass@3 of 40.0%, up from 38.9% for Grok 4.6, ranking sixth overall at $0.24 per trial, which is a fraction of the cost compared to Opus 4.8. Benchmark Senior SWE-Bench evaluates agents as compared to a senior engineer….

Sep 22, 2026 •
Learn more about Grok 4.7 on Senior SWE-Bench: Strong pass@3, Cheaper Cost per Trial
From Foundational Competency to Expert Performance: A Curriculum Approach to Model Development
From Foundational Competency to Expert Performance: A Curriculum Approach to Model Development

A student does not go from 1st to 12th grade in a single step. Each grade builds on a specific set of skills, and each one assumes the previous skills have already been mastered. Nobody learns calculus without algebra. When a student skips ahead anyway, what they end up with is memorization rather than understanding. The gaps show up later, usually in ways that are harder to trace back and fix.

Sep 15, 2026 •
Learn more about From Foundational Competency to Expert Performance: A Curriculum Approach to Model Development
OSWorld 2.0: Why Long-Horizon Computer-Use Agents Still Fail Four Out of Five Tasks
OSWorld 2.0: Why Long-Horizon Computer-Use Agents Still Fail Four Out of Five Tasks

Mengqi Yuan (XLANG Lab, University of Hong Kong) presents OSWorld 2.0, a benchmark of 108 long-horizon, real-world computer-use workflows where even frontier AI agents complete only 20.6% of tasks outright after 300+ steps each.

Sep 03, 2026 •
Learn more about OSWorld 2.0: Why Long-Horizon Computer-Use Agents Still Fail Four Out of Five Tasks
Fable 5.1 on Frontier Coding Tasks: Efficient Successes, Distinct Failure Modes
Fable 5.1 on Frontier Coding Tasks: Efficient Successes, Distinct Failure Modes

We evaluated Fable 5.1 on a series of frontier coding tasks from our proprietary Terminal-Bench+ dataset and compared the results against Opus 5. Fable remained competitive across most categories and was materially more efficient on successful runs, while its gap was concentrated in a small set of terminal-heavy and build/dependency tasks. Because category sizes are small and uneven, we treat…

Sep 01, 2026 •
Learn more about Fable 5.1 on Frontier Coding Tasks: Efficient Successes, Distinct Failure Modes
Terminal-Bench 4.0: Why Continuous Benchmarks Require Continuous QA
Terminal-Bench 4.0: Why Continuous Benchmarks Require Continuous QA

The speed of new frontier model releases keeps accelerating. Meanwhile benchmarks struggle to keep up and saturate quickly, often being left in the dust. Most benchmarks are static datasets with no active maintenance, causing them to lose value fast. Some benchmarks are looking to change this by becoming Continuous Benchmarks. Terminal-Bench is one of the most widely reported benchmarks on…

Aug 28, 2026 •
Learn more about Terminal-Bench 4.0: Why Continuous Benchmarks Require Continuous QA
1 2 … 40
Image

Join our newsletter

For expert advice, the latest research, and exclusive events.
By submitting this form, I acknowledge I will receive email updates from Snorkel AI, and I agree to the Terms of Use and acknowledge that my information will be used in accordance with the Privacy Policy.