RESOURCES

Blog

Ideas, updates, and practical guidance from the Snorkel team.

Image for Data 2.0 and the research era of AI data

Data 2.0 and the research era of AI data

Announcing Snorkel’s $350M fundraise to build the frontier lab for AI data.

September 22, 2026
All articles
Sort: Newest
Frontier AI is Accelerating. Open Benchmarks Need to Keep Up.
NEW
Frontier AI is Accelerating. Open Benchmarks Need to Keep Up.

Today, we’re excited to announce that we are expanding our Open Benchmarks Grants by 10x to a $30M commitment, to support researchers, domain experts, and open-source teams pushing the frontier of robust, open frontier AI evaluation. The expanded program includes: Through our initial $3M commitment to Open Benchmarks Grants (OBG) and our broader research collaborations, we’ve partnered with the teams…

Oct 07, 2026 •
Learn more about Frontier AI is Accelerating. Open Benchmarks Need to Keep Up.
MedPAIR: Measuring Whether Physicians and AI Agree on What Matters in Medical QA
NEW
MedPAIR: Measuring Whether Physicians and AI Agree on What Matters in Medical QA

Yuexing Hao presents MedPAIR, a dataset comparing which sentences physicians and LLMs find relevant in clinical questions. Humans and LLMs agree on only 50 to 60% of relevance labels.

Oct 02, 2026 •
Learn more about MedPAIR: Measuring Whether Physicians and AI Agree on What Matters in Medical QA
RL environments for LLM agents: Design, rewards, and validation
RL environments for LLM agents: Design, rewards, and validation

TLDR: An agent, by definition, can take actions and is more than just a language model. The model is only one part of the system, and is only one element that you are training. The environment determines what the agent can observe, what it can change, which actions are available, and what behavior receives a reward. For a coding agent,…

Sep 28, 2026 •
Learn more about RL environments for LLM agents: Design, rewards, and validation
Opus 5.5 vs Opus 5 vs Fable 5.1: Coding Benchmark Results
Opus 5.5 vs Opus 5 vs Fable 5.1: Coding Benchmark Results

We’ve now run three generations of frontier models through the same expert-created terminal-bench style task set: Fable 5.1 in September, Opus 5 in July, and, now Opus 5.5. For this analysis, the task set contained 200 trajectories, of 24 tasks, with every failure traced to a judge-confirmed root cause. A generational climb On the task set, pass@1 is 61.5% for…

Sep 23, 2026 •
Learn more about Opus 5.5 vs Opus 5 vs Fable 5.1: Coding Benchmark Results
Data 2.0 and the research era of AI data
Data 2.0 and the research era of AI data

Today, I’m excited to announce Snorkel’s $350M Series E financing at a $3.5B valuation, led by Insight and S32, with participation from Third Point, March, Blumberg, Allegis, Standard VC, Frontline, and existing investors Addition, Lightspeed, Greylock, GV, P7, Wells Fargo, Walden Catalyst Ventures, and Factory.

Sep 22, 2026 •
Learn more about Data 2.0 and the research era of AI data
Grok 4.7 on Senior SWE-Bench: Strong pass@3, Cheaper Cost per Trial
Grok 4.7 on Senior SWE-Bench: Strong pass@3, Cheaper Cost per Trial

Grok 4.7 was evaluated on Senior SWE-Bench, our benchmark for measuring whether coding agents work like senior engineers. It reaches a tasteful pass@3 of 40.0%, up from 38.9% for Grok 4.6, ranking sixth overall at $0.24 per trial, which is a fraction of the cost compared to Opus 4.8. Benchmark Senior SWE-Bench evaluates agents as compared to a senior engineer….

Sep 22, 2026 •
Learn more about Grok 4.7 on Senior SWE-Bench: Strong pass@3, Cheaper Cost per Trial
From Foundational Competency to Expert Performance: A Curriculum Approach to Model Development
From Foundational Competency to Expert Performance: A Curriculum Approach to Model Development

A student does not go from 1st to 12th grade in a single step. Each grade builds on a specific set of skills, and each one assumes the previous skills have already been mastered. Nobody learns calculus without algebra. When a student skips ahead anyway, what they end up with is memorization rather than understanding. The gaps show up later, usually in ways that are harder to trace back and fix.

Sep 15, 2026 •
Learn more about From Foundational Competency to Expert Performance: A Curriculum Approach to Model Development
OSWorld 2.0: Why Long-Horizon Computer-Use Agents Still Fail Four Out of Five Tasks
OSWorld 2.0: Why Long-Horizon Computer-Use Agents Still Fail Four Out of Five Tasks

Mengqi Yuan (XLANG Lab, University of Hong Kong) presents OSWorld 2.0, a benchmark of 108 long-horizon, real-world computer-use workflows where even frontier AI agents complete only 20.6% of tasks outright after 300+ steps each.

Sep 03, 2026 •
Learn more about OSWorld 2.0: Why Long-Horizon Computer-Use Agents Still Fail Four Out of Five Tasks
Fable 5.1 on Frontier Coding Tasks: Efficient Successes, Distinct Failure Modes
Fable 5.1 on Frontier Coding Tasks: Efficient Successes, Distinct Failure Modes

We evaluated Fable 5.1 on a series of frontier coding tasks from our proprietary Terminal-Bench+ dataset and compared the results against Opus 5. Fable remained competitive across most categories and was materially more efficient on successful runs, while its gap was concentrated in a small set of terminal-heavy and build/dependency tasks. Because category sizes are small and uneven, we treat…

Sep 01, 2026 •
Learn more about Fable 5.1 on Frontier Coding Tasks: Efficient Successes, Distinct Failure Modes
1 2 … 40
Image

Join our newsletter

For expert advice, the latest research, and exclusive events.
By submitting this form, I acknowledge I will receive email updates from Snorkel AI, and I agree to the Terms of Use and acknowledge that my information will be used in accordance with the Privacy Policy.