resources

Resource library

Explore our complete library of resources including blogs, benchmarks, research papers, and more.

Image for Data 2.0 and the research era of AI data
Blog

Data 2.0 and the research era of AI data

Announcing Snorkel’s $350M fundraise to build the frontier lab for AI data.

September 22, 2026
Image for Terminal-Bench 4.0: Why Continuous Benchmarks Require Continuous QA
Blog

Terminal-Bench 4.0: Why Continuous Benchmarks Require Continuous QA

Announcing a $3M commitment to launch Open Benchmarks Grants
August 28, 2026
Image for Grok 4.5 Testing Results: How SpaceXAI’s New Model Performs on Real Professional Work
Blog

Grok 4.5 Testing Results: How SpaceXAI’s New Model Performs on Real Professional Work

Announcing a $3M commitment to launch Open Benchmarks Grants
July 8, 2026
Image for Why coding agents need better data, evals, and environments
Blog

Why coding agents need better data, evals, and environments

Announcing a $3M commitment to launch Open Benchmarks Grants
May 11, 2026
Image for Closing the Evaluation Gap in Agentic AI
Blog

Closing the Evaluation Gap in Agentic AI

Announcing a $3M commitment to launch Open Benchmarks Grants

February 11, 2026
Image for Why Frontier Agents Fail Real Engineering Work: Two Terminal-Bench 3.0 Task Deep Dives
Blog

Why Frontier Agents Fail Real Engineering Work: Two Terminal-Bench 3.0 Task Deep Dives

Announcing a $3M commitment to launch Open Benchmarks Grants
August 24, 2026
Image for Building FinQA: An Open RL Environment for Financial Reasoning Agents
Blog

Building FinQA: An Open RL Environment for Financial Reasoning Agents

Announcing a $3M commitment to launch Open Benchmarks Grants
March 30, 2026
Image for The science of rubric design
Blog

The science of rubric design

Announcing a $3M commitment to launch Open Benchmarks Grants
September 11, 2025
Image for Benchtalks #3: We taught AI everything except how to learn
Blog

Benchtalks #3: We taught AI everything except how to learn

Featuring Parth Asawa (Continual Learning Bench)

June 25, 2026
of
Type: All Types
Sort: Newest
MedPAIR: Measuring Whether Physicians and AI Agree on What Matters in Medical QA
Blog
NEW
MedPAIR: Measuring Whether Physicians and AI Agree on What Matters in Medical QA

Yuexing Hao presents MedPAIR, a dataset comparing which sentences physicians and LLMs find relevant in clinical questions. Humans and LLMs agree on only 50 to 60% of relevance labels.

Oct 02, 2026 •
Learn more about MedPAIR: Measuring Whether Physicians and AI Agree on What Matters in Medical QA
Can Agents Design Libraries for Agents?
Agents increasingly build on code written by other agents, and they reimplement rather than reuse, growing the codebases later agents must work in. To measure how well agents design libraries for other agents, we introduce LibraryDesignBench, a two-phase benchmark in which an agent implements a full-featured library from a specification that defines required capabilities and potential use cases without prescribing the design. We evaluate the library through the correctness and simplicity of programs written by three user agents from different model families. The benchmark spans 242 expert-validated programming problems across 15 library-design tasks in four languages. On eleven of the...
Research Paper
NEW
Can Agents Design Libraries for Agents?

Agents increasingly build on code written by other agents, and they reimplement rather than reuse, growing the codebases later agents must work in. To measure how well agents design libraries for other agents, we introduce LibraryDesignBench, a two-phase benchmark in which an agent implements a full-featured library from a specification that defines required capabilities and potential use cases without prescribing…

Sep 30, 2026 •

Gabriel Orlanski, Alex L. Zhang, Avi Trost, Vincent Sunn Chen, Frederic Sala, Aws Albarghouthi, Ludwig Schmidt

Learn more about Can Agents Design Libraries for Agents?
RL environments for LLM agents: Design, rewards, and validation
Blog
NEW
RL environments for LLM agents: Design, rewards, and validation

TLDR: An agent, by definition, can take actions and is more than just a language model. The model is only one part of the system, and is only one element that you are training. The environment determines what the agent can observe, what it can change, which actions are available, and what behavior receives a reward. For a coding agent,…

Sep 28, 2026 •
Learn more about RL environments for LLM agents: Design, rewards, and validation
Opus 5.5 vs Opus 5 vs Fable 5.1: Coding Benchmark Results
Blog
Opus 5.5 vs Opus 5 vs Fable 5.1: Coding Benchmark Results

We’ve now run three generations of frontier models through the same expert-created terminal-bench style task set: Fable 5.1 in September, Opus 5 in July, and, now Opus 5.5. For this analysis, the task set contained 200 trajectories, of 24 tasks, with every failure traced to a judge-confirmed root cause. A generational climb On the task set, pass@1 is 61.5% for…

Sep 23, 2026 •
Learn more about Opus 5.5 vs Opus 5 vs Fable 5.1: Coding Benchmark Results
Data 2.0 and the research era of AI data
Blog
Data 2.0 and the research era of AI data

Today, I’m excited to announce Snorkel’s $350M Series E financing at a $3.5B valuation, led by Insight and S32, with participation from Third Point, March, Blumberg, Allegis, Standard VC, Frontline, and existing investors Addition, Lightspeed, Greylock, GV, P7, Wells Fargo, Walden Catalyst Ventures, and Factory.

Sep 22, 2026 •
Learn more about Data 2.0 and the research era of AI data
Grok 4.7 on Senior SWE-Bench: Strong pass@3, Cheaper Cost per Trial
Blog
Grok 4.7 on Senior SWE-Bench: Strong pass@3, Cheaper Cost per Trial

Grok 4.7 was evaluated on Senior SWE-Bench, our benchmark for measuring whether coding agents work like senior engineers. It reaches a tasteful pass@3 of 40.0%, up from 38.9% for Grok 4.6, ranking sixth overall at $0.24 per trial, which is a fraction of the cost compared to Opus 4.8. Benchmark Senior SWE-Bench evaluates agents as compared to a senior engineer….

Sep 22, 2026 •
Learn more about Grok 4.7 on Senior SWE-Bench: Strong pass@3, Cheaper Cost per Trial
From Foundational Competency to Expert Performance: A Curriculum Approach to Model Development
Blog
From Foundational Competency to Expert Performance: A Curriculum Approach to Model Development

A student does not go from 1st to 12th grade in a single step. Each grade builds on a specific set of skills, and each one assumes the previous skills have already been mastered. Nobody learns calculus without algebra. When a student skips ahead anyway, what they end up with is memorization rather than understanding. The gaps show up later, usually in ways that are harder to trace back and fix.

Sep 15, 2026 •
Learn more about From Foundational Competency to Expert Performance: A Curriculum Approach to Model Development
OSWorld 2.0: Why Long-Horizon Computer-Use Agents Still Fail Four Out of Five Tasks
Blog
OSWorld 2.0: Why Long-Horizon Computer-Use Agents Still Fail Four Out of Five Tasks

Mengqi Yuan (XLANG Lab, University of Hong Kong) presents OSWorld 2.0, a benchmark of 108 long-horizon, real-world computer-use workflows where even frontier AI agents complete only 20.6% of tasks outright after 300+ steps each.

Sep 03, 2026 •
Learn more about OSWorld 2.0: Why Long-Horizon Computer-Use Agents Still Fail Four Out of Five Tasks
Fable 5.1 on Frontier Coding Tasks: Efficient Successes, Distinct Failure Modes
Blog
Fable 5.1 on Frontier Coding Tasks: Efficient Successes, Distinct Failure Modes

We evaluated Fable 5.1 on a series of frontier coding tasks from our proprietary Terminal-Bench+ dataset and compared the results against Opus 5. Fable remained competitive across most categories and was materially more efficient on successful runs, while its gap was concentrated in a small set of terminal-heavy and build/dependency tasks. Because category sizes are small and uneven, we treat…

Sep 01, 2026 •
Learn more about Fable 5.1 on Frontier Coding Tasks: Efficient Successes, Distinct Failure Modes
1 2 … 70
Image

Join our newsletter

For expert advice, the latest research, and exclusive events.
By submitting this form, I acknowledge I will receive email updates from Snorkel AI, and I agree to the Terms of Use and acknowledge that my information will be used in accordance with the Privacy Policy.