We develop methods, benchmarks, and training systems that turn expert data into frontier AI
building benchmarks and collaborating with
Featured research
Vision and impact
We help labs advance frontier models by working with domain experts to design and build complex, realistic datasets that drive model performance.
Benchmarking & Evaluation
Build benchmarks that define and advance the AI frontier
Scaling Subject Matter Expertise
Define how subject matter experts encode their knowledge into data
RL, Training, & Data Valuation
Drive dataset development based on feedback from RL and model training
Community and open science
Open benchmarks, conversations, and research for real-world AI performance.


Open Benchmarks Grants
Backed by a $3M commitment, the program funds open-source datasets, benchmarks, and evaluation artifacts that shape how frontier AI systems are built and evaluated.


Benchtalks


Reading Group
DEEP RESEARCH Expertise
Technical advisors and distinguished affiliates
Browse research blogs and academic papers


Yuexing Hao presents MedPAIR, a dataset comparing which sentences physicians and LLMs find relevant in clinical questions. Humans and LLMs agree on only 50 to 60% of relevance labels.
Agents increasingly build on code written by other agents, and they reimplement rather than reuse, growing the codebases later agents must work in. To measure how well agents design libraries for other agents, we introduce LibraryDesignBench, a two-phase benchmark in which an agent implements a full-featured library from a specification that defines required capabilities and potential use cases without prescribing…


TLDR: An agent, by definition, can take actions and is more than just a language model. The model is only one part of the system, and is only one element that you are training. The environment determines what the agent can observe, what it can change, which actions are available, and what behavior receives a reward. For a coding agent,…


We’ve now run three generations of frontier models through the same expert-created terminal-bench style task set: Fable 5.1 in September, Opus 5 in July, and, now Opus 5.5. For this analysis, the task set contained 200 trajectories, of 24 tasks, with every failure traced to a judge-confirmed root cause. A generational climb On the task set, pass@1 is 61.5% for…


Today, I’m excited to announce Snorkel’s $350M Series E financing at a $3.5B valuation, led by Insight and S32, with participation from Third Point, March, Blumberg, Allegis, Standard VC, Frontline, and existing investors Addition, Lightspeed, Greylock, GV, P7, Wells Fargo, Walden Catalyst Ventures, and Factory.


Grok 4.7 was evaluated on Senior SWE-Bench, our benchmark for measuring whether coding agents work like senior engineers. It reaches a tasteful pass@3 of 40.0%, up from 38.9% for Grok 4.6, ranking sixth overall at $0.24 per trial, which is a fraction of the cost compared to Opus 4.8. Benchmark Senior SWE-Bench evaluates agents as compared to a senior engineer….


A student does not go from 1st to 12th grade in a single step. Each grade builds on a specific set of skills, and each one assumes the previous skills have already been mastered. Nobody learns calculus without algebra. When a student skips ahead anyway, what they end up with is memorization rather than understanding. The gaps show up later, usually in ways that are harder to trace back and fix.


Mengqi Yuan (XLANG Lab, University of Hong Kong) presents OSWorld 2.0, a benchmark of 108 long-horizon, real-world computer-use workflows where even frontier AI agents complete only 20.6% of tasks outright after 300+ steps each.
We evaluated Fable 5.1 on a series of frontier coding tasks from our proprietary Terminal-Bench+ dataset and compared the results against Opus 5. Fable remained competitive across most categories and was materially more efficient on successful runs, while its gap was concentrated in a small set of terminal-heavy and build/dependency tasks. Because category sizes are small and uneven, we treat…
October 8, 2026 | San francisco
A one-day, invite-only summit providing a first look at the benchmarks and research that will shape the frontier.





















