OpenAI’s GDPval grades frontier models against expert deliverables across 44 occupations in the industries that drive the most U.S. GDP. Snorkel AI’s GDPval+ dataset extends the same idea into harder, more interactive territory.
44
Knowledge-work
occupations across the
9 industries contributing
most to U.S. GDP
1,320
Expert-vetted tasks
in the full set (220
in the open-sourced
gold set)
~48%
Share of tasks where
the top model tied or
beat human expert
deliverables
GDPval is an evaluation from OpenAI, introduced in September 2025, that measures model performance on economically valuable, real-world tasks. Instead of academic exam questions or synthetic coding challenges, GDPval tasks are built from the actual work products of experienced professionals: a legal brief, an engineering blueprint, a customer support transcript, a nursing care plan. The name comes from how the task set was scoped, starting with Gross Domestic Product as the organizing concept, then drawing tasks from the occupations inside the industries that contribute most to it.
The full set spans 1,320 tasks across 44 occupations (a 220-task “gold” subset is open-sourced), each written and vetted by a professional with an average of 14 years of experience in their field. Deliverables aren’t plain text: they include reference files and produce documents, slides, diagrams, spreadsheets, and multimedia, much closer to what a model would actually be asked to hand back in a real job.
OpenAI selected the 9 industries that each contribute more than 5% of U.S. GDP, using Federal Reserve data, then picked the 5 highest-wage occupations within each industry that qualified as “predominantly knowledge work,” using Bureau of Labor Statistics wage data and O*NET task classifications, with a 60% knowledge-work threshold. That process yielded 44 occupations.
Each task went through roughly 5 rounds of expert review, including other task writers, occupational reviewers, and model-based validation, to confirm it was representative, feasible, and clearly scoped for grading.
Expert graders from the same occupation as each task blindly compare a model’s deliverable against the professional’s own solution, without knowing which is which, then rank them and label each AI output as better than, as good as, or worse than the human version. Task writers also produce detailed scoring rubrics for consistency. OpenAI additionally built an automated grader (an AI system trained to predict how a human expert would judge a deliverable), and released it as an experimental research tool at evals.openai.com, though it isn’t yet reliable enough to replace expert graders.
Snorkel AI publishes GDPval+, an expert-extended edition of GDPval inside the Snorkel Data Series. It’s built from original, expert-authored occupational tasks in the spirit of GDPval, then hardened in the direction OpenAI itself named as the benchmark’s biggest limitation: more interactive, multi-draft workflows instead of one-shot prompts, alongside longer horizons, multi-skill tasks, richer metadata, and difficulty tiers calibrated against current frontier models.
That means teams evaluating against GDPval-style tasks can go beyond the open gold set once models start saturating it, with a benchmark built to track the frontier rather than a fixed 2025 snapshot of it.