GDPval: Measuring AI on Economically Valuable, Real-World Work

Sep 25, 2025

OpenAI’s GDPval grades frontier models against expert deliverables across 44 occupations in the industries that drive the most U.S. GDP. Snorkel AI’s GDPval+ dataset extends the same idea into harder, more interactive territory.


What is GDPval?

GDPval is an evaluation from OpenAI, introduced in September 2025, that measures model performance on economically valuable, real-world tasks. Instead of academic exam questions or synthetic coding challenges, GDPval tasks are built from the actual work products of experienced professionals: a legal brief, an engineering blueprint, a customer support transcript, a nursing care plan. The name comes from how the task set was scoped, starting with Gross Domestic Product as the organizing concept, then drawing tasks from the occupations inside the industries that contribute most to it.

The full set spans 1,320 tasks across 44 occupations (a 220-task “gold” subset is open-sourced), each written and vetted by a professional with an average of 14 years of experience in their field. Deliverables aren’t plain text: they include reference files and produce documents, slides, diagrams, spreadsheets, and multimedia, much closer to what a model would actually be asked to hand back in a real job.

How occupations and tasks were chosen

OpenAI selected the 9 industries that each contribute more than 5% of U.S. GDP, using Federal Reserve data, then picked the 5 highest-wage occupations within each industry that qualified as “predominantly knowledge work,” using Bureau of Labor Statistics wage data and O*NET task classifications, with a 60% knowledge-work threshold. That process yielded 44 occupations.

  • Professional & technical services: Software developers, lawyers, accountants
  • Health care: Registered nurses, medical managers
  • Finance & insurance: Financial analysts, advisors
  • Manufacturing: Mechanical & industrial engineers
  • Government: Compliance officers, social workers
  • Real estate: Brokers, property managers
  • Retail trade: Pharmacists, store operations
  • Wholesale trade: Sales managers, reps
  • Information: Journalists, editors, producers

Each task went through roughly 5 rounds of expert review, including other task writers, occupational reviewers, and model-based validation, to confirm it was representative, feasible, and clearly scoped for grading.

How grading works

Expert graders from the same occupation as each task blindly compare a model’s deliverable against the professional’s own solution, without knowing which is which, then rank them and label each AI output as better than, as good as, or worse than the human version. Task writers also produce detailed scoring rubrics for consistency. OpenAI additionally built an automated grader (an AI system trained to predict how a human expert would judge a deliverable), and released it as an experimental research tool at evals.openai.com, though it isn’t yet reliable enough to replace expert graders.

Key findings

  • Frontier models are closing in on expert quality. Across the 220-task gold set, Claude Opus 4.1 was the top performer, with outputs rated as good as or better than human experts in just under half of tasks, especially strong on aesthetics like document formatting and slide layout. GPT-5 was close behind, particularly strong on accuracy and following complex, multi-step instructions.
  • Progress has been fast. Performance on GDPval tasks more than doubled, and by some measures tripled, from GPT-4o to GPT-5 in about a year, a clear, roughly linear improvement trend rather than a plateau.
  • Speed and cost gaps are large. Frontier models complete GDPval tasks on the order of 100x faster and 100x cheaper than human experts, measured purely on inference time and API billing, a figure that excludes the human oversight, iteration, and integration real deployments require.
  • It’s a one-shot evaluation, by design and by limitation. GDPval v1 doesn’t capture multi-draft revision or a model navigating ambiguity before starting work, the kind of back-and-forth real knowledge work usually involves. OpenAI has flagged this directly as the next thing to extend.

Snorkel AI’s role

Snorkel AI publishes GDPval+, an expert-extended edition of GDPval inside the Snorkel Data Series. It’s built from original, expert-authored occupational tasks in the spirit of GDPval, then hardened in the direction OpenAI itself named as the benchmark’s biggest limitation: more interactive, multi-draft workflows instead of one-shot prompts, alongside longer horizons, multi-skill tasks, richer metadata, and difficulty tiers calibrated against current frontier models.

That means teams evaluating against GDPval-style tasks can go beyond the open gold set once models start saturating it, with a benchmark built to track the frontier rather than a fixed 2025 snapshot of it.

Explore GDPval+ →

 Illution Back
Illution Front

For models that need to be right. Not just good enough.