OpenAI has unveiled a new benchmark, GDPval, designed to evaluate the performance of its AI models compared to human professionals across a range of industries and occupations. The test represents a first attempt to assess how close AI systems are to outperforming humans in economically valuable work—a key part of OpenAI’s mission to develop artificial general intelligence (AGI).
OpenAI reports that its GPT-5 model, along with Anthropic’s Claude Opus 4.1, “are already approaching the quality of work produced by industry experts.”
What GDPval Measures
GDPval focuses on nine of the U.S. industries that contribute most to GDP, including healthcare, finance, manufacturing, and government. It evaluates AI performance across 44 occupations, from software engineers to nurses and journalists.
For the initial version, GDPval-v0, experienced professionals were asked to compare reports generated by AI with reports produced by human peers and select the better output. For example, investment bankers were asked to create a competitive landscape for last-mile delivery companies and compare it to AI-generated reports. OpenAI then calculated a “success rate” for AI models by averaging their performance across all 44 roles.
AI Performance Insights
GPT-5-high, a more powerful variant of GPT-5, was rated equal to or better than human experts 40.6% of the time. Claude Opus 4.1 achieved a similar or superior rating in 49% of the tasks, which OpenAI attributes partly to Claude’s ability to produce visually appealing charts.
While these results are impressive, OpenAI notes that GDPval currently covers only a small subset of the actual work humans perform. Most professionals engage in far more complex, interactive, and dynamic tasks than the reports evaluated in the benchmark.
Implications for the Workplace
Researchers say that GDPval’s results suggest AI can increasingly assist human professionals, allowing them to focus on higher-value tasks. “As the model improves in these areas,” says OpenAI chief economist Dr. Aaron Chatterji, “people in these roles can use it to delegate portions of their work and engage in more meaningful activities.”
Tejal Patwardhan, OpenAI’s Director of Evaluations, points out that progress has been rapid: GPT-4o, released just over a year ago, scored only 13.7% (victories or ties against humans), while GPT-5 nearly triples that score—a trend she expects to continue.
Broader Context in AI Evaluation
Silicon Valley currently relies on a variety of benchmarks to measure AI progress, including AIME 2025 (competitive math problems) and GPQA Diamond (PhD-level scientific questions). However, many models are reaching performance ceilings on these tests. Researchers increasingly emphasize the need for benchmarks that assess AI’s capabilities on real-world tasks.
GDPval is a step in that direction. By testing AI in economically relevant industries, the benchmark offers insights into where AI can add tangible value. Researchers note that as benchmarks like GDPval evolve, they could become a critical tool for understanding AI’s impact on the workforce and guiding its responsible adoption.
Looking Ahead
OpenAI acknowledges that a more comprehensive version of GDPval will be necessary to make definitive claims about AI outperforming humans across industries. Nonetheless, the benchmark highlights the growing role of AI in professional settings, from finance to healthcare, and points toward a future where AI complements human labor rather than simply replacing it.
Looking for help building a product idea? Reach out to us through the form below. We help businesses like yours build and deliver big ideas. See our case studies for more.







