In a bold step toward its mission of building artificial general intelligence (AGI), OpenAI says GPT-5 stacks up to humans across a surprising range of jobs, according to its newly released benchmark test, GDPval. The company claims that its next-generation model is approaching the quality of work produced by industry experts in multiple professional domains—a milestone that could redefine the future of work.
On Thursday, OpenAI unveiled GDPval, a first-of-its-kind evaluation designed to measure how closely AI models like GPT-5 and Anthropic’s Claude Opus 4.1 compare to humans in economically significant roles. The early results suggest that AI’s capabilities are improving far faster than many expected, with the potential to support—and in some cases rival—human professionals.
AI Benchmarks Put GPT-5 in the Spotlight

OpenAI explains that GDPval currently covers nine industries contributing the most to America’s GDP, including healthcare, finance, manufacturing, and government. Within these industries, the benchmark evaluates 44 occupations—from software engineers and financial analysts to nurses and journalists—by comparing AI-generated reports with those created by human experts.
In the test’s first iteration, GDPval-v0, seasoned professionals were asked to review AI-produced reports against human-generated ones and then vote on which they believed was superior. OpenAI then calculated a “win rate” for the AI models across all occupations.
Results revealed that GPT-5-high, a more advanced version of GPT-5 with extra computing power, was rated equal to or better than human experts in 40.6% of cases. Anthropic’s Claude Opus 4.1 scored even higher, achieving parity or superiority in 49% of the tasks tested—a result OpenAI believes was influenced by Claude’s strong visual presentation skills.
From GPT-4o to GPT-5: Rapid Progress in AI Performance
Just 15 months ago, OpenAI’s GPT-4o scored 13.7% in the same benchmark, highlighting how quickly AI technology is advancing. Tejal Patwardhan, who leads OpenAI’s evaluations team, described this leap in performance as “encouraging,” adding that she expects the upward trend to continue as AI models become more sophisticated.
Dr. Aaron Chatterji, OpenAI’s chief economist, explained in an interview with TechCrunch that as models improve, professionals in various industries will be able to offload repetitive or time-consuming tasks to AI, freeing themselves to focus on higher-value work.
Why GDPval Matters for AI’s Future
The GDPval benchmark aims to test how close AI models are to performing “economically valuable work”—a core part of OpenAI’s vision for AGI. While most existing benchmarks, such as AIME 2025 for math problems or GPQA Diamond for scientific reasoning, focus on narrow skills, GDPval attempts to measure performance on real-world, high-impact tasks.
OpenAI acknowledges, however, that GDPval’s current scope is limited. The test primarily evaluates written outputs rather than interactive or dynamic workplace tasks, meaning future iterations will need to cover broader workflows before declaring that AI can truly outperform humans in the workplace.
Still, OpenAI believes benchmarks like GDPval will be essential as AI adoption accelerates across industries. As more companies integrate AI into everyday operations, these tests will help measure whether models like GPT-5 are ready for real-world applications beyond research labs.
The Bigger Picture: AI, Work, and Human Productivity
For now, OpenAI says GPT-5 stacks up to humans in a way that signals collaboration rather than replacement. By automating portions of work that require research, data synthesis, or written analysis, AI could help professionals redirect time and resources toward creativity, strategy, and decision-making.
With Anthropic’s Claude and other rivals also making rapid advances, the race to develop powerful yet practical AI systems is clearly accelerating. Benchmarks like GDPval are likely to become key indicators of which models are best suited for industries under pressure to modernize through automation.