Wire · founder news, decoded · opportunities
'Far from human level': AI models score below 25% on real-world job tasks, UC Berkeley study finds
◆ Published
20 July 2026
◆ Topic
opportunities
◆ Sectors
◆ Geography
◆ Source
◆ Verified
Fusion42 · 20 July 2026 · Fusion42 review
UC Berkeley's real-world professional task benchmark found leading AI models (including GPT-5.5) score below 25% on complex workflows across 55 industries, with 0% success on the hardest tier. The study shows current AI excels at repetitive tasks but struggles with sustained reasoning, execution reliability, and adaptive problem-solving.
This Wire brief sits within Fusion42's coverage of AI Frontier Models. Wire is Fusion42's founder-focused intelligence feed: each story is connected to the funds and startups it names — every one with a live profile on Raise or Scout — so founders can follow the capital and the momentum behind the headline rather than just the headline itself. Wire analysis is one of the live surfaces Arthur reasons over.
◆ The Wire takeaway
If you're building AI tools for professional workflows, you're still selling belief, not capability: the models fail at the complex, adaptive work that actually drives revenue in finance, law, and manufacturing. Build for the 25% of tasks that are routine and repeatable—that's where AI delivers today, and where job displacement will actually happen first.
◆ Related on Wire
- AI skills now listed in 73% of tech job postings | CIO Dive10 July 2026
- When Companies Call AI a 'Colleague,' Oversight Slips, Harvard Study Warns8 July 2026
- Race for AI dominance heats up as China releases high-performance open-source models22 July 2026
- China's latest open-source AI model further narrows gap with US rivals20 July 2026
- Chinese AI Model Takes US Tech Industry by Surprise With Abilities Rivaling Claude and ChatGPT17 July 2026
- The rise of AI in GST enforcement: Are algorithms becoming the new tax officers?14 July 2026
◆ Topics
AI Frontier Models · ai-capabilities · job-displacement · benchmark-study · automation-risk · workflow-automation