← Back

Wire · founder news, decoded · opportunities

'Far from human level': AI models score below 25% on real-world job tasks, UC Berkeley study finds

Published

20 July 2026

Topic

opportunities

Sectors

AI Frontier Models

Geography

United States

Source

Read at thecollegefix.com

Verified

Fusion42 · 20 July 2026 · Fusion42 review

UC Berkeley's real-world professional task benchmark found leading AI models (including GPT-5.5) score below 25% on complex workflows across 55 industries, with 0% success on the hardest tier. The study shows current AI excels at repetitive tasks but struggles with sustained reasoning, execution reliability, and adaptive problem-solving.

This Wire brief sits within Fusion42's coverage of AI Frontier Models. Wire is Fusion42's founder-focused intelligence feed: each story is connected to the funds and startups it names — every one with a live profile on Raise or Scout — so founders can follow the capital and the momentum behind the headline rather than just the headline itself. Wire analysis is one of the live surfaces Arthur reasons over.

The Wire takeaway

If you're building AI tools for professional workflows, you're still selling belief, not capability: the models fail at the complex, adaptive work that actually drives revenue in finance, law, and manufacturing. Build for the 25% of tasks that are routine and repeatable—that's where AI delivers today, and where job displacement will actually happen first.

Related on Wire

Topics

AI Frontier Models · ai-capabilities · job-displacement · benchmark-study · automation-risk · workflow-automation