← Back

Wire · opportunities

Frontier AI Models Fare Badly at Real-World Workplace Tasks: UCB Study

Published

28 July 2026

Topic

opportunities

Sectors

AI Frontier Models

Geography

United States

Source

Read at cxotoday.com

Verified

Fusion42 · 28 July 2026 · Fusion42 review

UC Berkeley's 'Agent's Last Exam' study finds frontier AI models (ChatGPT-5.5, Claude, Gemini) score below 25% on real-world professional tasks spanning 55 occupations, with 0% success on the hardest tier requiring sustained reasoning and domain expertise.

This Wire brief sits within Fusion42's coverage of AI Frontier Models. Wire is Fusion42's founder-focused intelligence feed: each story is connected to the funds and startups it names — every one with a live profile on Raise or Scout — so founders can follow the capital and the momentum behind the headline rather than just the headline itself. Wire analysis is one of the live surfaces Arthur reasons over.

◆ The Wire takeaway

The frontier models you're racing to integrate into workflows are failing at the complex, long-horizon tasks that actually matter—and zero of them can handle the hardest tier. Your moat isn't beating AI at your job; it's the domain expertise and sustained reasoning the models can't replicate yet, which buys you time to build defensible SaaS for the jobs AI will actually automate (the repetitive ones, not the ones that pay your bills).

Related on Wire

Topics

AI Frontier Modelsai-agentsworkplace-automationcapability-gapsustained-reasoningbenchmarking