Wire · opportunities
Frontier AI Models Fare Badly at Real-World Workplace Tasks: UCB Study
◆ Sectors
◆ Geography
◆ Source
◆ Verified
Fusion42 · 28 July 2026 · Fusion42 review
UC Berkeley's 'Agent's Last Exam' study finds frontier AI models (ChatGPT-5.5, Claude, Gemini) score below 25% on real-world professional tasks spanning 55 occupations, with 0% success on the hardest tier requiring sustained reasoning and domain expertise.
This Wire brief sits within Fusion42's coverage of AI Frontier Models. Wire is Fusion42's founder-focused intelligence feed: each story is connected to the funds and startups it names — every one with a live profile on Raise or Scout — so founders can follow the capital and the momentum behind the headline rather than just the headline itself. Wire analysis is one of the live surfaces Arthur reasons over.
◆ ◆ The Wire takeaway
The frontier models you're racing to integrate into workflows are failing at the complex, long-horizon tasks that actually matter—and zero of them can handle the hardest tier. Your moat isn't beating AI at your job; it's the domain expertise and sustained reasoning the models can't replicate yet, which buys you time to build defensible SaaS for the jobs AI will actually automate (the repetitive ones, not the ones that pay your bills).
◆ Related on Wire
◆ Topics