Wire · technology
8 GPUs to 2… Nota Lowers Enterprise AX Costs by Lightening Ultra-Massive AI
◆ Sectors
◆ Geography
◆ Source
◆ Verified
Fusion42 · 27 July 2026 · Fusion42 review
Nota has optimised Solar Open 2, a 250-billion-parameter language model, reducing GPU memory requirements from 8 H100s to 2 whilst maintaining tool-calling performance through quantization and pruning techniques. The model weight shrunk by 76.5% (500.6GB to 117.8GB), lowering the infrastructure cost threshold for enterprise AI deployment.
This Wire brief sits within Fusion42's coverage of AI Infrastructure and Generative AI. Wire is Fusion42's founder-focused intelligence feed: each story is connected to the funds and startups it names — every one with a live profile on Raise or Scout — so founders can follow the capital and the momentum behind the headline rather than just the headline itself. Wire analysis is one of the live surfaces Arthur reasons over.
◆ ◆ The Wire takeaway
If you're selling inference infrastructure to enterprises running large language models, your unit economics just got four times worse: Solar Open 2 now runs on 2 GPUs instead of 8, and this technique scales to any 250B+ model. You need to compete on something other than raw compute—or become the inference optimizer yourself.
◆ Related on Wire
◆ Topics