← Back

Wire · technology

8 GPUs to 2… Nota Lowers Enterprise AX Costs by Lightening Ultra-Massive AI

Published

27 July 2026

Topic

technology

Sectors

AI InfrastructureGenerative AI

Geography

South Korea

Source

Read at venturesquare.net

Verified

Fusion42 · 27 July 2026 · Fusion42 review

Nota has optimised Solar Open 2, a 250-billion-parameter language model, reducing GPU memory requirements from 8 H100s to 2 whilst maintaining tool-calling performance through quantization and pruning techniques. The model weight shrunk by 76.5% (500.6GB to 117.8GB), lowering the infrastructure cost threshold for enterprise AI deployment.

This Wire brief sits within Fusion42's coverage of AI Infrastructure and Generative AI. Wire is Fusion42's founder-focused intelligence feed: each story is connected to the funds and startups it names — every one with a live profile on Raise or Scout — so founders can follow the capital and the momentum behind the headline rather than just the headline itself. Wire analysis is one of the live surfaces Arthur reasons over.

◆ The Wire takeaway

If you're selling inference infrastructure to enterprises running large language models, your unit economics just got four times worse: Solar Open 2 now runs on 2 GPUs instead of 8, and this technique scales to any 250B+ model. You need to compete on something other than raw compute—or become the inference optimizer yourself.

Related on Wire

Topics

AI InfrastructureGenerative AImodel-lightweightinggpu-efficiencyenterprise-aicost-reductioninference-optimization