DigiNews

Tech Watch by Johan Denoyer

← Back to articles

The Economics of Open-Weight Inference

Quality: 8/10 Relevance: 9/10

Summary

Ornn Data’s paper argues that open-weight inference can extend the useful life of NVIDIA GPUs by enabling cheaper, self-hosted deployments compared with locked, subscription-based models. Across multiple GPU families and workloads, open-weight models often cost less per million tokens, though closed models remain cheaper in some regimes. The analysis shows older GPUs can remain economically viable for select workloads, with self-hosting reducing compute costs to as low as $0.12–$0.35 per million tokens, while forward-term pricing and occupancy data illustrate market dynamics in on-demand GPU rental. The study highlights workload flexibility and cost optimization as key drivers for open-weight adoption, while acknowledging limitations and the commercial interests behind the data.

🚀 Service construit par Johan Denoyer