Skip to main content

Command Palette

Search for a command to run...

Open-Weight vs Proprietary AI: Cost, Licensing, and Accuracy Compared

Updated
•8 min read•View as Markdown
Open-Weight vs Proprietary AI: Cost, Licensing, and Accuracy Compared
H
Father of two, tech lover. Building systems by day, raising curious minds by night.

Everyone is talking about open-weight AI models. You can download them. You can run them. You can even build products on them.

So they must be free, right?

Not even close.

This study breaks down what AI models actually cost in 2026 — split into open-weight and proprietary models — and shows why "open" is not the same as "free."

But before we dive in, let's understand the basic unit of AI pricing: the token.


Before We Start: What Is a Token?

AI models don't read words like humans do. They break text into small pieces called tokens.

A token is roughly 4 characters or about ¾ of a word in English.

Text Tokens
"Hello" 1
"Artificial intelligence" 2
"ChatGPT is amazing" 4
"I love learning about AI models" 7

When you use an AI model through an API, you pay per 1 million tokens. There are two types:

  • Input tokens — the text you send to the model

  • Output tokens — the text the model sends back

Output tokens almost always cost more than input tokens — usually 5 to 6 times more. Generating new text takes more computing power than reading existing text.

So when you see "\(1.25 input / \)10 output per 1M tokens," you pay $1.25 for every million tokens you send, and $10 for every million the model generates. Long answers get expensive fast.

Now that you understand tokens, let's look at the real costs.


Part 1: Open-Weight Models Are Not Free

Open-weight models are models whose parameters are publicly downloadable. Anyone can grab them.

But three hidden costs come with them:

  1. Hardware. Big models need big machines.

  2. Licenses. Many "open" licenses have strings attached.

  3. People. Someone has to run and maintain the servers.

The License Trap

Model License Commercial Use & Catch
DeepSeek R1 / V3 MIT Free, no restrictions
Gemma 4 Apache 2.0 Free, no restrictions
Phi-4 MIT Free, no restrictions
GLM-5.2 MIT Free, no restrictions
OLMo Apache 2.0 Free, no restrictions
Llama 4 Llama Community Free until 700M monthly users
Mistral Medium 3.5 Modified MIT Commercial supply restriction
Kimi K3 Custom Revenue share up to 30% for MaaS
Grok-2 Custom Free only under $1M/year revenue
StableLM Custom Free only under $1M/year revenue
Command R CC-BY-NC Research only. No commercial use

The lesson: MIT and Apache 2.0 licenses are genuinely free. Custom licenses are not. They look open, but they're really "free until you get big."

The Hardware Bill

Model Size Hardware Needed Rough Cost
Kimi K3 2.8T params Two 8-GPU H200 servers $320K–$420K+
Llama 4 Maverick 400B 8× H100 GPUs $80K–$120K
DeepSeek V3 671B 4–8 datacenter GPUs $50K–$80K
Gemma 4 31B 31B One gaming GPU ~$2,000
Phi-4 14B One gaming GPU ~$2,000

Small models really can be nearly free. Frontier models cannot.

The Break-Even Point

Renting an 8-GPU NVIDIA H100 node on a specialized cloud provider costs roughly $18,000 per month, though major hyperscalers can charge over $70,000 for the same setup.

Daily Usage API Cost/Month Self-Host Cost/Month Cheaper
1M tokens $150 $18,221 API
50M tokens $7,500 $18,221 API
100M tokens $15,000 $18,221 API
500M+ tokens — — Self-host

Simple rule: Unless you're processing over 100–500 million tokens per day, self-hosting a big open model costs more than just paying for an API.


Part 2: Proprietary Models — Pay As You Go

Proprietary models are rented, not owned. You send text, you get answers, you pay per token.

Model Company Input $/1M Output $/1M
GPT-5.5 OpenAI $5.00 $45.00
GPT-5 OpenAI $1.25 $10.00
GPT-5 mini OpenAI $0.25 $2.00
GPT-5 nano OpenAI $0.05 $0.40
Claude Opus 5 Anthropic $5.00 $25.00
Claude Sonnet 5 Anthropic $2.00 $10.00
Gemini 3.8 Flash Google $0.75 $3.75
DeepSeek V4 Pro DeepSeek $0.66 $1.98
Kimi K3 Moonshot AI $3.00 $15.00

The good news: no license negotiations, no user caps, no hardware. You just pay the bill.

Two things to watch:

  • Output costs more than input — usually 5–6× more.

  • Batch processing is 50% cheaper if you don't need answers immediately.


Part 3: Cost vs. Accuracy

Here's the big question: how much accuracy do you lose by going cheap?

General Knowledge (MMLU-Pro)

Model Type Score Input Cost
Qwen3.7 Max Proprietary 89.6% API only
Claude Opus 4.5 Proprietary 89.5% $5.00
Qwen3.5 397B Open weight 87.8% Self-host
Kimi K2.5 Open weight 87.1% $3.00
DeepSeek V4 Pro Open weight 87.1% $0.66
Gemma 4 31B Open weight ~85% Nearly free

The gap is about 2 points — but the price difference is 40×.

Real-World Coding (SWE-bench Verified)

This is the hard one. It measures whether a model can actually fix real software bugs.

Model Type Score
Claude Opus 5 Proprietary 96.0%
Claude Fable 5 Proprietary 95.0%
Ornith-1.5-397B Open weight 86.0%
DeepSeek V4 Pro Open weight 80.6%
MiniMax M3 Open weight 80.5%
Kimi K2.6 Open weight ~75%

Here the gap is real. For genuinely hard work — complex coding, multi-step reasoning — proprietary models still win.


Part 4: So What Should You Do?

Use open weights when:

  • Your task is simple (summarizing, sorting, basic chat)

  • You need data to stay on your own servers

  • You're processing over 100M tokens/day

  • You want to avoid vendor lock-in

Use proprietary models when:

  • The task is hard (complex reasoning, agentic coding)

  • Your volume is low or medium

  • You need legal protection against IP claims

  • You don't want to hire engineers to run servers

The best answer for most teams: use both. Route easy tasks to cheap open models. Send hard tasks to frontier proprietary models.


The Bottom Line

Open-weight models changed the game. They can cut your costs by 80–95% while giving up only a few points of accuracy.

But three things stay true:

  1. Downloading is free. Running is not. A frontier open model can cost $300K+ in hardware.

  2. "Open" licenses often have limits. Only MIT and Apache 2.0 are truly free.

  3. Someone has to maintain it. Budget for engineering time, not just GPUs.

The smartest teams in 2026 aren't choosing sides. They're mixing both.


References

  1. MillionMiner. "AI Server Price Guide 2026: Rent Vs Buy." September 9, 2026.

  2. IntuitionLabs. "Open-Weight AI Model Licenses: Commercial Use Rules Explained." September 5, 2026.

  3. Taskade. "10 Best Open-Source LLMs, August 2026." May 23, 2026.

  4. BenchLM. "Self-Hosting vs API: Break-Even Calculator for Open-Weight LLMs." July 3, 2026.

  5. OpenAI. "Pricing." OpenAI API, 2026.

  6. Anthropic. "Claude Opus 5 Pricing." July 24, 2026.

  7. Google. "What's new in Gemini 3.8 Flash." September 3, 2026.

  8. BenchLM. "MMLU-Pro Leaderboard & Scores — September 2026." September 9, 2026.

  9. BenchLM. "SWE-bench Verified Leaderboard (September 2026)." September 10, 2026.

  10. Meta. "Llama 4 Community License Agreement." April 5, 2025.

  11. Hugging Face. "Kimi K3 License." July 30, 2026.

  12. Reuters. "China's DeepSeek to make permanent 75% price cut on flagship V4-Pro AI model." May 23, 2026.

All pricing and benchmark data reflects publicly available information as of September 2026. Verify current terms and pricing before making procurement decisions.

A

"Open is not the same as free" is the right thesis, and starting from the token is the correct move since most cost confusion is really unit confusion.

The number that decides it in practice is utilisation. A rented GPU costs the same at 5% as at 95%, so self-hosting only wins above a fairly high floor of steady traffic, and bursty workloads pay for idle silicon overnight.

Two costs that rarely make these comparisons: the evaluation and rollback machinery you now own, since a serving change that regresses quality becomes your problem rather than a provider's, and the engineer-hours spent keeping it running, which at small scale are usually larger than the GPU bill.

M

The break-even table is the part I'd push on. $18K/month only makes sense if that node sits at roughly full utilization the whole month, and most agent workloads aren't anywhere near that -- they spike during business hours and sit mostly idle overnight. Once you factor in real utilization, or swap a reserved node for a per-second serverless GPU provider instead, the token volume where self-hosting wins drops well below the 100-500M/day line here, sometimes by an order of magnitude. Utilization rate deserves its own column next to token volume, not just token volume on its own.