Hearing open weights easily suggests cheap or free.
But downloadable weights and free inference are completely different stories.
Open-weight models cost money through APIs too
Artificial Analysis comparisons from September 2026 show open-weight models like GLM-5.3, Kimi K3, and Qwen3.8 with different token prices, speeds, and cost-per-task figures on provider APIs.
The interesting part: per-token price order and per-task cost order do not always match.
A model that emits more reasoning and output tokens can erase its unit-price advantage.
So at minimum, look at these together.
Intelligence / task success
Input price
Output price
Tokens per task
Time per task
Cost per task
Is self-hosting really free?
Hosting on your own GPUs can remove the API token bill. Other costs appear instead.
GPU rental / depreciation
+ idle capacity
+ inference server
+ autoscaling
+ observability
+ engineering time
+ model updates
So the comparison becomes:
Hosted API TCO
vs
Self-hosted TCO
CodeBridge Mini Lab: monthly break-even math
Write down your current workload for one month.
input tokens / month
output tokens / month
peak requests per second
required latency
API cost:
input_tokens × input_rate
+ output_tokens × output_rate
Self-host cost:
GPU hours
+ storage/network
+ operations staffing time
Then compute each at 30%, 100%, and 300% of the workload.
At low usage, idle GPUs can make APIs cheaper. At steady, large workloads, self-hosting can turn competitive.
Check active parameters alongside model size
MoE models can differ between total parameter count and the parameters activated during inference.
Artificial Analysis reports total and active parameters separately for models like GLM-5.3 and Kimi K3.
Rather than judging "2.8T must be slow" from one number, measure actual provider speed and latency.
The real upside of open weights is not only cost
Depending on your situation, stronger reasons exist.
- On-premises deployment
- Model modification and fine-tuning
- Control of the inference stack
- Less dependence on one provider
- Data location control
Meanwhile managed frontier APIs may ship new features and tool integrations faster.
Conclusion: open vs closed is not a one-line price decision
Open weights are a strong option, but the free model frame oversimplifies reality.
Compare these in real decisions.
Cost per success + required latency + operational complexity + control
With those four together, it becomes much clearer where self-hosting pays and where an API is the economical pick.
Further reading
- Judge AI model cost per successful task
- Why benchmark scores alone mislead you
- Luna to Sol to Astra model routing
References
Go deeper with a course
If you want a repeatable way to route each workload to the right model at the right cost, a guided course on working with multiple AI tools teaches the selection habit.