Hearing open weights easily suggests cheap or free.

But downloadable weights and free inference are completely different stories.

Open-weight models cost money through APIs too

Artificial Analysis comparisons from September 2026 show open-weight models like GLM-5.3, Kimi K3, and Qwen3.8 with different token prices, speeds, and cost-per-task figures on provider APIs.

The interesting part: per-token price order and per-task cost order do not always match.

A model that emits more reasoning and output tokens can erase its unit-price advantage.

So at minimum, look at these together.

Intelligence / task success
Input price
Output price
Tokens per task
Time per task
Cost per task

Is self-hosting really free?

Hosting on your own GPUs can remove the API token bill. Other costs appear instead.

GPU rental / depreciation
+ idle capacity
+ inference server
+ autoscaling
+ observability
+ engineering time
+ model updates

So the comparison becomes:

Hosted API TCO
vs
Self-hosted TCO

CodeBridge Mini Lab: monthly break-even math

Write down your current workload for one month.

input tokens / month
output tokens / month
peak requests per second
required latency

API cost:

input_tokens × input_rate
+ output_tokens × output_rate

Self-host cost:

GPU hours
+ storage/network
+ operations staffing time

Then compute each at 30%, 100%, and 300% of the workload.

At low usage, idle GPUs can make APIs cheaper. At steady, large workloads, self-hosting can turn competitive.

Check active parameters alongside model size

MoE models can differ between total parameter count and the parameters activated during inference.

Artificial Analysis reports total and active parameters separately for models like GLM-5.3 and Kimi K3.

Rather than judging "2.8T must be slow" from one number, measure actual provider speed and latency.

The real upside of open weights is not only cost

Depending on your situation, stronger reasons exist.

  • On-premises deployment
  • Model modification and fine-tuning
  • Control of the inference stack
  • Less dependence on one provider
  • Data location control

Meanwhile managed frontier APIs may ship new features and tool integrations faster.

Conclusion: open vs closed is not a one-line price decision

Open weights are a strong option, but the free model frame oversimplifies reality.

Compare these in real decisions.

Cost per success + required latency + operational complexity + control

With those four together, it becomes much clearer where self-hosting pays and where an API is the economical pick.

Further reading

References

Go deeper with a course

If you want a repeatable way to route each workload to the right model at the right cost, a guided course on working with multiple AI tools teaches the selection habit.