On October 2, NVIDIA put a computer on your desk. DGX Spark 64GB. Shipping October 23 through Acer, ASUS, Dell, Gigabyte, HP, and MSI, starting at $4,999.
The specs: GB10 Grace Blackwell chip, 64GB unified memory, up to 100B-parameter models on one box. Link two over ConnectX-7 and they share 128GB for up to 200B. NVIDIA reports up to 1.7x performance on Qwen 3.8 27B in the 2-node setup.
But spec talk is not the point. The question changed. From "can a small model run locally" to "is keeping a private agent always on cheaper than cloud."
What is inside: ecosystem over specs
| Item | Detail |
|---|---|
| Chip and memory | GB10 Grace Blackwell, 64GB unified memory |
| Model scale | Up to 100B solo, up to 200B with two linked units |
| Scaling | ConnectX-7 link plus Sync Cluster Assistant auto setup |
| OS and tools | DGX OS, Ollama, vLLM, llama.cpp, LM Studio, Agent Toolkit |
| Price and date | From $4,999, shipping October 23 |
What matters is the defaults, not the box. Ollama, vLLM, and PyTorch just run, and NVIDIA names "24-hour coding and research agent runs" as the use case. Local AI is moving from "occasionally tinkered hobby" to "always-on assistant." That is close to a declaration.
One pricing caveat: reports on October 3 say the 128GB model jumped toward $6,950 amid a RAM shortage, so recheck prices at launch. The calculation method below works whatever the numbers say.
The economics: is keeping it on cheaper
One formula before buying.
Monthly cloud GPU cost x actual usage hours
vs
(device price / months of use) + power + admin time
x privacy needs x always-on agent count
A worked example for feel. (Illustrative numbers, rerun with yours.)
Example A: 2 hours a day, solo use
- Cloud: hourly billing x 60 hours lands at tens of dollars a month
- Local: $4,999 / 36 months is about $139 a month plus power
→ Light use favors cloud
Example B: a team running 3 agents around the clock
- Cloud: always-on instances x 3 plus traffic, hundreds to thousands monthly
- Local: even two linked nodes amortize to hundreds monthly, and data never leaves
→ The more you keep on, and the more sensitive it is, the more local wins
Privacy multiplies in. Data staying inside rarely shows on the invoice, yet one contract line moves thousands of dollars. It is the isolation story from the security layers post pushed down into hardware.
Meeting open models: grade on your tasks, not the leaderboard
Local boxes run open weights. The picking method is in the open-weights comparison. Candidates like MiMo at 46, GLM-5.3 at 45, and Kimi K3 at 44 on the Intelligence Index.
And measurement went local too. AA-AgentPerf-Local, released September 29, times open-model agent runs on DGX Spark, Ryzen AI Halo, MacBook Pro (M5 Pro), and RTX 5090. "How many minutes does my box need to finish a job" is now an official bench. Use the p95 and concurrency method from the AIPerf post plus the time-to-task idea from the speed guide.
A local agent bar (sample):
[ ] Do your 5 representative tasks succeed 3 nights in a row
[ ] Is per-task time fine by next morning
[ ] Are noise, heat, and power acceptable (the desk-server reality)
[ ] Who updates models and backups (name the owner)
CodeBridge Mini Lab: measure for 2 weeks before buying
1. Pick 1 agent you already run in cloud
2. Log for 2 weeks:
- run hours (hours per day), projected monthly cost
- share of data that must not leave
- jobs you would run overnight
3. Call it:
- 8+ hours a day with 30%+ sensitive data → review local
- occasional with nothing sensitive → stay on cloud
- unclear → trial small local first (laptop or RTX)
It is the cost-per-success method times hours kept on. The box bills once, cloud bills monthly. Your usage pattern decides.
Conclusion: clock your hours before you shop
The order goes like this.
Measure usage hours, score sensitivity, count always-on agents, then look at price tags.
DGX Spark 64GB means "a 24-hour agent at home," not "100B on a desk." New question, new answer. Buy or skip comes second. Hours kept on comes first. Two weeks of logs give you the answer.
Further reading
- New open-weights race: how to pick MiMo, GLM-5.3, and Kimi K3
- AI model speed and latency guide
- Judge AI model cost per successful task, not token price
References
- NVIDIA Blog: DGX Spark 64GB Gives Developers More Ways to Build and Scale Local AI
- Artificial Analysis: AA-AgentPerf-Local
- NVIDIA: Benchmarking LLM Inference at Scale with AIPerf
Go deeper with a course
Local or cloud, running an agent overnight with confidence starts with verification loops and permission design. This course builds harnesses in Claude Code on real projects, exactly like the 24-hour agent story here.