On October 2, NVIDIA put a computer on your desk. DGX Spark 64GB. Shipping October 23 through Acer, ASUS, Dell, Gigabyte, HP, and MSI, starting at $4,999.

The specs: GB10 Grace Blackwell chip, 64GB unified memory, up to 100B-parameter models on one box. Link two over ConnectX-7 and they share 128GB for up to 200B. NVIDIA reports up to 1.7x performance on Qwen 3.8 27B in the 2-node setup.

But spec talk is not the point. The question changed. From "can a small model run locally" to "is keeping a private agent always on cheaper than cloud."

What is inside: ecosystem over specs

Item Detail
Chip and memory GB10 Grace Blackwell, 64GB unified memory
Model scale Up to 100B solo, up to 200B with two linked units
Scaling ConnectX-7 link plus Sync Cluster Assistant auto setup
OS and tools DGX OS, Ollama, vLLM, llama.cpp, LM Studio, Agent Toolkit
Price and date From $4,999, shipping October 23

What matters is the defaults, not the box. Ollama, vLLM, and PyTorch just run, and NVIDIA names "24-hour coding and research agent runs" as the use case. Local AI is moving from "occasionally tinkered hobby" to "always-on assistant." That is close to a declaration.

One pricing caveat: reports on October 3 say the 128GB model jumped toward $6,950 amid a RAM shortage, so recheck prices at launch. The calculation method below works whatever the numbers say.

The economics: is keeping it on cheaper

One formula before buying.

Monthly cloud GPU cost x actual usage hours
  vs
(device price / months of use) + power + admin time
  x privacy needs x always-on agent count

A worked example for feel. (Illustrative numbers, rerun with yours.)

Example A: 2 hours a day, solo use
- Cloud: hourly billing x 60 hours lands at tens of dollars a month
- Local: $4,999 / 36 months is about $139 a month plus power
→ Light use favors cloud

Example B: a team running 3 agents around the clock
- Cloud: always-on instances x 3 plus traffic, hundreds to thousands monthly
- Local: even two linked nodes amortize to hundreds monthly, and data never leaves
→ The more you keep on, and the more sensitive it is, the more local wins

Privacy multiplies in. Data staying inside rarely shows on the invoice, yet one contract line moves thousands of dollars. It is the isolation story from the security layers post pushed down into hardware.

Meeting open models: grade on your tasks, not the leaderboard

Local boxes run open weights. The picking method is in the open-weights comparison. Candidates like MiMo at 46, GLM-5.3 at 45, and Kimi K3 at 44 on the Intelligence Index.

And measurement went local too. AA-AgentPerf-Local, released September 29, times open-model agent runs on DGX Spark, Ryzen AI Halo, MacBook Pro (M5 Pro), and RTX 5090. "How many minutes does my box need to finish a job" is now an official bench. Use the p95 and concurrency method from the AIPerf post plus the time-to-task idea from the speed guide.

A local agent bar (sample):
[ ] Do your 5 representative tasks succeed 3 nights in a row
[ ] Is per-task time fine by next morning
[ ] Are noise, heat, and power acceptable (the desk-server reality)
[ ] Who updates models and backups (name the owner)

CodeBridge Mini Lab: measure for 2 weeks before buying

1. Pick 1 agent you already run in cloud
2. Log for 2 weeks:
   - run hours (hours per day), projected monthly cost
   - share of data that must not leave
   - jobs you would run overnight
3. Call it:
   - 8+ hours a day with 30%+ sensitive data → review local
   - occasional with nothing sensitive → stay on cloud
   - unclear → trial small local first (laptop or RTX)

It is the cost-per-success method times hours kept on. The box bills once, cloud bills monthly. Your usage pattern decides.

Conclusion: clock your hours before you shop

The order goes like this.

Measure usage hours, score sensitivity, count always-on agents, then look at price tags.

DGX Spark 64GB means "a 24-hour agent at home," not "100B on a desk." New question, new answer. Buy or skip comes second. Hours kept on comes first. Two weeks of logs give you the answer.

Further reading

References

Go deeper with a course

Local or cloud, running an agent overnight with confidence starts with verification loops and permission design. This course builds harnesses in Claude Code on real projects, exactly like the 24-hour agent story here.