On October 7, a one-paragraph improvement landed on the GitHub Changelog. Starting in CLI version 1.0.94-0, /model discovers supported models from a running Ollama instance.
It looks small but matters. Cloud and local models can now be picked inside the same terminal workflow. The docs draw one firm line though: picking a local model does not switch on offline mode. This post separates what works from what does not.
What changed: local models appear in /model
The new flow:
Before: /model → cloud models + manually configured models
Now: /model → cloud + configured + discovered Ollama models
1. Pick a discovered model
2. Review its provider and endpoint
3. Confirm "Add and use for this session" or "Add without switching"
4. Use it in the current session, no restart
Discovery is not registration. You pick, verify where it comes from, then decide whether this session uses it. Connection failures appear inside the picker with reasons, so you learn what is stuck without leaving the terminal.
Two prerequisites to memorize:
[ ] Ollama and the model must already be installed (this flow installs nothing)
[ ] Models must support tool calling and streaming
It installs no runtime and downloads no models. It picks usable ones from the Ollama already running.
Correcting the misunderstanding: local pick is not offline
The official sentence is blunt.
Picking a local model → does NOT enable offline mode
Picking a local model → does NOT disable GitHub telemetry
Offline is a separate choice: COPILOT_OFFLINE=true
Even in offline mode, a remote provider still receives prompts and code context over the network. If COPILOT_PROVIDER_BASE_URL points remotely, transmission happens regardless of the offline declaration. Full isolation holds only when the provider is local too. As the four-layer security post says, judge boundaries by actual transmission paths, not declarations.
| Choice | Telemetry | Prompt path |
|---|---|---|
| Cloud model | On | Cloud via GitHub |
| Local model (default) | On | Local Ollama (telemetry separate) |
| Local + COPILOT_OFFLINE=true | Off | Local only (when provider is local) |
| Remote BYOK + offline declared | Off | Sent to the remote provider (careful) |
The habit to build is not "local means safe" but "where does what leave to."
Background: from BYOK to automatic discovery
This did not appear from nowhere. In April, Copilot CLI opened BYOK and local-model support: plug a provider in with environment variables.
# The older manual way (BYOK) to attach local Ollama
export COPILOT_PROVIDER_BASE_URL=http://localhost:11434
export COPILOT_PROVIDER_TYPE=openai
export COPILOT_MODEL=llama3.2
Provider types are openai, azure, and anthropic, with openai covering OpenAI-compatible endpoints like Ollama, vLLM, and Foundry Local. Local Ollama needs no API key.
The October update pulls that manual setup into the picker. No config file, just discover, review, and add from /model. It mirrors the Copilot app's Settings > Model providers experience. One goal: keep model choice inside the workflow.
What comes next: Auto orchestration
The changelog's last lines preview what is next. Around month's end, intelligent routing should decide between local and cloud inference itself. Combined with the Microsoft Command Line blog, the picture forms.
Now: developers pick directly (local choice in /model)
Next: Auto picks (a HydraFusion orchestrator weighing
performance, cost, and latency)
├─ fits on-device → local (e.g. MAI Code 1.1 Flash)
└─ needs scale → cloud
On RTX Spark machines like Surface Laptop Ultra,
local coding plus hardware acceleration owns the edge experience
That is the automated version of the model routing post. The human-picked picker arrives first, the machine-picked orchestrator second. Remember the order. You need direct-pick experience before automation, so you notice when Auto behaves oddly.
Note that the same week's release (v1.0.94-3) added Claude Haiku 5.5 to model selection. Local discovery and a cheap cloud model now share one picker. Keep the Haiku 5.5 pricing beside it.
Trying it: env vars and a comparison run
Setup follows the official flow.
# 1. Check Ollama is running
ollama list
# 2. Discover inside the CLI (1.0.94-0+)
/model
# → pick a discovered local model → review provider/endpoint → add
# 3. Declare offline separately if needed
export COPILOT_OFFLINE=true
CodeBridge Mini Lab: one fix, two models
1. Pick 1 code fix (e.g. refactor 1 function + pass tests)
2. Run once on cloud → log: accuracy (tests pass?) and latency
3. Run once on local (Ollama) → log: accuracy and latency
4. One comparison line:
[ ] same accuracy → default to local (data control + cost)
[ ] local wrong → classify where (tool calls? long context? reasoning?)
[ ] ambiguous → write an escalation rule to Auto or a top model
5. Offline check:
[ ] confirm no COPILOT_OFFLINE=true was assumed from a local pick
[ ] confirm the provider URL is local, not remote (the path matters!)
It attaches one model-comparison line to the Agent and Plan flow from the vibe-coding with Copilot post.
Conclusion: picking became the workflow
One line to close.
Not that local got easier, but that cloud and local are now picked in the same place.
The real change is not Ollama support itself. Model choice moved out of config files into the terminal session, with a warning attached. A local pick is not offline mode. Check the transmission path separately.
Today's job is small. Run the same fix once on cloud and once on local. Log accuracy and latency. That table decides whether to hand the next pick to Auto or keep choosing yourself.
Further reading
- Vibe coding with GitHub Copilot
- Why Claude Haiku 5.5 is 90% cheaper
- Why you must not run agents without permissions in the computer-use era
References
- GitHub Changelog: discover local models in GitHub Copilot CLI (Oct 7, 2026)
- GitHub Docs: adding LLM models to GitHub Copilot CLI (BYOK)
- Microsoft Command Line: local models and sandboxed tools (Oct 8, 2026)
- GitHub Changelog: Copilot CLI BYOK and local models (Apr 7, 2026)
Go deeper with a course
To attach Agent and Plan modes to a real project including model choice as flow, this course applies Copilot to Java and Spring projects, exactly like the comparison experiment here.