AI architecture • 005 · 8 October 2026 · Inference routing
Local or cloud?
Where should AI inference happen?
Maddipalli Gopalakrishna · AI / ML Engineer
45 seconds · music only · routed requests on screen are illustrative
Not every AI request needs to travel to the cloud. And not every AI request should run locally.
What Microsoft and GitHub announced
On 7 October 2026, Microsoft and GitHub said GitHub Copilot will “determine when a task is best handled by on-device intelligence and when it should leverage cloud-scale models”. On NVIDIA RTX Spark PCs such as Surface Laptop Ultra, that includes a local version of MAI Code 1.1 Flash, a quantized coding model. Developers can let Copilot's Auto mode route work between local and cloud models, which considers task context and cache state, or they can pick a local model or an OpenAI-compatible local endpoint themselves.
GitHub frames this as the next step for Project HydraFusion, its multi-model orchestrator (still a research preview): orchestration across compute environments, not just across models. Microsoft describes the wider Windows direction as “hybrid intelligence”, with agents that run locally when needed and connect to the cloud when appropriate.
Two lines from GitHub's post are worth keeping in mind. An agent's shell commands “inherit the access of the account running them. Moving inference onto the device doesn't change that.” And “local inference does not make the session offline.” Where a model runs and what an agent is allowed to do are separate questions.
From one route to a decision
The usual architecture has a single execution environment for every request:
user → application → cloud model API → responseA hybrid design adds a decision before execution:
request
→ policy check
→ inference router ─┬─ local model
└─ cloud model
→ output validation
→ response
→ observabilityChoosing the model is only part of the design. Choosing where that model executes becomes another architecture decision, with its own failure modes. The router isn't simply choosing the cheapest model. It's selecting an execution environment that satisfies the workload's requirements.
Five routing dimensions
| Dimension | What the router has to know |
|---|---|
| Capability | Can this model finish this task at the quality we need? Route on measured task success per task type, not on parameter count. |
| Latency | Time to first token, decode throughput and end-to-end task time. Local removes the network round trip, but cold loads, a growing KV cache and memory pressure can still make it slower. |
| Privacy and policy | Some data shouldn't leave the device, and a local path can help with residency rules. It doesn't replace access control, encryption, sandboxing, log policy or model supply-chain checks. |
| Cost | Cloud means API spend, GPUs, networking and operations. Local means hardware, energy, memory, model updates and fleet complexity. Compare total cost of ownership. |
| Reliability | Network loss, device resource pressure, model unavailability, timeouts and version skew. Fallback is allowed only inside policy. |
None of these has a fixed winner. Local isn't automatically faster, cheaper or more private; each depends on the workload, the hardware and how the system is built. GitHub's own post makes the memory point: model weights are only part of the budget once the context and KV cache grow over an agent session.
Policy before routing
The order matters. Data classification decides which routes are allowed at all; the router then chooses among the allowed routes using measured quality, latency budget, cost budget, device resources and model health. A sketch:
allowed = policy.allowed_routes(data_classification) # deterministic, first
if not allowed: return REJECT
candidates = [r for r in allowed
if measured_quality(r, task_type) >= threshold
and healthy(r) and resources_fit(r)]
if not candidates:
return DEFER or HUMAN_REVIEW # never widen 'allowed' to fall back
route = best(candidates, latency_budget, cost_budget)
log(route, inputs, model_version, policy_version)A learned scorer can rank candidates, but it should never be able to add a route that policy removed. A single LLM deciding every route is not a design I'd trust here. Some illustrative cases:
| Request | Route |
|---|---|
| A · Classify a short document | Low complexity, approved for local processing, local model meets the quality bar → LOCAL. |
| B · Analyze a multi-document problem | Needs long context and stronger reasoning, cloud permitted → CLOUD. |
| C · Summarize a confidential document | Restricted data, cloud prohibited → LOCAL. If local can't meet the task: defer, human review or a policy-approved failure. Never a silent hop to the cloud. |
| D · Respond while offline | No network, local model available → LOCAL. Otherwise fail gracefully; a cloud-only task is deferred, not forced onto a model that can't do it. |

The router needs evals too
A wrong route is a quiet failure. The answer still arrives, just from the wrong place, at the wrong cost, or against policy. Treat the router like any other production component:
| Measure | Why |
|---|---|
| Routing accuracy | Did the router pick the route a labelled eval set says it should have? |
| Task success per route | The same task run locally and in the cloud, scored the same way. |
| Latency per route | Including cold starts and long-context cases, not just warm averages. |
| Fallback rate | How often the preferred route fails, and what happens next. |
| Policy compliance | Restricted requests that reached a disallowed route. The target is zero, and it should be tested, not assumed. |
Keep an evaluation matrix per task type (local quality, cloud quality, latency on each route, cost, policy eligibility, selected route), fill it with your own measurements, and re-run it when a model, quantization, runtime or policy version changes.
Model selection determines capability. Inference routing determines how that capability is delivered.
Where would you draw the line in your systems: which requests should never leave the device?