News · Science & Technology
GitHub Copilot is going local — but Microsoft won’t say what gets sent to the cloud
Inference is the model's process of interpreting a coding request and producing an answer, such as code or a shell command. With local inference, that work runs on the developer's computer. With cloud inference, Copilot sends the task to a remote model. GitHub plans to let Copilot choose between them automatically. In Copilot CLI, the Copilot app, and VS Code, developers can use Auto routing or select a local model. One local option is MAI Code 1.1 Flash through Windows ML. Developers can also use Axiom-compatible local endpoints. Microsoft says routing considers task context and cache state, including during multi-turn sessions. The distinction matters for privacy, speed, hardware, and connectivity. Selecting a local model keeps inference on the device, but tools may still contact external services. Microsoft also has not explained how much repository context Auto sends to cloud models, so local and fully offline are not equivalent.
Based on reporting by The New Stack
What does it mean for GitHub Copilot to run inference locally or in the cloud?
Inference is the model's process of interpreting a coding request and producing an answer, such as code or a shell command. With local inference, that work runs on the developer's computer. With cloud inference, Copilot sends the task to a remote model. GitHub plans to let Copilot choose between them automatically.
In Copilot CLI, the Copilot app, and VS Code, developers can use Auto routing or select a local model. One local option is MAI Code 1.1 Flash through Windows ML. Developers can also use Axiom-compatible local endpoints. Microsoft says routing considers task context and cache state, including during multi-turn sessions.
The distinction matters for privacy, speed, hardware, and connectivity. Selecting a local model keeps inference on the device, but tools may still contact external services. Microsoft also has not explained how much repository context Auto sends to cloud models, so local and fully offline are not equivalent.
How will Copilot's Auto routing decide whether a coding task uses a local model or a cloud model?
Auto routing is part of Project HydraFusion, which already selects models for coding tasks. GitHub is expanding it to decide not only which model to use, but also where that model runs. A task may therefore move between local and cloud inference as the session develops.
Microsoft says Copilot will consider task context and cache state. This applies during multi-turn sessions, where earlier requests, files, and tool results can change what the agent needs. Auto routing is available in Copilot CLI, the Copilot app, and VS Code. Developers can alternatively choose a local model themselves.
The exact routing policy remains unclear. Microsoft has not said what context crosses to the cloud, whether developers can inspect routing decisions, or whether Auto can be locked to local inference. That uncertainty matters most to teams with strict data-handling rules, because they cannot yet verify every routing outcome.
What repository code, conversation history, or tool results might still be sent to the cloud when Copilot uses Auto routing?
When Auto routes a task to a cloud model, the central unanswered question is how much working context travels with it. Microsoft has not specified the amount of repository context or conversation history sent to cloud models. That leaves teams unable to determine exactly which code or discussion may leave their environment.
The uncertainty also covers information accumulated during an agent session. Copilot considers task context and cache state, while the cache grows as the agent reads files and receives tool results. The article does not establish that every cached item or tool result is transmitted, but it identifies these as part of the session's working state.
Microsoft also has not said whether developers can inspect routing decisions or restrict Auto to local inference. Selecting a local model keeps inference on the device, but Auto's cloud behavior remains insufficiently documented for strict data-handling policies. Teams therefore still lack a clear account of repository data exposure.
How much computer memory does the local MAI Code 1.1 Flash model require, and why can a 53GB model need far more than 53GB of memory?
MAI Code 1.1 Flash is compressed to 53GB, but that is only the model's weight storage. Microsoft measured peak memory use of 75.5GB with a 256K-token context on an NVIDIA RTX Spark Windows PC. That machine offers up to 128GB of unified memory, making it suitable for the initial rollout.
Memory also holds the operating system, running applications, inference runtime, and key-value cache. The cache stores information needed during an ongoing session. It grows as Copilot reads files and receives tool results. Longer coding sessions therefore need more memory, even though the model's weights remain 53GB.
Most developer laptops have only 16GB or 32GB of RAM, according to the article. Those machines cannot comfortably match the reported setup. Developers must budget well beyond 53GB, especially for long contexts and tool-heavy sessions. The memory requirement is therefore a major practical limit on local Copilot use.
Why does choosing a local model not necessarily make a Copilot session fully offline?
Choosing a local model controls where inference happens, not everything the Copilot agent can do. The model may run entirely on the developer's computer, while the agent's tools continue making network requests. External services can therefore remain part of the session.
For example, Copilot could use local inference while a tool accesses an external service or communicates through a network-enabled command. The article states that local model selection does not stop the agent from reaching external services or making network requests through its tools. Inference location and tool connectivity are separate controls.
Developers who need a fully local session must also restrict what those tools can access. Sandboxing can apply operating-system restrictions to shell commands and supported local MCP servers, but remote MCP servers remain outside the local process sandbox. Local inference improves where model computation occurs; it does not by itself guarantee that code, results, or requests stay on the device.
How do Copilot's sandbox controls protect shell commands, file tools, local MCP servers, and remote MCP servers differently?
Copilot does not protect every tool in the same way. Shell commands and, where supported, local MCP and language servers receive operating-system-enforced restrictions. These policies apply whether the coding task uses a local model or a cloud model. The MXC library provides the sandboxing system.
Built-in file tools work differently. They run inside the agent process, so the agent harness checks each request against the sandbox policy rather than relying on OS-level isolation. Remote MCP servers are different again: they remain outside the local process sandbox. The article therefore separates local execution from remote service boundaries.
The platform also changes the enforcement mechanism. Windows uses the BaseContainer tier of the ProcessContainer backend, macOS uses Seatbelt, and Linux uses bubblewrap. These controls can limit local commands and supported servers, but they do not automatically make every tool offline or isolated. Remote MCP servers especially require separate consideration.
What are model quantization and speculative decoding, and how do they make large language models smaller or faster on local hardware?
Quantization reduces the numerical precision used to store model weights. That makes a large model take less space and fit on smaller hardware. Microsoft used mixed-precision quantization at roughly 3.3 bits per weight, shrinking MAI Code 1.1 Flash to 53GB, an 80% reduction from its bfloat16 cloud version.
The model also uses speculative decoding to improve speed. A smaller drafter proposes blocks of tokens, and the main model verifies them. When proposals are accepted, the main model can generate output more efficiently than producing every token alone. This approach addresses speed rather than primarily reducing storage.
These techniques make local inference more practical, but memory remains a serious constraint. Microsoft measured 75.5GB of peak memory at a 256K-token context, including more than the weights. The quantized model scored 70.8% on SWE-Bench Verified versus 72.6% for full precision, while benchmark results still include important limitations.
Key Facts:
📌 Local inference runs the model on the developer's device.
📌 Cloud inference runs the model remotely.
📌 Local inference does not make the session offline.
📌 Project HydraFusion already selects models for coding tasks.
📌 Auto routing considers task context and cache state.
📌 Developers can choose a local model instead.
📌 Microsoft has not disclosed how much repository context Auto sends.