Skip to content
Review

GitHub Copilot Review (2026): Computer Use, Free Local Sandboxing, and Two Model Retirement Waves

Published June 27, 2026 · Updated October 8, 2026 · by Pondero Reviews

4.5

The short version

GitHub Copilot can now drive desktop apps (public preview), local sandboxing went GA at no extra cost, Claude Haiku 5.5 joined the picker, and six more models retire October 19. What changed since September, which switches to flip, and whether the 4.5 rating holds.

Pros

  • ✓Local sandboxing is generally available in Copilot CLI, the Copilot app, and VS Code Agent Host sessions at no additional cost, with filesystem, network, and credential limits that enterprise-managed settings can make mandatory (per GitHub changelog, Oct 7 2026)
  • ✓Computer use lets Copilot click, type, scroll, and drag through desktop apps on macOS and Windows, which reaches GUI-only and legacy software that has no API, CLI, or MCP surface (per GitHub changelog, Oct 1 2026)
  • ✓Claude Haiku 5.5 is GA from Pro up, aimed at subagents, quick edits, and terminal tasks, and Claude Sonnet 5.5 also reaches Pro, so the $10 tier gets two current Anthropic models (per GitHub changelog, Sept 28 and Oct 7 2026)
  • ✓Copilot code review can now be requested over the REST and GraphQL APIs, so a team can trigger reviews from its own scripts and internal tools (per GitHub changelog, Oct 2 2026)
  • ✓Pricing held flat: Free $0, Pro $10, Pro+ $39, Max $100, Business $19 per granted seat, Enterprise $39 per granted seat (per GitHub Copilot plans docs, pulled Oct 8 2026)

Cons

  • ✕Computer use is public preview, asks for approval before it controls each app unless you choose Always allow, and on macOS needs Accessibility and Screen Recording permissions granted first (per GitHub changelog, Oct 1 2026)
  • ✕Local sandboxing is off by default, and GitHub's docs describe it as lighter-weight OS-level containment rather than a separate VM or container (per GitHub Docs, cloud and local sandboxes, pulled Oct 8 2026)
  • ✕Six more models retire October 19 (Gemini 3.7 Flash, GPT-5.5, GPT-5.4, GPT-5.4 mini, GPT-5 mini, Grok 4.5), the second retirement wave in seventeen days after four models left on October 2 (per GitHub changelog, Sept 18 and Oct 2 2026)
  • ✕Balanced became the default code review effort level on September 28, and GitHub estimates a Balanced review at $0.25 to $5 in AI credits versus $0.05 to $1 on Lite, so teams that never touched the setting now pay more per review (per GitHub changelog, Oct 2 2026, and GitHub code review docs, pulled Oct 8 2026)
  • ✕Agent mode still trails Cursor Composer on whole-repo multi-file refactors, the gap carried from prior reviews

GitHub Copilot Review (2026): Computer Use, Free Local Sandboxing, and Two Model Retirement Waves

Copilot can now take over your mouse. On October 1, GitHub put computer use into public preview in the Copilot CLI and the Copilot app, so the agent can read a desktop app's screen, click, type, scroll, and drag through software that has no API, no CLI, and no MCP server (GitHub changelog, Oct 1). Six days later, local sandboxing went generally available at no additional cost, fencing what agent-run commands can read, write, and reach on the network (GitHub changelog, Oct 7). Our September 7 version of this review rated Copilot 4.5 and predated both. The re-call after this month: 4.5 out of 5, held.

Computer use and local sandboxing are separate switches, and both ship turned off (computer use docs, sandbox docs, both pulled 2026-10-08). Enabling one does not enable the other. If you only flip computer use, the agent gains reach into your desktop while its shell commands still run with your full user account. Flip sandboxing first.

What changed since September 7

Ten dated items landed between September 18 and October 7. Read the release status on each, because the most interesting capability is a preview and most of the secret-detection item is not live for Copilot users yet.

  • Computer use, public preview (Oct 1). Copilot can drive desktop apps on macOS and Windows from the CLI or the Copilot app. Enable it with /computer on, or under Settings in the app (GitHub changelog, Oct 1). Our news brief covers the launch.
  • Local sandboxing, GA (Oct 7). Generally available in Copilot CLI, the Copilot app, and VS Code sessions that use Agent Host, included at no extra cost (GitHub changelog, Oct 7). It moved up from the public preview GitHub opened on June 2 (GitHub changelog, June 2).
  • Claude Haiku 5.5, GA (Oct 7). Anthropic's newest lightweight model, on Pro, Pro+, Max, Business, and Enterprise, billed at provider list pricing under usage-based billing (GitHub changelog, Oct 7).
  • A purpose-built secret-detection model (Oct 7). GitHub's fine-tuned model reads surrounding code to find likely credentials, including passwords with no recognizable token format. Today it powers AI-detected alerts for GitHub Secret Protection and Advanced Security customers at no extra charge. The Copilot /security-review secret checks are "available soon in private preview" and will bill AI Credits, so Copilot buyers have nothing to buy yet (GitHub changelog, Oct 7).
  • Local model discovery in the CLI (Oct 7). From CLI version 1.0.94-0, /model lists models from a running local Ollama instance next to the cloud models (GitHub changelog, Oct 7).
  • Four models retired (Oct 2). Gemini 3.5 Flash, Gemini 3.6 Flash, Kimi K2.7 Code, and Claude Opus 4.7 left every Copilot surface (GitHub changelog, Oct 2).
  • Code review API and a new default effort level (Oct 2). Reviews can be requested over REST and GraphQL, and Balanced replaced Lite as the default effort level on September 28 (GitHub changelog, Oct 2).
  • GPT-6.1 Sol, GA (Sept 29). OpenAI's newest model for agentic coding and terminal work, on Pro+ and up (GitHub changelog, Sept 29).
  • Claude Sonnet 5.5, GA (Sept 28). Anthropic's model for well-scoped feature work and bug fixes, from Pro up (GitHub changelog, Sept 28).
  • Six more retirements scheduled (Sept 18). Gemini 3.7 Flash, GPT-5.5, GPT-5.4, GPT-5.4 mini, GPT-5 mini, and Grok 4.5 go on October 19 (GitHub changelog, Sept 18).

This is a capability month, so the verdict below gets re-weighed on what Copilot can do and what it can be trusted to touch.

More reach and more containment, shipped in the same week

Read the two October headliners together. Computer use widens what the agent can act on. Local sandboxing narrows what its commands can damage. GitHub shipped them six days apart, and for a team deciding how much autonomy to hand an agent, the pairing matters more than either feature alone.

Computer use is for the work that never got an integration. GitHub's docs give the examples: reviewing data in a legacy desktop application, updating a presentation, entering records into GUI-only software, and moving information between apps as one workflow (computer use docs, pulled 2026-10-08). To see the screen, the agent reads the operating system's accessibility tree and takes screenshots when it needs visual context. GitHub's page also tells you to prefer an API, MCP server, terminal command, or browser tool whenever one exists, because those return structured results. Treat computer use as the fallback for the internal Windows app your finance team still runs, not as the default path for anything with an endpoint.

Approval works per application: when Copilot wants to control an app it asks first, and you can allow it for the session, save the approval, or deny it (computer use docs, pulled 2026-10-08). Saved "Always allow" decisions are stored locally and apply to both the CLI and the app on that machine. GitHub's own warning is blunt: the agent can pick the wrong control or type into the wrong field, and it advises against Always allow for apps holding sensitive data or high-impact actions.

Local sandboxing runs Copilot's shell commands, file search, and by default local MCP and language servers inside an operating-system-level sandbox (sandbox docs, pulled 2026-10-08). The CLI exposes controls for filesystem paths, outbound and local network access, whether Git and gh credentials are visible, the macOS keychain, and per-command exceptions. Backends differ by OS: Seatbelt on macOS 15 or later, bubblewrap 0.5.0 or later on Linux, and recent Windows 11 builds. GitHub is candid about the strength, placing it at the lighter-weight end of isolation, restricting what a process can read, write, and reach, without a separate virtual machine or container. Two more caveats from the same page: remote MCP servers are never sandboxed, and the CLI's built-in file tools run in-process, so they honor the sandbox policy on a best-effort basis rather than under OS enforcement.

What lifts sandboxing above a developer preference is the enterprise lever. Enterprise-managed settings can require sandboxing and enforce policies developers cannot weaken (GitHub changelog, Oct 7). Organization-managed settings can also disable computer use (GitHub changelog, Oct 1).

SwitchDefaultTurn onWho can lock itCostStatus
Local sandboxingOff/sandbox enable in the CLI; project settings in the appEnterprise-managed settings can require itNo extra costGA, Oct 7
Computer useOff/computer on in the CLI; Settings in the appOrg or enterprise managed settings can disable itUses AI credits like any CLI or app sessionPublic preview, Oct 1
Sourcessandbox docssandbox docs, Oct 1Oct 7, Oct 1Oct 7, plansOct 7, Oct 1

The order we would run on a fresh Copilot CLI session follows GitHub's own docs. Sandbox first, then computer use, then confirm computer use is on:

/sandbox enable
/computer on
/computer show

Commands per the sandbox docs and the Oct 1 changelog. Sandboxing persists across future sessions until you run /sandbox disable. Turn computer use back off with /computer off, and stop a runaway action by pressing Esc twice in the CLI.

The usage-billing math, with two new levers

Every Copilot plan includes a monthly credit allowance, and all model usage draws from it: $15 on Pro, $70 on Pro+, $200 on Max (Copilot plans page, pulled 2026-10-08). Frontier models still meter at provider list pricing. GPT-6 Astra and Claude Fable 5.1 did in September (Sept 4 changelog, Sept 1 changelog), and October's Haiku 5.5, Sonnet 5.5, and GPT-6.1 Sol do the same (Oct 7, Sept 28, Sept 29).

Haiku 5.5 is the first new lever, and it saves credits. It is built for subagents, quick edits, and terminal tasks, and GitHub says that in its early testing it matched Claude Sonnet 5 on many coding tasks with significantly fewer tokens and steps (GitHub changelog, Oct 7). Fewer tokens per task means more tasks per credit. Point routine agent turns at it and the Pro pool stretches further than it did a month ago.

Code review cuts the other way. Balanced became the default code review effort level on September 28 (GitHub changelog, Oct 2), and the gap is wide. GitHub's code review docs (pulled 2026-10-08) estimate a typical review at $0.05 to $1 in AI credits on Lite and $0.25 to $5 on Balanced, before GitHub Actions minutes. Both ends of the range are a fivefold jump. A team that reviews every PR and never touched the setting now pays the Balanced rate. Overrides run enterprise, org, repo, and personal, each level overriding the one above; our code review API news brief walks through it.

On Business and Enterprise, included credits pool org-wide with admin controls for spending (Copilot plans page, pulled 2026-10-08). Set the review effort level on purpose, route the subagent work to Haiku 5.5, and keep GPT-6 Astra for the long autonomous runs it was built for.

The model picker: what to default by task

Since September 7 the picker added Haiku 5.5, Sonnet 5.5, and GPT-6.1 Sol, and lost four models. Here is the pick by task rather than by benchmark, with each model's source and lowest plan.

Task typeReach forLowest planWhy
Subagents, quick edits, terminal tasksClaude Haiku 5.5ProGitHub reports Sonnet 5 parity on many coding tasks with fewer tokens (Oct 7)
Everyday chat, routine boilerplateGemini 3.8 FlashProIntroductory pricing through Dec 31 (Sept 3)
Well-scoped features and bug fixesClaude Sonnet 5.5ProBuilt for everyday scoped work (Sept 28)
Agentic terminal workflowsGPT-6.1 SolPro+OpenAI's newest; GitHub reports fewer tokens and steps than earlier GPT-6 and GPT-5.6 models (Sept 29)
Long-horizon autonomous agent runsGPT-6 AstraPro+Plans and validates as it goes (Sept 4)
Deep codebase research, long feature devClaude Fable 5.1Pro+Requires the retention policy enabled first (Sept 1)
Hard architectural callsClaude Opus 5.5See models docsGitHub's named replacement for the retired Opus 4.7 (Oct 2)
Cheap high-volume turnsKimi K3See models docsGitHub's named replacement for Kimi K2.7 Code (Oct 2)

September's Fable 5.1 caveat still applies. It requires data retention on by default so Anthropic can run its safety classifiers, GitHub says retained data is not used to train Anthropic's models, and the model policy ships off, so a Business or Enterprise admin has to opt in (Sept 1 changelog). If your compliance posture assumes zero retention, keep Fable 5.1 as a deliberate opt-in.

Local models are now one menu away too. In the CLI, /model lists models from a running Ollama instance, though GitHub is explicit that choosing one does not turn on offline mode or disable telemetry; offline mode stays a separate COPILOT_OFFLINE=true setting (GitHub changelog, Oct 7).

Two retirement waves in seventeen days

The October 2 wave is done. Four models left chat, inline edits, ask and agent modes, and completions, and GitHub's suggested replacements were Gemini 3.8 Flash, Kimi K3, and Claude Opus 5.5 (GitHub changelog, Oct 2). The October 19 wave is next, and it is bigger.

ModelRetiresGitHub's suggested replacement
Gemini 3.5 FlashOct 2 (done)Gemini 3.8 Flash
Gemini 3.6 FlashOct 2 (done)Gemini 3.8 Flash
Kimi K2.7 CodeOct 2 (done)Kimi K3
Claude Opus 4.7Oct 2 (done)Claude Opus 5.5
Gemini 3.7 FlashOct 19Gemini 3.8 Flash
GPT-5.5Oct 19GPT-5.6 Sol
GPT-5.4Oct 19GPT-5.6 Sol
GPT-5.4 miniOct 19GPT-5.6 Luna
GPT-5 miniOct 19GPT-5.6 Luna
Grok 4.5Oct 19Grok 4.6
SourceOct 2, Sept 18Oct 2, Sept 18

The two waves handle the swap differently. For October 2, GitHub told Enterprise admins they might need to enable the alternative models through model policy themselves (Oct 2 changelog). For October 19, the suggested replacements turn on automatically for Business and Enterprise under default model enablement, unless an admin has turned off the global default or disabled that model (Sept 18 changelog). If your org switched the global default off to control spend, you are back on manual duty.

One detail Free-plan readers should catch: the plans page, pulled October 8, still lists GPT-5 mini among the models the Free tier gets (Copilot plans page). GPT-5 mini is on the October 19 list, so expect the Free picker to change within two weeks.

Copilot code review can approve pull requests

Since September 1, Copilot adds an approval assessment to every review and, once an admin enables it, can submit an approval that counts toward the required-approvals rule (GitHub changelog, Sept 1). It is off by default, and control runs from enterprise to org to repository, down to the file paths Copilot may approve. New commits after a Copilot approval dismiss it, the same as a human's. With the October 2 API, a team can now request those reviews from its own tooling (Oct 2 changelog). Our recommended setup has not changed: approvals on low-risk paths like docs and test fixtures, human sign-off on core logic.

Pricing, still flat

The sticker held again (Copilot plans docs, pulled 2026-10-08). What varies is the credit meter, and this month two settings move it: the review effort level and which model your subagents use.

PlanPrice/moIncluded allowanceOctober note
Free$02,000 completions, limited chat and agent useGPT-5 mini in its picker retires Oct 19
Pro$10$15 in creditsGets Haiku 5.5 and Sonnet 5.5
Pro+$39$70 in creditsAdds GPT-6.1 Sol, GPT-6 Astra, Fable 5.1
Max$100$200 in creditsSame picker, largest pool
Business$19 per granted seatPooled creditsFree local sandboxing; Oct 19 swaps auto-enable
Enterprise$39 per granted seatPooled creditsCan require sandboxing; can disable computer use
Sourceplans docsplans pageOct 7, Sept 29, Sept 18

All prices and allowances pulled from GitHub's plans page and plans docs on 2026-10-08. New models roll out gradually, so a fresh seat may not show the full picker on day one.

Rating

Copilot holds at 4.5 out of 5 in October 2026, unchanged from September.

The case to move up is the strongest it has been. Free local sandboxing, with enterprise enforcement that developers cannot weaken, answers the main objection security teams raise about agent autonomy, and it shipped GA rather than as a preview (Oct 7 changelog). Computer use reaches a class of software no coding agent touched before (Oct 1 changelog).

Four things hold it back. Computer use is a public preview whose own docs warn it can click the wrong control. Sandboxing ships off and is OS-level containment, not a VM (sandbox docs). Ten models retiring across two waves in seventeen days is real churn for any team with pinned workflows (Sept 18, Oct 2). The editing gap also persists: agent mode still trails Cursor Composer on whole-repo multi-file refactors. If computer use reaches GA with the approval model intact, that is the change most likely to move this score up.

The pick, by who you are

October moves the decision from which model to route where, toward which switches to flip and in what order.

Solo developer. Pro+ at $39 is still the working tier if you want GPT-6 Astra, GPT-6.1 Sol, or Claude Fable 5.1 (Copilot plans page, 2026-10-08). What changed is the floor. Pro at $10 now carries Claude Haiku 5.5 and Sonnet 5.5, which makes it a real option for a developer whose work is scoped features and quick fixes (Oct 7, Sept 28). New sign-ups are being enabled gradually, per GitHub's plans page (checked 2026-10-08), so a brand-new account may have to wait a few days. On either tier, run /sandbox enable before you try /computer on. The flip condition is unchanged: if your day is all-day whole-repo editing, Cursor's Composer still edits cleaner on multi-file refactors. Our Cursor review covers its tiers and its own model churn.

Small team on Business ($19 per seat). The pick holds, with three admin tasks for this week. First, decide the code review effort level, since Balanced is now the default (Oct 2 changelog) and uses more credits per review than Lite (code review docs). Second, check whether your org turned off the global model default, because that decides whether the October 19 replacements turn on by themselves (Sept 18 changelog). Third, have every developer who runs the CLI or the app enable local sandboxing; it costs nothing (Oct 7 changelog). Start from the Copilot plans page.

Enterprise on Enterprise ($39 per seat). October widens Copilot's governance lead. Enterprise-managed settings can now require local sandboxing with policies developers cannot weaken, and managed settings can switch computer use off until it leaves preview (Oct 7 changelog, Oct 1 changelog). Our action list: require sandboxing org-wide, disable computer use except for a named pilot group, set the review effort level per org, and confirm the October 19 swaps landed in the picker. Re-check the Cursor comparison only for teams whose work is editor-bound. For everyone whose work runs through GitHub, October keeps Copilot the standardize-here pick.

Review history

This review is updated in place as the product changes. Earlier versions:

  • October 8, 2026: Computer Use, Free Local Sandboxing, and Two Model Retirement Waves (current version)
  • September 7, 2026: GPT-6 Astra, Four Models Retiring, and the Usage Billing Math
  • August 24, 2026: Agent Plugins 1.0, Grok 4.6, and Whether Slack Changes the Verdict
  • July 8, 2026: Vision GA, Kimi K2.7, and Whether the Credits Math Changed
  • June 27, 2026: Four Tiers, 1M Context, and the New Credits Math

Ready to try it?

Try copilot →