Service eligibility and regional restrictions
PuppyIP serves only compliant overseas businesses and their authorized personnel. Proxy services are not available in mainland China. The service may only be used for lawful business activities outside mainland China. Use of this service within mainland China is prohibited.
Hosting a proxy IP or server overseas does not change these restrictions. The service must not be provided to end users in mainland China through relaying, forwarding, sharing or resale. Before use, read the Terms of Service.
Key Takeaways
- HydraFusion is available as a research preview through /experimental in GitHub Copilot CLI. GitHub says all Copilot plans can try it.
- Enable it by running /update, /experimental on, then /model and selecting HydraFusion (Research Preview).
- Usage sums tokens consumed by each model HydraFusion calls, charged at each model's standard rates. A multi-model workflow cannot be costed as one request.
- GitHub's TerminalBench, DeepSWE and CheckpointBench results are controlled offline evaluations, not proof of better, faster or cheaper results in every repository.
- It currently best suits substantial, well-scoped, first-turn, single-prompt tasks. Long multi-turn sessions remain a future focus.
- Research-preview models, routing, names and availability can change. Return to a known fixed model and disable experimental if quality, cost or latency crosses limits.
HydraFusion is runtime orchestration, not a new fixed model ID
GitHub's official blog introduced Project HydraFusion on September 4, 2026. A user selects HydraFusion once in Copilot CLI, then the runtime chooses Single, Cascade or Critique: direct completion by one model; an efficient draft escalated through a quality gate; or read-only review by another model family followed by revision.
The same prompt therefore does not guarantee the same model, number of calls or latency every time. Results, model pools, workflows, naming and availability can change in research preview. Production automation needing reproducible audits should retain CLI version, execution time, plan, task, output differences and actual usage, not merely the HydraFusion entry-point name.
Three activation commands, after checking policies and versions
The official sequence is /update for the latest Copilot CLI, /experimental on, then /model to select HydraFusion (Research Preview). If absent, first record CLI version, signed-in account, Copilot plan and organization policy. A missing research-preview entry is not automatically a network fault.
GitHub says all Copilot plans can access the preview, but administrators may still control clients, features and model use. 401, 403, a missing model list and an unknown command can concern identity, organizational policy, product availability and client versions respectively. Complete read-only checks before upgrading the client or contacting an administrator.
Costs must include every workflow branch
HydraFusion usage is based on tokens consumed by the models it calls at their standard rates. Cascade may try a cheaper model before escalation; Critique adds independent review and revision; retries and fallbacks may add usage. One visible response does not mean one model call on the bill.
Pilot records should include input size, output tokens, workflow stages, total time, official usage, estimated cost, retries and final acceptance. Set a token budget, wall-clock timeout and cancellation conditions for each run. Without full usage, multiplying one model's price by one input/output pair is not an actual cost calculation.
Reading official benchmarks: gains and counterexamples
GitHub reports that in controlled offline evaluation against Claude Opus 5, HydraFusion's verified task quality on TerminalBench 2.1 was 4.9 percentage points higher with 67% lower estimated cost. On DeepSWE, quality was 1.5 points lower and cost 36% lower; on CheckpointBench, quality was 0.1 points lower and cost 65% lower.
Those figures depend on benchmark versions, model pools, routing, pricing assumptions and a shared medium reasoning setting. A 67% cost reduction is not a discount for every task, and one benchmark's quality gain cannot be generalized to your repository. Decide using the same internal tasks' correctness, change scope, review hours, latency and total cost.
A/B test clearly scoped first-turn single-prompt tasks
GitHub currently recommends substantial, well-scoped, first-turn, single-prompt coding tasks. Select ten to twenty tasks with automated tests or explicit human acceptance, and give HydraFusion and a known fixed model identical repository commits, tool permissions, timeouts and prompts. Do not mistake environment differences for model differences.
Compare first-pass rate, final tests, wrongly edited file counts, human repair time, first-byte and total latency, tokens and total cost. Do not let both approaches edit the same working tree. Use isolated branches or disposable worktrees and retain human confirmation before external writes.
Do not combine login failures and model quality into one network issue
If /experimental and /model appear normally but answer quality, edit scope or cost fails acceptance, investigate preview behavior and task fit; another exit will not improve it. For Copilot CLI DNS, TLS, 407 or timeout failures, check proxy variables, Kerberos and enterprise CAs with the GitHub Copilot enterprise proxy guide, and preserve evidence using the proxy connection checklist.
A fixed exit can make authorized GitHub/Copilot network paths more consistent. It cannot grant preview eligibility, bypass organization policies, improve model quality or reduce official token prices. Check identity and permissions first for 401/403; enter the network layer for DNS/TLS/407/timeouts.
Stop and roll back: treat the experiment switch as a real switch
Stop expansion immediately for total cost above budget, excessive P95 latency, falling test pass rates, edits outside allowed files, excessive variation across similar tasks, unexplained retries or unknown external-write outcomes. Preserve CLI version, task ID, redacted logs, usage and diffs. Do not rerun potentially consequential tasks while their outcome is unknown.
Switch back to a verified fixed model through /model, disable /experimental if needed, and restore the isolated worktree or commit from before the task. Expand only after multiple consecutive batches meet quality, cost, latency and permission boundaries. Reestablish baselines after research-preview changes rather than reusing old conclusions.
Sources
Frequently Asked Questions
Is HydraFusion a new Copilot model?
It is not a fixed model. It is a research-preview runtime orchestration entry point that selects Single, Cascade or Critique workflows and may call multiple models.
How do I enable HydraFusion in Copilot CLI?
Follow /update, /experimental on, then select HydraFusion (Research Preview) in /model. If missing, check version, account, plan and organizational policies first.
Is it available on every Copilot plan?
GitHub's blog says the research preview covers all Copilot plans. Still verify organization policies, client versions and staged availability in your account.
Is HydraFusion always 67% cheaper?
No. 67% is an estimate against Opus 5 in a specific controlled TerminalBench 2.1 evaluation. Quality and costs differ across benchmarks; retest actual business tasks and tokens.
Which tasks should try it first?
GitHub recommends substantial, clearly scoped, first-turn single-prompt coding tasks, with automated tests, budgets, timeouts and recoverable worktrees.
Can a fixed IP reveal HydraFusion or improve quality?
There is no such guarantee. Fixed exits address connections supported by DNS, TLS, 407 or timeout evidence, not plans, policies, preview eligibility, model quality or official prices.