GitHub has released a research preview called Project HydraFusion, which combines AI models from several providers at runtime to produce a single answer. It can send a draft to a model from a different family for review, or hand a stubborn task up to a stronger model, and it decides all of this out of the developer's sight. In internal evaluations it held quality on par with Claude Opus 5 while lowering estimated cost by as much as 67 percent.

Choosing the approach, not just the model

GitHub Copilot already had Auto model selection, which reads a request and assigns the best-suited model. HydraFusion pushes that idea a step further: what gets selected is no longer a single model but the shape of the work itself. For each request it builds an execution plan and assigns roles, a drafter, a critic, and an escalation path if one is needed, drawing from models across multiple providers.

Nothing extra is asked of the user. You pick HydraFusion from the model list, and the workflow is settled internally while balancing quality, cost, and latency. GitHub frames this as part of a broader effort toward automated semantic routing between local, cloud, and compound models.

Three execution patterns

HydraFusion treats workflow selection as an optimization problem, using capability signals for reasoning, code generation, debugging, and tool use to pick the lightest approach that still clears the quality bar. Three patterns are available today.

Single has one model solve the task outright. For straightforward requests, skipping extra legs is both faster and cheaper.

Cascade lets an efficient model draft first, then a quality gate decides whether to accept the result or escalate to a stronger model. Easy requests stay cheap, and harder ones keep an exit route.

Critique brings in a model from a different family as a read-only reviewer, after which the drafting model revises once. The reviewer runs in an isolated, tool-less context, so nothing in the repository changes during review.

Intermediate drafts stay hidden until a result is final. The reasoning is that work which may still be discarded should not look finished, though GitHub acknowledges that waiting without visibility is a real trade-off and lists better progress reporting as an open item.

Benchmarks: close to Opus 5 at a fraction of the cost

Fixed routing policies were measured across three agentic coding benchmarks. Claude Opus 5 and GPT-5.6 Sol served as baselines, all models were evaluated at the same reasoning level, and the cost figures account for every leg, including drafting, critique, revision, escalation, retry, and fallback.

Benchmark Cost vs. Opus 5 Quality vs. Opus 5
TerminalBench 2.1 67 percent lower 4.9 points higher
DeepSWE 36 percent lower 1.5 points lower
CheckpointBench 65 percent lower 0.1 points lower

On TerminalBench 2.1, which covers multi-step work in terminal environments, HydraFusion beat the baseline on quality while spending a third as much. DeepSWE, which requires navigating large codebases and landing end-to-end fixes, leaves it 1.5 points short, but a 36 percent cost reduction for that gap will be worth taking in plenty of situations. CheckpointBench is an internal set built from real Copilot sessions, each anchored to a public repository and a fixed commit so that runs can be replayed.

The routing policy itself was not hand-tuned. GitHub used beam search to explore candidate policies. The development record also notes that two harness failures between August 11 and August 25 produced invalid runs, which were excluded before the gains continued.

How to try it

HydraFusion is available on every GitHub Copilot plan as an experimental feature in GitHub Copilot CLI. Billing follows the tokens each model consumes, at that model's standard rate.

/update
/experimental on
/model

Pick HydraFusion (Research Preview) from the list and you are set. For now the sweet spot is well-scoped coding work that fits in a single prompt. Support for longer, iterative sessions is the next step, and GitHub notes that names, behavior, and availability may all shift as the research continues.

Summary

HydraFusion is not an argument about which model is strongest. It is an argument that combining models well can deliver the same outcome for less. Beating the baseline on TerminalBench 2.1 while spending a third as much suggests there is still room in that direction. These are research preview numbers and will not map cleanly onto day-to-day development, but they do show the competition in coding assistants widening from raw model performance to how the work is orchestrated.