Sakana AI's Fugu orchestration model, which bundles multiple AI systems behind a single API, has gained a new endpoint built specifically for cyber defense. Called Fugu-Cyber, it was published on July 21, 2026, and posts higher scores than frontier models such as GPT-5.5-Cyber and Mythos-Preview on the field's main security benchmarks. Access requires an application and manual review, and pricing runs through a usage-based token plan.
Many Models, One Endpoint
Fugu-Cyber rests on the multi-agent orchestration approach that Fugu has used from the start. A pool of specialized agents runs underneath, but developers see only one endpoint. A request comes in, the system picks the mix of agents that fits the problem, and the work proceeds across multiple steps.
The design has a purpose beyond raw performance. Handing a business process to a single vendor's model means price changes, discontinuation, or shifts in quality land directly on operations. By switching among several models internally, Fugu thins out that dependency. Fugu-Cyber keeps that character while its internals have been rebuilt around the complexity of modern cyber defense.
Ahead of Frontier Models on Two Benchmarks
The numbers are clear. On CyberGym, which measures the ability to verify real-world vulnerabilities, Fugu-Cyber scored 86.9%. GPT-5.5-Cyber reached 85.6% and Mythos-Preview 83.1% on the same metric. On CTI-REALM, which measures the ability to turn raw threat intelligence reports into working detection rules, Fugu-Cyber scored 72.1%, ahead of GPT-5.5-Cyber at 67.3% and Mythos-Preview at 68.5%.
The two benchmarks look at different sides of enterprise defense. CyberGym asks whether a system can read through a large, tangled codebase and confirm that a flaw actually holds. CTI-REALM asks whether fragmentary threat information arriving from outside can be translated into something a monitoring stack can use. Together they reflect how defensive work runs on two tracks: reading code, and turning intelligence into operations.
Benchmark Scores Are Only a Starting Point
What stands out is that Sakana AI itself spells out the limits of those scores. The company argues that debate over the cyber capabilities of frontier models has run hot, and it pushes back on the notion that gaining access to a powerful model resolves an organization's security problems on its own.
Its evidence is the situation inside large organizations, major financial institutions included. Even with a model that has state-of-the-art capability, a company without specialized internal talent and without deep integration into its own proprietary code will not get as far as uncovering and patching real vulnerabilities. Sakana AI says it has run into the same wall repeatedly while working with large Japanese enterprises.
A raw model run in isolation produces false positives. It flags something as a possible vulnerability without understanding the specifics of a live production environment. So Sakana AI emphasizes a flow in which a suspicious finding is validated by sub-agents specialized in security, then confirmed by a human, and only judged as a flaw that would actually trigger before a patch is proposed. The model is a powerful component, not the solution itself.
Application-Only Access and an Updated Usage Policy
The Fugu-Cyber API is offered under the token plan. Access is not open to anyone, however. Users must submit an application form describing their intended use case along with verified contact information, and the company's team reviews and approves each request individually. The acceptable usage policy has also been updated to prohibit offensive misuse.
Because offensive and defensive techniques share the same ground, a model sharpened for defense can be turned toward attack just as easily. The weight of the process is the flip side of that. Anyone considering the API for production work will need a concrete operational plan in place before applying.
Summary
Fugu-Cyber is a new Sakana AI API endpoint that keeps the multi-model orchestration structure while specializing in cyber defense. It scored 86.9% on CyberGym and 72.1% on CTI-REALM, ahead of competing frontier models. What the company stresses, though, is not the scores themselves but a practical boundary: the technology works in the field only when paired with specialized talent and verification processes. For anyone weighing adoption, how to structure internal operations matters as much as which model to pick.
