Data center operator Equinix has announced Equinix Inference Exchange, a distributed AI inference program for enterprises. It brings NVIDIA's validated architectures and Together AI's inference platform together on top of the company's own data centers across 77 metros worldwide. General availability is scheduled for the first quarter of 2027.
The Argument Has Shifted from Training to Where Inference Runs
Conversations about generative AI tend to center on model performance, but what actually matters on enterprise ground is placement: where a finished model gets executed. Response latency, cost per token, and whether data ever leaves the country are largely decided by the location where inference runs.
Equinix is proposing to solve that placement problem through interconnection infrastructure rather than a single cloud. The company operates more than 280 data centers across 77 metros, with 230 cloud on-ramps and over 10,500 connected businesses. Eight of the top 10 AI model providers and nine of the top 10 AI cloud providers are deployed on its facilities, and that vendor-neutral position has long been its selling point.
The announcement was made in Redwood City, California, at Equinix Horizon, the company's first event for customers and partners.
Power, GPUs, and Model Platform Split Into Three Layers
What makes Inference Exchange easy to follow is how cleanly the responsibilities are divided into three layers.
Equinix handles the foundation: power, advanced cooling, and operational support services. These connect through Equinix Fabric, the company's interconnection service, to the clouds, networks, and AI providers a workload needs.
On top of that, NVIDIA supplies its validated Enterprise Reference Architectures along with infrastructure optimized for AI factories. The goal is to maximize throughput and minimize cost per token, sparing enterprises the work of designing a configuration from scratch.
The top layer, the model platform, is where Together AI enters as a newly added partner. Its platform supports more than 200 open source models and offers both a multi-tenant configuration that gains efficiency through shared use and a single-tenant configuration for workloads that require dedicated resources.
Vipul Ved Prakash, co-founder and CEO of Together AI, commented to the effect that enterprises should not be forced to choose between model performance and operational flexibility. Raj Mirpuri, vice president of the global AI cloud infrastructure ecosystem at NVIDIA, said the combination lets enterprises put intelligence closer to their data, applications, and customers.
Three Intended Use Cases
Equinix names three use cases.
The first is metro edge inference, running inference near users and data to squeeze out latency. The second is migration to open models, giving enterprises that are moving away from closed proprietary models for cost control or to avoid vendor lock-in a production-ready destination. The third is sovereign AI, letting companies in regulated industries or specific regions run workloads in locations that satisfy data residency and sovereignty requirements.
All three are framed around where a workload can sit rather than how fast it runs, which speaks to the problems enterprises hit once AI moves from experiment to production. Nick Patience, who covers AI platforms at The Futurum Group, noted that as workloads spread across providers and data sources, performance, cost, and governance become management issues.
Fabric One Takes On the Connectivity Side
At the same event, Equinix also announced Equinix Fabric One, a managed service that automates connectivity itself. Specify what you want to connect to and under what requirements, and the service works out and provisions the necessary links, including routing, encryption, redundancy, and failover. Requests can come through a portal or API, and also via automation workflows, agent-based requests, or natural language prompts.
The foundation is OpenAPI 3.0 Interconnect, an open source connectivity specification developed jointly by AWS and Google Cloud. The aim is a state where applications and AI agents can discover, request, and provision connections programmatically without human involvement, a design clearly aimed at enterprises shifting toward agentic architectures.
Fabric One enters beta in the second half of 2026, with general availability planned in North America during 2027. Paired with Inference Exchange in the first quarter of 2027, both pieces land next year.
Summary
Equinix Inference Exchange reframes AI inference as a question of which facility in which metro, rather than which cloud. Equinix brings power and connectivity, NVIDIA brings the configuration templates, and Together AI brings the open model runtime, which cuts down what enterprises have to assemble themselves. Availability starts in the first quarter of 2027, and the real picture will emerge next year alongside Fabric One on the connectivity side.
