Equinix has announced Inference Exchange, a distributed AI-inference program it says will be available beginning in the first quarter of 2027. The September 2 announcement combines NVIDIA Enterprise Reference Architectures, Together AI’s inference platform and Equinix’s data-center and interconnection footprint. The immediate product is not a new foundation model. It is a deployment proposition: put inference near the data, users and applications that need it, while retaining a choice of providers and models.
That framing addresses an enterprise problem that is easy to understate. An organization can choose a capable model yet still struggle with data movement, latency, residency rules, security boundaries, network cost and integration with existing clouds. Inference is where a model responds to live requests, so a system that performs well in centralized testing can create disappointing user experience when prompts, retrieval data and applications are separated by geography or policy controls. Equinix is positioning infrastructure location as part of the AI architecture rather than a back-office procurement decision.
The company says Together AI’s platform supports more than 200 open-source models and describes three intended modes: metro-edge inference, migration from closed to open models, and sovereign AI workloads. Together AI would offer multitenant environments for shared efficiency and dedicated single-tenant environments for workloads requiring dedicated capacity. The model-choice element is strategically important. It can give enterprises options around cost, customization and vendor dependence, but it also creates a wider governance task. Every model route, deployment mode and data boundary needs consistent policy enforcement, logging and evaluation.
Equinix cites more than 280 data centers across 77 metros, 230 cloud on-ramps and more than 10,500 interconnected businesses. It also says eight of the top ten AI model providers and nine of the top ten AI clouds are deployed with Equinix. Those scale claims explain the appeal of a neutral exchange: an enterprise can connect to several clouds, networks and AI providers without building every link independently. They do not establish throughput, unit cost or latency for any specific workload. Those measures will need to be validated in production once the service is available.
The strongest use case is not necessarily the flashiest. Regulated organizations often need to know where data is processed and which party operates each layer. A distributed architecture can help place inference within a chosen region, but it does not automatically make a workflow compliant. Data residency depends on the full path, including prompts, retrieved context, logs, backups, support access, model-provider policies and cross-border failover. Enterprises should treat “sovereign AI” as an architectural objective requiring legal, security and operational verification, not as a property conferred by geographic proximity alone.
The program also sharpens a trade-off around open-model flexibility. Moving between models can reduce lock-in and support task-specific optimization. It can also multiply the work required to evaluate safety, accuracy, licensing, update cadence and vulnerability response. A common interconnection fabric may simplify routing, but it cannot decide which models are suitable for a high-stakes workflow. The governance layer has to travel with the workload: policy controls, identity management, audit trails, retrieval protections and performance monitoring should be portable across locations and providers.
Inference Exchange is therefore best understood as an infrastructure thesis awaiting execution. Equinix has announced the program; it has not yet delivered a generally available service, benchmarked end-to-end customer results or disclosed commercial terms. The Q1 2027 availability target, integration across three organizations and dependency on power, cooling and network capacity are material delivery variables. If the platform works as described, it could make deployment location a practical lever for latency, cost and governance. If it does not, enterprises may inherit another layer of coordination without resolving their core AI-operating-model challenge.