Lining up faster GPUs does not make an AI cluster as fast as the spec sheet suggests. Much of the gap sits in the network between them rather than in the silicon itself. On September 14, Cornelis introduced Active Compute Fabric, an architecture that gives the wiring real processing capability so that GPUs spend less time waiting. The company also disclosed 205 million USD in funding and a collaboration with Qualcomm Technologies.
Roughly Half of GPU Time Is Spent Waiting
Cornelis frames the problem as accelerator utilization. As models and clusters have grown, inter-node communication and synchronization have come to dominate the clock rather than computation. Memory and storage became workload-aware over the past decade, while networking largely stayed in its original role of moving data from one endpoint to another.
The company put numbers on the loss using published industry data. In a system with 100,000 GPUs, about half of all GPU hours go to waiting for data, which works out to roughly 1.68 billion USD a year in wasted capacity and 500 GWh of electricity, about as much as 48,000 US households consume annually. The model assumes 4 USD per GPU hour, 8,400 operating hours a year, and roughly 50 percent unproductive time.
Exchange rate: 1 USD = 155 JPY
From Moving Data to Acting on It
Active Compute Fabric tries to absorb that idle time inside the network. It rests on three pillars: lossless transport that does not drop packets, in-fabric acceleration, and programmable compute. Because operations can run on data while it is in flight, collective operations that would otherwise stall a GPU can be pushed into the fabric instead.
Cornelis puts weight on the programmable part. AI algorithms and software change on short cycles, so a fabric with its functions burned into hardware risks becoming the bottleneck within a few years. A design that accepts new functions later lets customers get more life out of accelerators they have already bought.
The architecture sits on industry standards. Scale-up connections inside the rack use UALink and ESUN, while scale-out between racks follows Ultra Ethernet specifications, keeping customers free to choose accelerators. CEO Lisa Spelman said customers want complete rack-scale solutions without being locked into a single vendor or architecture.
Qualcomm Technologies appeared alongside Cornelis in the announcement. Tony Pialis, EVP and GM of the Data Center business, noted that as AI systems keep scaling, moving data efficiently across the rack becomes as important as the compute itself. Low utilization shows up directly in the economics of inference, so treating compute, memory, and networking as separate problems no longer works. Both executives took the keynote stage together at AI Infra Summit.
The idea that the network is a first-order design decision rather than an afterthought echoes arguments other vendors made at the same event. As the industry shifts toward rack-scale architectures, how much choice a platform leaves open is becoming a point of differentiation.
Funding Goes to the Next Generation
The 205 million USD round supports the move into scale-up networking and the launch of the next product generation. Joel Whitley, Partner at IAG Capital Partners, put open-standard scale-up and scale-out networking for AI at more than 55 billion USD of opportunity by 2030, and argued that the network will determine how much of the overall AI build-out delivers real return.
On the product side, the 400 Gb/s CN5000 is shipping today. The SuperNIC offers one or two 400 Gb/s ports on a PCIe Gen5 x16 card, the switch aggregates 48 ports at 400 Gb/s for 38.4 Tb/s full duplex, and the director class chassis reaches 576 ports and 230.4 Tb/s. The 800 Gb/s CN6000 is sampling with customers, with broader availability expected in the fourth quarter of 2026. CN6000 is an 800G SuperNIC supporting Omni-Path, RoCEv2, and Ultra Ethernet over PCIe 6.0.
Cornelis traces its origins to the spinout of Intel's high-performance interconnect business and continues the Omni-Path lineage. Its OPX software stack is open source and built on libfabric, so applications written for MPI, OpenSHMEM, and RCCL run without code changes. In Lenovo's validation on ThinkSystem SC750 V4 with Intel Xeon 6, node-to-node throughput measured 378 to 391 Gb/s, averaging 385 Gb/s, at roughly 1.1 microseconds of latency.
Summary
Active Compute Fabric reframes the network as part of the compute system rather than passive wiring. If the estimate of 1.68 billion USD a year lost to waiting at 100,000-GPU scale holds, cutting idle time is a better investment than adding raw bandwidth. With CN5000 shipping and CN6000 arriving in volume late in 2026, and with UALink and Ultra Ethernet underneath, the open question is whether this becomes a practical alternative in a market NVIDIA currently dominates.
