Intel used Hot Chips 2026 to lay out the internals of Crescent Island, a data center GPU built specifically for AI inference. It pairs up to 480GB of LPDDR5X with a 350W air-cooled PCIe form factor. By not assuming a liquid-cooled rack, the design goes straight after the running cost of agentic AI.

Betting on memory capacity to fix inference economics

Crescent Island is built around 32 Xe cores based on Xe3P, along with 256 XMX engines that handle matrix math. Attached to that is up to 480GB of LPDDR5X. The whole card fits inside a 350W budget and connects over PCIe.

What stands out is that Intel leads with memory capacity rather than raw throughput. In production inference, memory is consumed not only by the model weights but also by the KV cache that holds conversation history. The longer the context window and the more AI agents running at once, the more capacity the workload demands. When a card runs short, the only option is to spread the job across more GPUs, which means more boards, more racks and more power.

Intel's pitch is that fitting larger models, longer contexts and more concurrent agents onto a single card brings down the cost per token. In the company's own framing, the goal is to reframe the economics of real-time inference.

Why the 350W air-cooled line matters

Cooling is the other axis. Crescent Island is meant to drop into air-cooled data centers as they exist today. Recent AI accelerators routinely draw 700W to over 1,000W per card, and deploying them brings liquid-cooling plumbing and upgraded power delivery along with it. That upfront investment has been a real barrier for companies that simply want to try AI at scale.

A 350W air-cooled card can be added without tearing up the server room. Choosing LPDDR5X over HBM follows the same logic: lower cost per gigabyte and lower power. Bandwidth does not match HBM, but for inference workloads where capacity is the binding constraint, the trade-off is reasonable. With memory supply tight worldwide, avoiding HBM is itself a supply-side advantage.

Diamond Rapids and Wildcat Lake shared the stage

The Hot Chips session covered three architectures rather than Crescent Island alone.

On the server CPU side is Diamond Rapids, the next-generation Xeon. It offers up to 256 cores with 1.28GB of last-level cache, 16 memory channels running at 12800 MT/s, 128 lanes of PCIe Gen6 and CXL 3.0 support. It is built on Intel 18A-P, a power- and performance-enhanced process, and uses Foveros Direct 3D packaging with UCIe-S interconnect. Advanced Performance Extensions (APX) and an enhanced Advanced Matrix Extensions (AMX) round out the design. Intel positions it as the orchestration layer that coordinates fleets of agents.

Wildcat Lake covers client and edge, shipping as Intel Core Series 3 processors. Built on Intel 18A, it combines 2 performance cores and 4 efficiency cores with Xe3 integrated graphics that include XMX acceleration, plus an NPU rated at up to 17 TOPS. It supports LPDDR5X-7467 and carries Wi-Fi 7 and Bluetooth 6.0. It is also the first Intel processor to use UCIe, which makes cost-effective multi-chip packages practical for mainstream systems.

Putting its own manufacturing front and center

All three share the same foundation: the Intel 18A process family, Foveros Direct 3D packaging and UCIe, the industry standard for chiplet interconnects. Intel CTO Pushkar Ranade said agentic AI is changing design from the transistor and package all the way up to full system architecture, and that the task ahead is tightly integrating general-purpose compute with purpose-built acceleration so systems can scale within real-world power, cost and deployment constraints.

For a company that also runs a foundry business, proving out these technologies in its own products doubles as a sales argument to external customers. Bundling rack-to-edge silicon into a single announcement reads as a deliberate demonstration of that manufacturing stack.

Summary

Crescent Island is not chasing peak compute. It is a GPU tuned to keep inference running within 480GB of capacity and a 350W air-cooled envelope. Because it does not require liquid cooling, it lowers the barrier to deployment and becomes a practical option for organizations that want more concurrency out of existing facilities. Together with Diamond Rapids and Wildcat Lake, it shows which layers Intel intends to serve as agentic AI spreads.