Meta has confirmed it has started testing physical units of its own AI inference chip, code-named Arke (internally MTIA 450), with data center deployment targeted for the first half of 2027. A next-generation chip that is likewise aimed at inference, code-named Astrid (MTIA 500), is said to be close to finishing its design as well. The dual push is aimed at cutting Meta's reliance on NVIDIA GPUs while keeping the ballooning costs of AI infrastructure under control.

Two Chips, Both Built Around Inference

Meta's custom silicon effort builds on the MTIA chip family, which already includes MTIA 300 and earlier generations. Arke (MTIA 450) is the third generation in the inference-focused line, purpose-built to run the trained models behind Facebook, Instagram, and WhatsApp as efficiently as possible per watt. It is not designed to compete in workloads that demand split-second responsiveness; instead, its whole design center is driving down the cost of running huge volumes of inference.

Astrid (MTIA 500) is the fourth generation. Its design is reportedly about a month from completion, with broader data center rollout expected by the end of 2027. The MTIA parts are said to be capable of handling a wide range of workloads, training included, but Meta has signaled that for now it intends to use both Arke and Astrid primarily for generative AI inference.

Twelve Prototypes Delivered September 1, Within 2-3% of Simulations

Twelve prototype Arke chips were delivered on September 1, 2026, and early testing reportedly came within just 2-3% of Meta's pre-launch simulations. The chips were tested not only with Meta's own models but also with models from DeepSeek and Alibaba, and reportedly ran without issue from day one. Validating the chip against third-party models, not just Meta's own, suggests an ambition to build a genuinely general-purpose inference platform.

Manufacturing is handled by TSMC, with Broadcom collaborating on chip design. The Broadcom partnership is expected to continue through 2029, positioning this as a multi-generation effort rather than a one-off project.

A Combined Training-and-Inference Chip, "Olympus," Scrapped

Alongside the Arke and Astrid news, Meta has reportedly canceled development of "Olympus," a chip meant to handle both training and inference in one design. According to Meta's own estimates, folding both functions into a single chip would raise manufacturing costs by roughly 30%, a gap the company apparently judged untenable once chip deployment reaches gigawatt scale.

Yee Jiun Song, the Vice President of Engineering overseeing Meta's custom silicon program, was quoted as saying that once a company starts building up gigawatts and gigawatts of capacity, cost becomes something it cares about a great deal. Narrowing Arke and Astrid to inference and optimizing both around that single job appears to reflect that thinking directly.

Over a Gigawatt in 12 Months, and Less Reliance on NVIDIA

Meta has committed to deploying more than one gigawatt of computing capacity built on its custom chips within any given 12-month period, with further acceleration expected after that. The goal is to keep using costly NVIDIA GPUs for training while shifting inference workloads onto its own silicon, bringing down the overall cost of its AI infrastructure. As demand for generative AI services keeps pushing up data center power needs and capital spending, optimizing costs through custom chip design is quickly becoming a shared challenge across major AI companies, not just Meta.

Summary

Meta has begun testing physical units of its new inference-focused AI chip, Arke (MTIA 450), with deployment targeted for the first half of 2027. Its companion chip, Astrid (MTIA 500), is nearing design completion and is expected to roll out by the end of 2027, also with inference as its main job. While prototype units matched simulations to within 2-3%, Meta scrapped its combined training-and-inference chip, Olympus, over cost concerns. With a goal of deploying more than a gigawatt of custom chip capacity every 12 months, Meta's approach is drawing attention as an attempt to both reduce its dependence on NVIDIA and rein in the cost of AI infrastructure.

References