NVIDIA has published a detailed look at Vera, a CPU it positions as a new category built for the agentic AI era[1]. Described as a "max single-threaded CPU at scale," Vera is designed so that all 88 cores run at full performance simultaneously, delivering 1.8x the sustained per-core performance of x86 in workloads that emulate agent execution. Perplexity, the AI search company, is also looking to deploy Vera in its upcoming production system.
Why Single-Threaded Performance Matters in the Agent Era
An AI agent runs in a loop: the model reasons about the next step, the CPU executes the actual work — tool calls, code execution, data processing — and the result feeds back into the model's next decision. Each step depends on the result of the one before it, so adding more cores cannot make any single loop advance faster. What determines loop speed is per-core performance[1].
NVIDIA argues that today's data center CPUs are not optimized for this workload. Since the rise of the cloud, CPU makers have evolved toward lowering the cost per rentable core. Increasing core counts came at the expense of silicon area for what makes cores fast, such as high-performance memory fabrics, and the shift to chiplet architectures created a "chiplet tax" in which each core can no longer access the full memory performance of the chip[1].
From the AI factory perspective, a slow CPU translates directly into idle time for expensive GPUs. The faster agent work completes, the higher GPU utilization and revenue become — which is why NVIDIA contends AI factories need a CPU where every core delivers uncompromised single-threaded performance[1].
The Olympus Core and 1.2TB/s of Memory Bandwidth
At the heart of Vera is Olympus, NVIDIA's custom-designed CPU core. It delivers 50 percent higher instructions per cycle (IPC) than the previous-generation NVIDIA Grace. The memory subsystem provides up to 1.2TB/s of LPDDR5X bandwidth at less than 40 watts of memory power, and a monolithic compute die paired with 3.4TB/s of core-to-core bandwidth keeps all 88 cores supplied with data even under full load. NVIDIA says that core-to-core bandwidth is 3x greater than any other data center CPU[1].
As a result, Vera delivered 1.8x the sustained per-core performance of x86 in loaded CPU workloads representing agentic execution. Those gains compound across tool calls, code executions, data-processing steps and verification passes, raising the overall throughput of an AI factory[1].
1.5x in Perplexity's Tests, Up to 6x in Data Workloads
The adoption examples are concrete. Perplexity tested Vera on the agentic work it runs every day: a real coding workflow — cloning a repository and running its test suite in sandboxes — completed about 1.5x faster than x86, and concurrent sandboxes started up to 1.9x faster. The company is now looking to deploy Vera in its upcoming production system[1].
On the data side, partners measured 3x faster large-scale SQL analytics with Starburst and up to 6x lower latency in real-time streaming with Redpanda, both against leading x86 server CPUs. Vera is also the same CPU that hosts the GPUs in NVIDIA Vera Rubin and powers the NVIDIA BlueField-4 STX storage processor, allowing an entire AI factory to run on one architecture and one toolchain[1].
Next-Generation "Rosa" CPU and "Rigel" Core Previewed
NVIDIA also confirmed its CPU roadmap will continue. The next-generation Rosa CPU will feature the new Rigel core, an Arm v9.2 design that keeps the same silicon footprint as Olympus while pushing per-core performance further through better instruction delivery, a larger L2 cache and more efficient memory handling[1].
Summary
The design philosophy NVIDIA laid out for Vera marks a shift away from the core-count race toward making every core run at maximum speed. With the high-IPC Olympus core, 1.2TB/s of memory bandwidth and 3.4TB/s of core-to-core bandwidth, Vera achieves 1.8x the per-core performance of x86 in agent workloads, backed by real measurements such as 1.5x in Perplexity's coding workflow and 3x in SQL analytics. In data centers where agents take center stage, the very yardstick for evaluating CPUs may be about to change.
Source: https://blogs.nvidia.com/blog/nvidia-vera-max-single-threaded-cpu-at-scale/
