NVIDIA Hand-Delivers Its First Custom CPU to Anthropic, OpenAI, and Oracle โ The Vera Era Begins

News Summary
NVIDIA has officially begun delivering its first-ever custom CPU โ the Vera CPU โ to leading AI laboratories, marking a pivotal milestone in the company's push into the agentic AI era. In mid-May 2026, NVIDIA hand-delivered the inaugural units to Anthropic, OpenAI, SpaceXAI, and Oracle Cloud Infrastructure, signaling that large-scale production of the chip is fully underway.
What Is the NVIDIA Vera CPU?
The Vera CPU is NVIDIA's first wholly custom-designed central processing unit, built from the ground up for the demands of agentic AI workloads โ systems where AI models must reason, plan, retrieve information, and take actions autonomously over extended periods. Unlike training-focused accelerators, Vera is optimized for the sustained, latency-sensitive inference pipelines that agentic applications require.
Each Vera CPU die packs 88 CPU cores, code-named Olympus, which implement the ARM v9.2-A architecture. The design is simultaneous multithreading (SMT) capable, supporting 176 threads through NVIDIA's proprietary spatial multi-threading technology. Compared to its predecessor, the Grace CPU, Vera delivers twice the performance efficiency and operates approximately 50% faster than traditional rack-scale CPUs across agentic inference benchmarks.
The Historic First Deliveries
NVIDIA Vice President of Hyperscale and High-Performance Computing Ian Buck personally hand-delivered the first Vera CPU systems to three of the world's top AI research organizations โ Anthropic in San Francisco, OpenAI in Mission Bay (San Francisco), and SpaceXAI in Palo Alto โ on a Friday in mid-May 2026 (Pacific Time). Oracle Cloud Infrastructure (OCI) in Santa Clara received its delivery the following Monday morning (Pacific Time).
At Anthropic, James Bradbury, the company's Head of Compute, accepted the handoff in a conference room overlooking San Francisco Bay. Buck walked Bradbury through the server architecture built around the new processor. Bradbury said: "We're excited to see Vera emerge as a promising part of the ecosystem when solving for agentic workloads."
At OpenAI, Sachin Katti, Head of Compute Infrastructure, personally received the delivery and thanked Buck for bringing the system directly to their offices.
Architecture and Platform Context
The Vera CPU is a central component of NVIDIA's broader Vera Rubin platform, announced at CES 2026 in January and detailed further at GTC 2026 (March 2026, Pacific Time). The Vera Rubin platform represents a complete redesign of NVIDIA's AI computing stack, combining:
- The Vera CPU โ 88-core ARM-based processor optimized for orchestration and agentic reasoning
- The Rubin GPU architecture โ next-generation graphics processing units with HBM4 memory
- NVLink 6 interconnects โ ultra-high-bandwidth chip-to-chip and system-to-system communication fabric
- Integrated Groq inference accelerators โ tightly coupled into the system for low-latency token generation
Jensen Huang, NVIDIA's CEO, confirmed at GTC 2026 that the first complete Vera Rubin rack was already operational at Microsoft Azure, offering cloud customers early access to evaluate the platform's capabilities.
Production Scale and Commercial Availability
Vera is now in full production. NVIDIA and its contract manufacturing partners โ Foxconn, Quanta, and Wistron โ are scaling system and rack assembly rapidly, with large-scale shipments targeting the third quarter of 2026 (Q3 begins July 1, 2026 Eastern Time).
OCI has committed to deploying hundreds of thousands of Vera CPUs starting in 2026, citing the relentless growth in agentic AI workload demand across enterprise and cloud customers. Vera-based systems will be commercially available from major OEM partners โ including Cisco, Dell, HPE, Lenovo, and Supermicro โ in the second half of 2026.
Why Agentic AI Needs a Purpose-Built CPU
Traditional AI infrastructure is built around GPU-centric pipelines where training or batch inference dominates the compute profile. Agentic AI changes this equation. Agents run continuously, make rapid sequential decisions, manage large context windows, call external tools, and orchestrate sub-tasks โ all of which place heavy demands on system memory bandwidth, I/O throughput, and CPU-side orchestration logic.
NVIDIA designed Vera specifically to address these bottlenecks. The chip's high core count, expanded memory bandwidth, and deep integration with NVIDIA's NVLink fabric make it well-suited for the role of "command and control" in large-scale AI factories where thousands of GPU accelerators must be coordinated efficiently.
Industry Significance
This delivery marks the first time NVIDIA has shipped a CPU it designed entirely in-house, extending its hardware portfolio beyond GPUs and networking chips. Analysts view this as a strategic move to offer customers a fully integrated, NVIDIA-optimized compute stack โ reducing reliance on third-party CPUs and enabling tighter co-design between processing, memory, and interconnect layers.
The Vera CPU delivery also underscores the rapid maturation of agentic AI as a commercial workload category. Partnerships with Anthropic, OpenAI, SpaceXAI, and OCI at the very first production run demonstrate strong industry confidence that purpose-built AI infrastructure will define the next wave of large-scale AI deployment.