Skip to content

The Edge of the Cyber World See the latest

Apps

CoreWeave to Offer NVIDIA Vera, First AI Agent CPU [2026]

CoreWeave confirmed on September 30, 2026 that it will bring NVIDIA Vera to its cloud, a processor NVIDIA is calling the first CPU built specifically for AI agents. The announcement lands alongside CoreWeave’s rollout of the NVIDIA Vera Rubin NVL72 rack-scale system, and it marks a shift in how cloud providers think about the “boring” half of AI infrastructure: the CPU cores that manage sandboxes, tool calls, and orchestration around the GPUs that get all the attention.

For years, CPUs in AI data centers have been treated as plumbing, cheap, interchangeable cores that shuttle data to GPUs and otherwise stay out of the way. Vera flips that assumption. NVIDIA designed it around a workload that barely existed three years ago: autonomous AI agents that plan, call tools, spin up sandboxed environments, and loop through reasoning steps thousands of times a minute. CoreWeave, the Nvidia-backed cloud provider that has built its business on GPU rentals, is now betting that agent-heavy customers will pay for CPU capacity engineered the same way its GPU capacity is: for a specific job, not a general one.

What NVIDIA Vera actually is

NVIDIA is positioning Vera as a companion chip to its GPUs rather than a general-purpose server CPU competing head-on with Intel Xeon or AMD EPYC on raw core-for-core throughput. The company describes it as purpose-built for agentic workloads, meaning tasks that involve repeated reasoning, tool use, orchestration, and constant interaction with external systems, according to NVIDIA’s official blog post announcing the CoreWeave partnership. That framing matters. A CPU tuned for agent workloads isn’t optimized for the same thing a database server or a web backend needs. It’s optimized for spinning up isolated execution environments fast, handling bursty concurrent requests, and staying close to the GPU memory fabric so agent loops don’t stall waiting on data.

CoreWeave plans to run Vera bare metal at rack scale, which means customers get direct access to the CPU infrastructure instead of a heavily virtualized, shared environment. That’s a deliberate choice. Bare-metal access strips out hypervisor overhead, which matters when the workload is thousands of short-lived agent sandboxes spinning up and tearing down in rapid succession rather than a handful of long-running virtual machines.

Chen Goldberg, Executive Vice President of Product and Engineering at CoreWeave, framed the shift bluntly. “General-purpose infrastructure can hinder agentic AI; Vera is the inaugural CPU specifically crafted to enhance its performance,” Goldberg said, according to a CoreWeave announcement carried by Seeking Alpha. Goldberg added that the company’s platform already supports Vera through existing tooling: “Our platform inherently supports Vera with tools like CoreWeave Sandboxes right from the start,” she said in the same announcement.

Inside the Vera Rubin NVL72 rack

Vera doesn’t ship alone. It’s the CPU half of NVIDIA’s Vera Rubin NVL72 platform, a rack-scale system that pairs 72 NVIDIA Rubin GPUs with 36 Vera CPUs in a single integrated unit, according to reporting cited by CoreWeave and NVIDIA. NVIDIA says a full rack packs 128 Vera CPUs and 11,264 CPU cores, a configuration the company claims can support more than 11,000 concurrent execution environments when allocating roughly one CPU core per environment. That’s the core pitch: agent workloads don’t need a handful of beefy cores, they need a very large number of cores that can each spin up, execute, and shut down an isolated environment on demand.

The rack ties everything together with a 20.7 TB unified HBM4 memory domain, letting the GPUs and CPUs work against a shared memory pool rather than shuttling data across slower interconnects. NVIDIA connects the components using sixth-generation NVLink, rated at 216 TB/s, alongside ConnectX-9 SuperNICs and BlueField-4 DPUs for networking and data processing. CoreWeave has previously deployed BlueField-4 as part of its AI agent safety stack, a system we covered in detail here when NVIDIA first bet on the DPU as a security layer.

CoreWeave says it completed what it calls an industry-first bring-up and validation of Vera Rubin NVL72 on CoreWeave Cloud, positioning itself as the earliest production cloud deployment of the platform among major providers. NVIDIA’s own performance claims for Vera Rubin are aggressive: the company says the platform delivers up to 10 times higher inference performance per watt than the previous-generation Blackwell NVL72 system, and that some large mixture-of-experts training workloads could use as little as one-quarter the GPU count compared with Blackwell. NVIDIA also claims the platform can cut cost per million tokens to one-tenth of Blackwell’s. All three figures are vendor claims tied to specific workloads and benchmark conditions rather than independently verified, workload-agnostic performance numbers, and no public rental pricing for Vera or Vera Rubin NVL72 has been released.

Why CPUs suddenly matter again in the AI stack

The GPU has been the star of every AI infrastructure story since the generative AI boom took off. CPUs got treated as an afterthought, a cost center to minimize rather than an area to invest R&D dollars into. Vera is NVIDIA’s argument that this assumption breaks down once workloads shift from single large model calls to agentic loops. An AI agent doesn’t just generate one response and return it. It plans a task, calls external tools or APIs, executes code in a sandbox, evaluates the result, and loops back through the reasoning process, often dozens or hundreds of times per task. Each of those steps needs somewhere to run that isn’t the GPU, and doing it inefficiently means the expensive GPU sits idle waiting on CPU-bound orchestration work.

NVIDIA’s own testing found more than 3x faster agent sandbox startup times when running on Vera CPUs compared with the systems CoreWeave used previously, according to NVIDIA’s blog post on the partnership. That’s a meaningful number for anyone running agent fleets at scale: sandbox startup latency directly determines how many concurrent agent tasks a cluster can serve, and how fast an autonomous system can react to a new request. A CoreWeave co-founder, identified only as Nick in a published interview, put the strategic logic plainly: “You’re absolutely going to see a bunch of Vera CPUs crammed next to a bunch of Vera Rubin servers,” he said, according to a published interview with CoreWeave executives.

How Vera compares to the rest of the AI CPU field

Vera isn’t launching into an empty field. Arm has been pushing its own AI-optimized silicon through the Neoverse CSS N4 platform, which we detailed in our look at Arm’s AGI CPU and its challenge to AMD and NVIDIA. AMD, meanwhile, has been posting strong data center numbers of its own, with data center revenue reported to have jumped 107% year over year, a trend we covered in our AMD data center revenue report. Intel’s Xeon lineup remains the incumbent default in most enterprise data centers, though it hasn’t been positioned specifically around agentic AI workloads the way Vera has.

The distinction that separates Vera from Arm’s Neoverse-based chips and AMD’s EPYC lineup is scope. Arm and AMD are selling general-purpose data center CPUs that happen to work well for AI-adjacent tasks. NVIDIA is selling a CPU that exists because a GPU platform needs it, bundled into a rack that only makes sense as a unit. That’s consistent with NVIDIA’s broader strategy of building vertically integrated systems rather than standalone components, the same logic behind its two-layer AI agent safety silicon push announced earlier this year.

Platform Vendor Primary focus Notable claim
NVIDIA Vera NVIDIA Agentic AI sandboxes, orchestration 3x faster agent sandbox startup vs. prior systems
Vera Rubin NVL72 NVIDIA Rack-scale AI training/inference Up to 10x inference-per-watt vs. Blackwell NVL72
Neoverse CSS N4 (Arm-based) Arm licensees General AI data center compute Positioned against AMD/NVIDIA data center share
EPYC (data center) AMD General-purpose server compute Data center revenue up 107% year over year
Xeon Intel General-purpose enterprise compute Incumbent default, not agent-specific

CoreWeave’s bigger bet on agentic infrastructure

CoreWeave has built its entire business model around being first to deploy NVIDIA’s newest silicon, and the Vera announcement follows that pattern closely. The company says it completed an industry-first bring-up and validation of the Vera Rubin NVL72 platform on its own cloud, a claim that, if accurate, gives CoreWeave a head start over hyperscalers like AWS, Google Cloud, and Microsoft Azure in offering the hardware commercially. That speed advantage has been central to CoreWeave’s pitch to investors and customers alike since its IPO, and it’s part of why the company built out data center capacity aggressively even as questions about power availability persist across the industry, an issue we examined in our coverage of Blackstone’s $95 million purchase of a power-rich Nvidia-linked building.

The Vera announcement also plugs directly into CoreWeave’s existing Sandboxes product, a tool built for running isolated code execution environments for AI agents. According to Goldberg’s comments, that integration was designed in from the start rather than bolted on after the fact, which suggests CoreWeave anticipated agent-native infrastructure demand well before Vera CPUs were available to deploy. That kind of foresight matters commercially: cloud providers that can offer agent-specific infrastructure without forcing customers to re-architect their stacks have an edge in a market where enterprises are still figuring out how to operationalize AI agents at scale.

The supply chain and memory angle

Rack-scale systems like Vera Rubin NVL72 depend heavily on memory supply, and that supply chain has been under visible strain. Micron discontinued its 2GB GDDR7 modules earlier this year, a move that squeezed supply for RTX 50-series GPUs, according to our earlier reporting on the Micron GDDR7 discontinuation. While Vera Rubin uses HBM4 rather than GDDR7, the broader memory market remains tight, and NVIDIA’s decision to build a 20.7 TB unified HBM4 domain into every NVL72 rack puts additional pressure on an already constrained supply of high-bandwidth memory. That constraint could shape how quickly CoreWeave, or any other cloud provider, can scale Vera deployments beyond initial early-access customers.

NVIDIA’s market position gives it leverage most competitors can’t match when negotiating memory supply agreements. The company’s valuation has outpaced even Apple’s, a gap we broke down in our look at the market cap difference between Apple and Nvidia. That scale advantage matters when a company needs to lock in HBM4 capacity months or years ahead of a product launch, something smaller cloud and chip competitors struggle to do at the same volume or price.

What this means for developers and enterprises

For engineering teams building AI agent products, the practical question isn’t whether Vera is interesting, it’s whether it changes cost or latency enough to justify migrating workloads. Right now, that’s hard to answer with certainty. NVIDIA has not published independent, workload-agnostic CPU benchmarks for Vera, only platform-level claims tied to specific test conditions like the 3x sandbox startup improvement and the 10x inference-per-watt figure for the full Rubin NVL72 rack. No public hourly pricing, reserved-instance rate, or per-token cost has been released for CoreWeave’s Vera offering, which means enterprise buyers evaluating a switch don’t yet have the numbers needed to run a real cost comparison against existing infrastructure.

What is clear is the direction of travel. Agent workloads are pushing infrastructure vendors to rethink assumptions that held for a decade of cloud computing, chiefly that CPUs are interchangeable commodity resources. If Vera’s sandbox startup gains hold up under independent testing, the CPU choice behind an AI agent platform could become as consequential a decision as the GPU choice already is, particularly for companies running large fleets of concurrent agents where startup latency compounds across thousands of parallel tasks.

Historical context: from GPU wars to full-stack silicon

NVIDIA’s push into CPU silicon isn’t new. The company has been building Grace and Grace Hopper CPU platforms for several years as part of its broader ambition to control every layer of the AI compute stack, not just the GPU. Vera represents a narrowing of that ambition into a more specific niche: rather than another general-purpose data center CPU, it’s a chip whose entire design brief is agentic AI. That’s a notable evolution from NVIDIA’s earlier CPU efforts, which competed more directly with Arm-based server chips on general compute metrics. The shift mirrors what happened with GPUs a decade ago, when NVIDIA moved from selling general graphics cards to selling purpose-built AI accelerators as machine learning workloads diverged from traditional graphics rendering.

CoreWeave’s own history reinforces the pattern. The company started as a cryptocurrency mining operation before pivoting entirely to GPU cloud services during the generative AI boom, and it has since built its identity around being the fastest mover on new NVIDIA hardware. Its willingness to commit to unproven agentic-CPU infrastructure before independent benchmarks exist fits that established playbook: move first, let scale and NVIDIA’s roadmap validate the bet later.

Market impact and investor reaction

The Vera announcement arrives at a moment when investors are scrutinizing AI infrastructure spending more closely than at any point in the current buildout cycle. Every major cloud provider has poured capital into GPU capacity over the past two years, and the emergence of agent-specific CPU demand gives NVIDIA and CoreWeave a new growth narrative beyond simply selling more GPUs. For NVIDIA, Vera extends its total addressable market into a category, CPU infrastructure, that it has never dominated the way it dominates GPUs. For CoreWeave, being the first named cloud partner for Vera reinforces the company’s positioning as NVIDIA’s preferred early-access partner, a relationship that has underpinned much of CoreWeave’s growth since its public listing.

Competitors will be watching closely. AWS, Microsoft Azure, and Google Cloud have their own custom silicon programs, including Amazon’s Graviton and Trainium chips and Google’s Axion CPUs, but none has publicly marketed a CPU explicitly built around AI agent sandboxing the way NVIDIA has framed Vera. If enterprise demand for agent infrastructure grows the way NVIDIA is betting it will, expect rival clouds to either license similar NVIDIA hardware or accelerate their own agent-specific CPU roadmaps within the next several quarters.

Predictions: where Vera and agentic CPUs go from here

  • Pricing transparency arrives within two quarters. CoreWeave will likely publish concrete Vera rental rates once early-access customers move to general availability, giving the market its first real cost-per-agent benchmark.
  • Rival clouds respond with their own agent-tuned CPUs. Expect AWS, Google Cloud, and Microsoft Azure to at minimum reposition existing custom silicon, Graviton, Axion, and Cobalt, around agentic workloads even without a hardware redesign.
  • Independent benchmarks will complicate NVIDIA’s claims. The 10x inference-per-watt and 3x sandbox startup figures are vendor-reported; third-party labs and academic groups will likely publish more conservative numbers once Vera Rubin systems are widely accessible.
  • Memory supply becomes the real bottleneck. With HBM4 demand already elevated and GDDR7 supply tightening, as seen in the Micron GDDR7 exit, Vera Rubin’s rollout pace will likely be gated more by memory availability than by CPU manufacturing capacity.
  • Agent-specific silicon becomes a standard product category. Within 12 to 18 months, expect “CPU for AI agents” to become a recognized market segment with competing offerings from Arm licensees and AMD, not just an NVIDIA marketing term.

The competitive stakes for CoreWeave

CoreWeave’s willingness to be the launch partner for unproven agentic-CPU hardware carries real risk alongside the upside. If Vera underdelivers relative to NVIDIA’s benchmark claims once independent testing catches up, CoreWeave’s reputation as the fastest and most reliable early-access NVIDIA partner takes the hit, not just NVIDIA’s. That’s a meaningful exposure for a company whose valuation has been built substantially on the premise that being first to market with NVIDIA’s newest hardware translates directly into customer wins and revenue growth.

At the same time, the downside of not moving first is arguably larger in a market this competitive. Enterprises building agent products today are actively choosing infrastructure partners, and a cloud provider that can offer purpose-built agent CPUs alongside GPUs, with sandboxing tools already integrated, has a differentiated pitch that’s hard for a general-purpose cloud provider to match quickly. CoreWeave’s bet is that agent-native infrastructure demand will grow fast enough to reward early movers, even before the pricing and independent-benchmark questions get fully answered.

Vera Rubin NVL72 spec Reported figure
GPUs per rack 72 NVIDIA Rubin GPUs
CPUs per rack 36 (rack pairing) / 128 Vera CPUs cited by NVIDIA for full configuration
CPU cores per rack 11,264
Concurrent environments supported 11,000+ (approx. 1 core per environment)
Unified memory domain 20.7 TB HBM4
Interconnect NVLink 6, rated at 216 TB/s
Networking/DPU ConnectX-9 SuperNICs, BlueField-4 DPUs
Claimed inference-per-watt gain Up to 10x vs. Blackwell NVL72
Claimed cost per million tokens Up to 10x lower vs. Blackwell
Public rental pricing Not yet released

What’s still unverified

It’s worth being precise about what’s confirmed versus what’s still a vendor claim. CoreWeave and NVIDIA have confirmed the partnership, the Vera CPU’s existence, its positioning around agentic workloads, and the broad architecture of the Vera Rubin NVL72 rack. What hasn’t been independently verified includes the detailed core microarchitecture of Vera itself, standalone CPU benchmark results outside NVIDIA’s own testing, actual customer pricing, and real-world performance under production agent workloads at scale. The 3x sandbox startup figure and the 10x inference-per-watt claim both come from NVIDIA and CoreWeave’s own announcements rather than third-party benchmarking labs, and buyers evaluating the platform should treat those numbers as directional rather than guaranteed.

Frequently asked questions

What is NVIDIA Vera?
Vera is a CPU that NVIDIA describes as the first processor built specifically for AI agent workloads, designed to handle sandboxing, orchestration, and tool-use tasks that support GPU-based AI models rather than compete with GPUs directly.

How is Vera different from a regular server CPU like Intel Xeon or AMD EPYC?
Vera is optimized around spinning up large numbers of isolated, short-lived execution environments quickly, which NVIDIA says produces more than 3x faster agent sandbox startup times compared with prior systems, rather than optimizing for general-purpose enterprise workloads.

When will CoreWeave offer NVIDIA Vera?
CoreWeave announced the addition of Vera to its cloud portfolio on September 30, 2026, alongside its deployment of the Vera Rubin NVL72 platform, though a specific general-availability date and public pricing have not been released.

What is the Vera Rubin NVL72 platform?
It’s a rack-scale AI computing system that combines NVIDIA Rubin GPUs with Vera CPUs, a unified 20.7 TB HBM4 memory domain, and NVLink 6 interconnects, built for agentic AI, reasoning models, large-scale inference, and mixture-of-experts training.

How much does NVIDIA Vera cost to rent on CoreWeave?
No public rental price has been disclosed. CoreWeave and NVIDIA’s announcements describe hardware capabilities and efficiency claims but do not include an hourly rate, reserved-instance pricing, or minimum commitment terms.

Is NVIDIA Vera the same as NVIDIA Grace?
Vera and Grace are both NVIDIA CPU platforms, but Vera is positioned specifically around agentic AI workloads as part of the Vera Rubin platform, representing a narrower, more specialized design brief than NVIDIA’s earlier Grace CPU efforts.

Are NVIDIA’s performance claims for Vera independently verified?
No. The 3x sandbox startup improvement and the 10x inference-per-watt figures come from NVIDIA and CoreWeave’s own announcements and testing, not from independent third-party benchmarking labs.

Which companies compete with NVIDIA in AI-focused CPU silicon?
Arm licensees offer AI-optimized platforms like Neoverse CSS N4, AMD sells EPYC data center CPUs that have posted strong recent revenue growth, and Intel’s Xeon lineup remains a widely deployed enterprise default, though none is marketed specifically around AI agent sandboxing the way Vera is.

Related Coverage

Source: Tech Insider