Jensen Huang wasn't even in the room, and Nvidia still managed to make the week about Nvidia. While the CEO was in Japan shaking hands with robotics firms, his lieutenants were running a technical workshop in Santa Clara that had one quiet, unmistakable message: we don't just want to power your AI—we want to own every chip inside the box.

The centerpiece is Vera Rubin, Nvidia's successor to its Grace Blackwell superchip system. On the surface, it's a GPU story. Underneath, it's a CPU land grab.

What Vera Rubin Actually Is (And Why the CPU Ratio Matters)

Think of Vera Rubin as a purpose-built AI rack, not just a chip. The flagship NVL72 configuration pairs 72 Rubin GPUs with 36 Vera CPUs—one CPU for every two GPUs—all packed into a single liquid-cooled platform that Nvidia is pitching as "plug-and-play."

That CPU-to-GPU ratio isn't arbitrary. As AI workloads shift toward agentic systems—models that don't just respond but plan, orchestrate, and call external tools—the CPU becomes load-bearing infrastructure. Someone has to manage data routing, networking stacks, and the orchestration logic that keeps agents from running into each other. That someone used to be someone else's chip.

Nvidia wants it to be theirs.

The Agentic AI Inflection Point Is Nvidia's Opening

For years, the GPU was the star and the CPU was the boring co-star nobody wrote think-pieces about. That's changing fast. Agentic AI pipelines are CPU-hungry in ways that pure inference workloads aren't—they involve more branching logic, more I/O, more orchestration overhead.

  • Training and inference: Still GPU-dominated, still Nvidia's core business.
  • Agentic orchestration: CPU-dependent, increasingly strategic.
  • Networking and data flow: Where CPUs and NICs fight for relevance in the rack.

Ian Buck, Nvidia's VP of accelerated computing and the architect of the CUDA software ecosystem, framed it bluntly at the workshop: "We're on a road map to crank out new architectures, not just GPUs but CPUs. We're going to keep innovating, because it's do this or die in Silicon Valley." That's either a genuine strategic pivot or the most dramatic way anyone has ever described a product roadmap. Possibly both.

OpenAI Already Has One. The China Angle Is Wilder.

Nvidia confirmed that OpenAI is already running a Vera Rubin rack—which is less of a validation and more of a reminder that OpenAI is essentially Nvidia's most important customer and most important marketing asset simultaneously.

More interesting: Nvidia has reportedly been pitching the standalone Vera CPU to Chinese customers, with potential availability as early as August. That's notable given the labyrinthine export control environment around advanced AI chips. CPUs don't carry the same restriction profile as high-end GPUs, which may be precisely the point—a foot in the door where the front entrance is blocked.

The "Plug-and-Play" Claim Deserves Scrutiny

Nvidia executives made a point of calling the NVL72 rack more plug-and-play than earlier systems. That's doing a lot of work for four words. "Plug-and-play" in data center terms usually means "easier than before, still requires a team of specialists, significant integration work, and a cooling infrastructure overhaul."

Liquid cooling at rack scale is not a weekend project. The operational complexity of running NVL72-class hardware—power density, thermal management, network fabric—doesn't disappear because the chassis ships pre-assembled. The real question is whether Nvidia's software stack (CUDA, NIM microservices, the whole ecosystem) creates enough lock-in that buyers stop asking about the integration burden. Historically, the answer has been yes.

AMD's Conference Was Right Around the Corner. Funny Timing.

Nvidia timed these disclosures to land just before AMD's annual product event in San Francisco. That's not a coincidence—that's a media cycle judo move. Drop your benchmarks first, own the conversation, make AMD's announcements feel reactive. It's a tactic as old as competitive tech marketing, and Nvidia has gotten extremely good at it.

AMD is a real competitor, particularly on price-performance for certain inference workloads. But AMD doesn't have a CPU story that integrates this tightly with a GPU roadmap. Intel has Gaudi but limited mindshare in the hyperscaler AI stack. Nvidia's actual moat isn't the hardware—it's CUDA, and everything Vera Rubin does deepens that dependency.

Hot Take

Nvidia selling CPUs isn't about CPUs. It's about making the data center a Nvidia-native environment where swapping out any component—GPU, CPU, networking—feels like pulling a thread that unravels the whole sweater. The Vera Rubin system is less a product launch and more an architectural declaration of intent.

Prediction: Within three years, at least two major hyperscalers will publicly announce partial Nvidia CPU deployments, framed as "optimization" decisions—while privately wrestling with the vendor lock-in they've just deepened. The ones who don't will be the ones who've invested hardest in custom silicon (think Google TPUs, AWS Trainium). Everyone else is walking into Nvidia's ecosystem with their eyes open and their procurement teams slightly nervous.

What This Means If You're Actually Building With This Stuff

  • If you're a hyperscaler: The NVL72 rack density is genuinely compelling, but model your 3-year TCO with lock-in risk baked in.
  • If you're a mid-size AI company: The standalone Vera CPU availability is worth watching—it may offer a cheaper on-ramp into Nvidia's ecosystem before full rack commitments.
  • If you're building agentic systems: CPU orchestration overhead is a real bottleneck. Understanding how Vera's CPU handles multi-agent I/O will matter more than peak GPU FLOPS.
  • If you're AMD or Intel: You have a narrowing window to make the integration story compelling before the Nvidia flywheel gets another full rotation.

The Taiwanese snacks Jensen Huang left on the desks at his executive briefing center are a nice human detail. But the real story is that Nvidia is methodically converting a GPU monopoly into a full-stack infrastructure monopoly—one rack, one CPU, one roadmap announcement at a time.

Your Turn

If Nvidia successfully owns both the GPU and CPU layers of AI infrastructure, what leverage do cloud buyers and AI developers actually have left—and is custom silicon the only realistic escape hatch?