# NVIDIA Vera Rubin and Blackwell Ultra: Setting a New Bar for Agentic AI Performance per Watt
## NVIDIA Unveils Vera Rubin and Blackwell Ultra at GTC 2025
NVIDIA’s next-generation data-center roadmap has become a focal point for the AI infrastructure conversation, centered on two names: **Blackwell Ultra** and **Vera Rubin**. According to accounts circulating around GTC 2025, NVIDIA CEO Jensen Huang introduced both the Blackwell Ultra GPU and the Vera Rubin chip during the conference — a claim that the sources reviewed here do not independently confirm, and which should be treated as reported rather than settled. The framing that ties the two products together comes most directly from NVIDIA’s own developer materials, which present Vera Rubin and Blackwell as a combined step forward for [agentic AI performance per watt](https://developer.nvidia.com/blog/nvidia-vera-rubin-and-blackwell-set-a-new-standard-for-agentic-ai-performance-per-watt/). Independent corroboration of the specific launch details below remains open.
Where possible, the discussion here separates what NVIDIA itself asserts from figures that have not been cross-checked against primary sources.
## Jensen Huang’s GTC 2025 Keynote: What Was Announced
Reports attribute the unveiling of Blackwell Ultra and Vera Rubin to Jensen Huang’s GTC 2025 keynote, but the exact contents of that keynote are not verified by the sources available here. Community discussion consistent with a GTC-timed announcement can be found in [contemporaneous forum threads](https://news.ycombinator.com/item?id=43409280), though such threads are commentary rather than authoritative confirmation and should not be read as establishing the specifics.
What can be stated with more confidence is the through-line NVIDIA itself emphasizes: positioning its current and next-generation silicon around efficiency for multi-step, agentic workloads rather than raw throughput alone. The particulars of the keynote — dates, on-stage demos, and precise product framing — are not corroborated in the material reviewed here and are left open.
## Vera Rubin: Expected in Late 2026 with a 3.3x Performance Boost
Two figures dominate coverage of Vera Rubin: a **late-2026 availability window** and a **3.3x performance boost**. Both are presented here as uncorroborated. Sources reviewed do not confirm the timing, and the 3.3x figure lacks a verified baseline — it is unclear from the available material against which product, workload, or metric (training, inference, tokens per watt, or per-rack throughput) any such multiplier would be measured. Performance multipliers of this kind are typically vendor-supplied and highly configuration-dependent, so the number should be read as a reported target rather than an independently validated result.
Absent a primary specification, the safe reading is that Vera Rubin is *expected* around late 2026, with a performance uplift *said to be* on the order of 3.3x — with both claims still requiring confirmation from NVIDIA’s own published specifications.
## Vera Rubin Ultra: The 2027 Successor Targeting a 14x Performance Improvement
Beyond Vera Rubin, a successor referred to as **Vera Rubin Ultra** is reported to be slated for **2027**, aiming for a **14x performance improvement**. As with the figures above, the sources here do not independently verify the 2027 timing, the “Ultra” positioning, or the 14x target, and the baseline for that multiplier is not established. A 14x claim, if accurate, would again depend heavily on the comparison point and workload, and it is not clear from available evidence whether it refers to a full rack-scale system, a single accelerator, or a specific agentic-inference benchmark.
These figures are therefore best characterized as an unconfirmed roadmap signal. Nothing in the reviewed material settles whether Vera Rubin Ultra’s specifications, ship date, or performance headline will hold as described.
## Blackwell Today, Rubin Next: How the Rollout Reaches U.S. Customers
A recurring narrative frames the rollout as sequential: **Blackwell chips already in use by U.S. customers, with Rubin following soon after**. The reviewed sources do not independently confirm the deployment status of Blackwell among U.S. customers, nor the specific sequencing that would place Rubin “next.” This should be treated as an unverified characterization of the rollout rather than a documented deployment timeline.
Any account of *which* customers, *what* volumes, or *when* Rubin transitions from roadmap to shipping product is outside what can be substantiated here and remains open. The generational ordering — Blackwell in market, Rubin approaching — reflects the reported roadmap shape, not a confirmed schedule.
## A New Standard for Agentic AI Performance per Watt
The most concrete anchor for the “new standard” framing is NVIDIA’s own developer material, which explicitly presents Vera Rubin and Blackwell as [setting a new standard for agentic AI performance per watt](https://developer.nvidia.com/blog/nvidia-vera-rubin-and-blackwell-set-a-new-standard-for-agentic-ai-performance-per-watt/). It is important to attribute this precisely: the claim that NVIDIA *positions* its parts this way is grounded in NVIDIA’s own publication, but that is company framing, not an independently verified benchmark result. The reviewed sources do not corroborate the underlying performance-per-watt figures, so the marketing framing and the measured reality should be kept distinct.
Performance per watt is a meaningful axis for data-center economics: at scale, power and cooling constraints often bind before raw compute does, so efficiency gains translate directly into cost and deployable capacity. Whether NVIDIA’s specific efficiency claims meet the “new standard” bar is not something the available evidence establishes, and independent measurement would be needed to confirm it.
## Why Agentic AI Needs It: From Single-Turn Inference to Multi-Step Workflows, Tools, and Subagents
Underlying the efficiency emphasis is a shift in how inference is consumed. The broader industry description — echoed in NVIDIA’s own framing — is that AI agents have expanded inference from single-turn interactions into multi-step workflows that reason across several steps, invoke external tools, and coordinate subagents. This characterization aligns with NVIDIA’s [agentic AI positioning](https://developer.nvidia.com/blog/nvidia-vera-rubin-and-blackwell-set-a-new-standard-for-agentic-ai-performance-per-watt/), though the reviewed material does not independently quantify how much this pattern increases per-task compute or token generation.
The intuition is straightforward even where the numbers are unconfirmed: a single agentic task may issue many model calls, retrieve context repeatedly, and spawn subordinate agents, each of which consumes tokens and energy. If that pattern holds, efficiency per watt becomes more consequential than in a single-turn chatbot regime, because the total work per user request rises substantially. That directional argument is reasonable, but it is presented here as rationale rather than as a measured claim — the sources reviewed do not confirm the magnitude of the shift, and the specific hardware figures that would tie it to Vera Rubin or Blackwell remain uncorroborated.
For the networking layer behind agentic AI infrastructure, see NVIDIA BlueField-4 and Scale-In Networking for Agentic AI Factories.
