← Back
SiTech Team⏱️ 7 წთ. საკითხავი

NVIDIA Vera Rubin GPU Architecture — 336 Billion Transistors, 10x Agentic Throughput Over Blackwell

NVIDIA Vera Rubin GPU Architecture — 336 Billion Transistors, 10x Agentic Throughput Over Blackwell

NVIDIA unveils the Vera Rubin platform — 336 billion transistors, HBM4 memory, and third-generation Transformer Engine delivering up to 10x more agentic throughput per unit of energy than Blackwell for agentic AI workloads.

The Era of Agentic AI — Introducing NVIDIA Vera Rubin

What began as discrete AI model training and human-facing chat interfaces has evolved into always-on AI factories dedicated to producing intelligence at scale. These factories are now tasked with powering agentic workflows that reason, plan, use tools, verify intermediate results, and execute complex multistep tasks across vast contexts. Agentic workloads are not defined by a single prompt and response, but by sustained inference across many reasoning steps.

They demand low per-step latency, high decode throughput, efficient long-context attention, large KV cache capacity, and the ability to scale models across tightly coupled GPU domains. The data center must be reimagined as a single unit of compute — a vision realized with the NVIDIA Vera Rubin platform, announced on July 21, 2026.

Rubin GPU — 336 Billion Transistors of Compute Density

At the core of the platform is the NVIDIA Rubin GPU, designed to deliver up to 10x more agentic throughput per unit of energy than NVIDIA Blackwell. This is not merely a benchmark figure — it represents a fundamental generational leap, enabling AI factories to produce more useful tokens within the same energy envelope.

The Rubin GPU is constructed from 336 billion transistors, featuring 224 streaming multiprocessors (SMs) and 896 Tensor Cores. The chip is fabricated from reticle-limited compute dies unified on a single package through the high-speed NV-HBI (NVIDIA High-Bandwidth Interface) inter-die link. This approach achieves the high density and efficiency required for modern AI workloads.

The architecture organizes compute resources into Graphics Processor Clusters (GPCs) with a large centralized L2 cache. The GigaThread Engine coordinates work, MIG Control partitions the GPU for multiple workloads, and NV-DEC accelerates decoding. Together, these capabilities turn raw compute density into sustained utilization across the diverse and dynamic workloads that define large-scale agentic systems.

Third-Generation Transformer Engine — 50 Petaflops of NVFP4

One of Rubin's crown-jewel innovations is the third-generation Transformer Engine. This engine adapts precision across numerical formats, enabling the Rubin GPU to deliver up to 50 petaflops of NVFP4 inference performance while preserving accuracy. Enhanced Tensor Cores with expanded precision flexibility work alongside this engine to accelerate agentic workloads efficiently.

Rubin doubles Tensor Core throughput per clock by doubling the amount of data it can process along the K dimension. This optimization benefits not only throughput-bound kernels but also memory-bound and latency-bound kernels. As model execution is split across many GPUs, each GPU often receives a smaller slice of output work while the reduction dimension remains large. The benefit of a larger K dimension is fewer K loop iterations — a GEMM requiring four K iterations on Blackwell can be completed in just two iterations on Rubin.

HBM4 Memory — 288 GB at 22 TB/s Bandwidth

Rubin integrates up to 288 GB of HBM4 memory, driven by dedicated HBM controllers and 12-Hi stacks, delivering up to 22 TB/s of peak bandwidth. HBM4 doubles interface width relative to HBM3e and, combined with a new memory controller, deep co-engineering with the memory ecosystem, and tighter compute-memory integration, this subsystem provides a 2.8x bandwidth increase over Blackwell and Blackwell Ultra.

High-capacity, high-bandwidth HBM4 is critical for hosting multi-trillion-parameter models, extending context length without KV cache offload, and supporting high-concurrency, long-horizon inference workloads. The enhanced Tensor Memory Accelerator (TMA) manages high-efficiency movement across complex data layouts, ensuring data reaches compute cores with minimal latency.

NVLink 6 and NVLink-C2C — Blazing Fast Interconnects

In the Rubin architecture, communication infrastructure is as vital as compute power itself. NVIDIA NVLink 6 provides 3,600 GB/s scale-up bandwidth to the NVLink Switch for all-to-all GPU-GPU communication. NVLink-C2C delivers 1,800 GB/s for coherent CPU-GPU communication, while x16 PCIe Gen 6 provides up to 256 GB/s of host connectivity.

Rubin introduces counted writes for device-initiated NVLink communication, streamlining synchronization for GPU-to-GPU data transfers. This allows the receiving GPU to track transfer completion more efficiently, providing a lower-latency mechanism that keeps computation moving instead of waiting on synchronization operations.

Architecture Optimized for Agentic AI

The Rubin GPU is purpose-built for the execution patterns of agentic AI — reasoning, planning, tool use, and long-context processing. This demands not only peak compute performance but also efficient data movement, long-context attention processing, and smooth transitions between dependent kernels.

Rubin accelerates attention by combining activation sparsity with adaptive compression and improved softmax throughput. The attention pipeline begins with a dense QK^T computation to generate intermediate attention scores. Rubin can then load that intermediate data from Tensor Memory into a structured 2:4 sparse compressed form, reducing write cost and storage requirements. This compressed representation reduces work in two key places: softmax and the second attention GEMM, improving tokens/watt without requiring the surrounding model pipeline to change.

Rubin also increases exponential throughput — including 2x FP32 and 4x BF16/FP16 throughput versus Blackwell — helping softmax keep pace with faster matrix operations. Additionally, Rubin enables more fine-grained coordination between dependent kernels, allowing consumer work to begin earlier as required input data becomes available, rather than waiting for a larger set of producer work to complete.

Vera CPU — Olympus Cores for Maximum Performance

The Vera Rubin platform also includes the Vera CPU, built on Olympus cores designed for maximum single-thread performance. This CPU provides tight coordination with the GPU via NVLink-C2C, which is especially critical for agentic workflows where CPU and GPU constantly exchange data and control signals.

Vera Rubin NVL72 — Rack-Scale AI Supercomputer

Vera Rubin NVL72 extends the Rubin GPU into a rack-scale execution domain for multi-trillion-parameter and multi-rack agentic workloads. Its third-generation MGX rack architecture combines cable-free compute and switch trays, 45°C liquid cooling, dynamic rack-scale power steering, and Intelligent Power Smoothing to keep compute, networking, cooling, and power operating as one unified system.

With DSX MaxLPS, Vera Rubin power supplies use state-of-charge Intelligent Power Smoothing to absorb transient power swings, reducing average power consumption by approximately 10% and 50 ms peak power by approximately 20%. This enables operators to provision up to 40% more GPUs within the same power budget at energy-efficient operating points, with minimal impact on workload performance. Spectrum-6 networking arrives for gigascale AI factories, and hot-swappable NVLink switch trays with improved RAS capabilities further support resilient operation at scale.

Confidential Computing with TEE-I/O

For large-scale agentic AI deployments, data security is paramount. NVIDIA introduces Confidential Computing with TEE-I/O, designed to secure data at rest, in transit, and in use across the AI factory. This ensures that sensitive data — whether proprietary model weights, customer information, or research data — remains protected throughout its lifecycle.

Bristol Myers Squibb — Life Science's Most Advanced AI Factory

To demonstrate the real-world potential of the Rubin platform, Bristol Myers Squibb is building life science's most advanced AI factory on Rubin. This move underscores a key message: Vera Rubin is not just another GPU generation — it is a platform that will enable entire industries to move to the next level of AI adoption, from drug discovery to genomics and beyond.

Conclusion

The NVIDIA Vera Rubin GPU architecture represents a watershed moment in AI infrastructure evolution. With 336 billion transistors, HBM4 memory, the third-generation Transformer Engine, and up to 10x agentic throughput over Blackwell, it not only meets the demands of today's AI workloads but lays the foundation for tomorrow's agentic systems. Backed by 300 global partners and major cloud providers including CoreWeave, Google Cloud, Microsoft Azure, and Mistral, Vera Rubin is already ramping up worldwide. With Vera Rubin, NVIDIA continues its vision of AI factories that run continuously, more efficiently, more securely, and at a grander scale than ever before.

📖 Source