If AI keeps getting smarter, can the infrastructure behind it keep up?
Artificial intelligence is rapidly moving beyond simply generating answers. Modern AI systems are becoming capable of reasoning through problems, using external tools, retrieving information, executing tasks, and deciding what to do next. This is the idea behind Agentic AI, where AI can work through multiple steps to achieve a goal rather than simply responding to a single prompt.
This evolution is also making AI workloads more complex and demanding. An AI agent may need to perform several rounds of inference, access memory, retrieve data, call APIs, execute code, and interact with other systems or agents. Supporting these workloads requires more than just powerful GPUs. CPUs, memory, networking, storage, and data movement all play an important role in keeping the entire system running efficiently.
As AI continues to advance, the infrastructure supporting it must evolve as well. The focus is shifting from simply making individual components more powerful to designing complete systems where compute, memory, networking, storage, and data movement can work together efficiently at scale.
This broader shift in AI infrastructure sets the stage for NVIDIA Vera Rubin. Rather than viewing AI computing as a collection of individual components, Vera Rubin takes a more integrated approach, bringing together compute, memory, networking, and other technologies at large scale. It offers a glimpse into how the infrastructure behind AI is evolving to meet the demands of increasingly capable and autonomous workloads.
The Technology Behind NVIDIA Vera Rubin
NVIDIA Vera Rubin is more than a collection of powerful chips. At the center of the platform is the Vera Rubin POD, a large-scale AI computing system made up of five specialized rack-scale systems. Together, these systems combine computing, memory, networking, and storage technologies into a unified infrastructure designed for demanding AI workloads. NVIDIA describes the POD as an AI supercomputer built for the growing requirements of the agentic AI era.
One of the key systems within the POD is NVIDIA Vera Rubin NVL72, which integrates 72 Rubin GPUs with 36 Vera CPUs. The Rubin GPUs provide the accelerated computing power for AI workloads, while the Vera CPUs are purpose-built to support data movement, agentic reasoning, and other CPU-side workloads around AI applications. High-speed technologies such as NVLink 6 connect the processors and allow them to communicate efficiently, helping the system operate as a tightly integrated computing platform.
The POD also includes technologies that handle the data movement, networking, and storage requirements around AI computation. NVIDIA BlueField-4 powers the BlueField-4 STX storage platform, which hosts NVIDIA CMX, an AI-native context memory tier designed for long-context, multi-turn, and agentic AI inference. NVIDIA Spectrum-6 SPX provides the high-speed networking infrastructure that connects the different systems within the POD and the broader AI infrastructure.
The result is an architecture that goes beyond simply adding more GPUs to individual servers. By connecting multiple specialized systems into a larger platform, the Vera Rubin POD is designed to provide the compute, memory, storage, bandwidth, and connectivity needed for large-scale AI workloads.
Vera Rubin Architecture in Simple Terms
Think of the Vera Rubin POD as a large, highly organized school. The Rubin GPUs are like the students handling difficult assignments, while the Vera CPUs act like teachers and staff managing the tasks around them. Memory and storage are like the school’s books and filing systems, providing the information they need, while networking technologies act like the communication system connecting classrooms and buildings. The POD brings all of these parts together so they can work as one coordinated environment. Instead of simply having more powerful computers, the goal is to make sure every part can communicate, share information, and contribute efficiently to the overall workload.
Vera Rubin Built for Agentic AI
The architecture of Vera Rubin becomes particularly important as AI moves toward more agentic workloads. Unlike systems that simply generate a response, AI agents can work through multiple steps to complete a task. They may reason about a problem, retrieve information, use tools, execute code, and evaluate results before taking the next action. This repeated processing makes agentic workloads more demanding than traditional AI inference.
One challenge is managing the growing amount of context an AI agent needs during these tasks. As an agent goes through more steps, it may need to retain information from previous actions, retrieved data, and tool outputs. NVIDIA CMX helps address this by providing a shared context tier that extends GPU memory and is optimized for KV cache, supporting longer and more complex AI workloads.
Speed is another important factor. Since an agent may perform several rounds of inference before completing a task, even small delays can add up. Groq 3 LPX is designed to support low-latency inference, while the high-speed interconnects and networking technologies across Vera Rubin help move data efficiently between computing, memory, storage, and other resources.
Ultimately, Vera Rubin is designed around the changing requirements of AI. As AI systems become more capable of reasoning, maintaining context, using tools, and performing multiple operations, the infrastructure supporting them must become more capable as well. Vera Rubin’s integrated approach is designed to provide the computing, memory, inference, and communication capabilities needed to support increasingly complex and autonomous AI workloads.
What Makes Vera Rubin Stand Out?
With the key components of Vera Rubin and its role in supporting Agentic AI now established, the bigger question is: what actually makes it different from the AI infrastructure that came before it?
Vera Rubin is not simply a new generation of GPUs. NVIDIA’s focus is on improving the efficiency of the entire AI system, particularly as workloads shift toward reasoning and Agentic AI.
More AI Performance With Fewer GPUs
One of the clearest differences can be seen in how much computing infrastructure is required to accomplish the same task. NVIDIA says Vera Rubin NVL72 can train large mixture-of-experts models using one-fourth the number of GPUs required by NVIDIA Blackwell. This means that scaling AI does not necessarily have to mean adding GPUs at the same rate as the workload grows.
For large AI deployments, this can have a significant impact. Fewer GPUs can mean less infrastructure to deploy, power, cool, and manage while still achieving the required level of compute.
Designed for the Economics of AI
Another area where Vera Rubin stands out is efficiency. NVIDIA positions the platform around performance per watt and cost per token rather than raw compute performance alone.
According to NVIDIA, Vera Rubin NVL72 can deliver up to 10x higher inference throughput per watt and one-tenth the cost per million tokens compared with NVIDIA GB200 NVL72 for the workloads highlighted by NVIDIA. These improvements are particularly relevant for inference, where AI systems may need to operate continuously and serve large numbers of users or agents.
This shift changes the way organizations evaluate AI infrastructure. As models become more capable and inference becomes a larger part of AI workloads, the question is no longer just how much a system can compute. It is also how efficiently it can produce useful AI output.
Built for the Agentic AI Era
Vera Rubin also stands out because NVIDIA designed it specifically around the changing behavior of AI workloads. Agentic systems can generate much more context and perform many more inference steps than traditional single-turn applications.
NVIDIA reports that the Rubin GPU can deliver up to 10x more agentic throughput per unit of energy than Blackwell on its internal workload. The platform also combines Rubin GPUs with technologies such as Vera CPUs, Groq 3 LPX, and CMX to address the compute, inference, and context requirements of these workloads.
A More Integrated Approach to AI Infrastructure
Perhaps the biggest difference is the level at which NVIDIA is optimizing the system. Blackwell established NVIDIA’s rack-scale approach to AI computing, while Vera Rubin extends that approach through deeper co-design across compute, networking, storage, power, and cooling. NVIDIA’s approach increasingly treats AI infrastructure as a coordinated system that extends from individual chips and racks to entire AI factories.
That makes Vera Rubin significant beyond its individual specifications. Its improvements are aimed at reducing the amount of infrastructure, energy, and cost required to deliver AI at scale.
The Future of AI Infrastructure
NVIDIA Vera Rubin represents a broader shift in how AI infrastructure is being designed. As AI systems become more capable of reasoning, using tools, managing larger contexts, and performing increasingly complex tasks, infrastructure must evolve beyond simply providing faster GPUs. Compute, memory, networking, storage, power, cooling, and data movement all need to work together efficiently to support AI at scale. Vera Rubin demonstrates this direction by bringing these technologies together into an integrated platform designed for the demands of modern and agentic AI workloads.
What I find particularly significant about Vera Rubin is its systems-level approach to AI infrastructure. Its design shows how performance increasingly depends not only on individual processors, but on how effectively compute, memory, networking, power, cooling, and other components work together. To me, this highlights that the future of AI will depend not only on more capable models and faster processors, but also on the engineering of the infrastructure that enables them to operate efficiently at scale.
References:
https://developer.nvidia.com/blog/nvidia-vera-rubin-pod-seven-chips-five-rack-scale-systems-one-ai-supercomputer/
https://developer.nvidia.com/blog/inside-nvidia-rubin-gpu-architecture-powering-the-era-of-agentic-ai/
https://nvidianews.nvidia.com/news/vera-rubin-full-production-agentic-ai-factory
https://www.nvidia.com/en-us/data-center/vera-rubin-nvl72/
https://www.nvidia.com/en-us/data-center/technologies/rubin/















