header-langage
简体中文
繁體中文
English
Tiếng Việt
한국어
日本語
ภาษาไทย
Türkçe
Scan to Download the APP

SemiAnalysis Analysis: Rubin Ultra VRAM Reduced to 192GB, Will NVIDIA Also Compromise with HBM?

Read this article in 11 Minutes
VRAM reduced from 1TB to 192GB, power consumption down to 1800W
TL;DR
· SemiAnalysis reports that NVIDIA is showcasing a downgraded variant of the key customer preview Rubin Ultra, with the mainstream SKU featuring 192GB of VRAM.
· This specification has not been officially confirmed by NVIDIA, with key trade-offs being made in terms of HBM availability, power consumption, and NVLink 576 system deployment.
· The downgrade does not signal a decrease in AI demand; the high-power version may still be retained, but data center power capacity and the supply chain remain limiting factors.


A report released by SemiAnalysis on July 29th stated that NVIDIA is previewing the revised Rubin Ultra to key customers. According to their customer preview, the mainstream SKU will utilize 8-Hi HBM4 with 192GB of VRAM, falling short of the market's previously more aggressive expectations for Rubin Ultra.


This is not the final product specification officially confirmed by NVIDIA. The Rubin GPU unveiled in NVIDIA's technical blog on July 21st features dual compute dies, adopts HBM4 with a maximum of 288GB, 22TB/s bandwidth, and 50 PFLOPS NVFP4 performance, supporting the assessment that "Rubin is the next-generation key product." However, the 192GB mainstream version of Rubin Ultra should still be considered based on SemiAnalysis's disclosed customer preview.


For AI data center customers, VRAM capacity, HBM availability, single-card power consumption, and rack deployment difficulty will directly impact how training and inference systems are built, how quickly they are built, and how much they cost. If Rubin Ultra limits the mainstream VRAM to 192GB, it appears more like NVIDIA prioritizing a version that is easier to deliver at scale under limited HBM and power conditions, rather than pushing the individual GPU parameters to the maximum.


192GB VRAM Falls Below Early Expectations, Mainstream Version to Ensure Deployability First


Previously, at GTC 2025 and in materials from various brokerages and media, Rubin Ultra was often described as a next-generation product with higher VRAM, higher power consumption, and a more complex package, including more aggressive concepts such as HBM4E, 1TB, NVL576, and others.


The customer preview specifications provided by SemiAnalysis this time are notably more conservative: the mainstream version features 8-Hi HBM4 with 192GB of VRAM. This change should be understood as a downward adjustment compared to previous market expectations and early roadmaps, rather than NVIDIA formally announcing a specification and then retracting it.


VRAM is not an isolated parameter. Large-scale model training and long-context inferences require more high-speed memory. The larger the VRAM, the more model slices, caches, and data a single GPU can handle. However, HBM is also one of the tightest and most expensive components in current AI hardware. Using slightly less HBM in the mainstream version may help NVIDIA allocate limited supply to more GPUs and systems.


This is also where the concept of "downspec" can be easily misinterpreted. Current information does not directly support framing it as AI demand slowing down. A more plausible explanation is that in an environment where both HBM and power are constrained, mainstream products need to strike a balance between performance, cost, supply, and deployment.


Compute Narrative Shifts to NVL576, Customers Purchase the Whole System


According to SemiAnalysis, the peak theoretical FLOPs of the Rubin Ultra mainstream version are largely maintained, with adjustments mainly focusing on memory capacity, power consumption, and system form factor, rather than simply weakening computational power.


The emphasis of this approach is on NVL576. In simple terms, NVIDIA is still combining more GPUs through high-speed interconnects to form a larger unified computing domain, allowing customers to transition from single-card parameters to the system capabilities of entire racks or clusters. The report mentions that a planned system configuration is comprised of 8 interconnected 72-GPU racks, forming the NVL576 scale.


This is more practical for cloud providers and large AI labs. What limits the expansion of training clusters is not just how powerful a single GPU is, but also factors such as rack power consumption, cooling, power supply, networking, HBM supply, delivery pace, and data center construction. While a higher single-card memory capacity is certainly valuable, if it leads to higher power consumption, higher BOM costs, and more challenging rack deployment, customers may not be able to fully utilize it.


The main changes in Rubin Ultra are more about preserving computational performance and system scalability, while keeping memory and power consumption within a range that is easier to mass-produce, supply, and deploy.


1800W Becomes the Standard, Power Constrains the Extreme Configuration


Power consumption is another key figure in this shift in strategy.


SemiAnalysis states that the chip-level power consumption of the Rubin Ultra mainstream version is around 1800W, similar to the previous Rubin 1800W Max-Q configuration, but lower than the previously discussed 2300W version. The report also mentions that NVIDIA may still offer 2600W to 2800W Max-P high-power versions, and there may be a 1200W SKU designed for lower computing loads, such as token decode.


These details have not yet been officially confirmed by NVIDIA and should still be considered based on the report's perspective. But they point to the same real-world issue: even if the chip itself can run at higher power levels, customers' data centers may not be able to reliably support them.


The higher the power consumption of a single GPU, the greater the power supply and cooling requirements for the rack. Many customers' concerns are not "whether they want to buy a more powerful GPU" but whether their existing facilities, power connections, and cooling systems can support it. SemiAnalysis has previously mentioned that Rubin has two categories of power consumption configurations: 2300W Max-P and 1800W Max-Q. The Max-Q variant emphasizes performance per watt, and customers will choose to operate at a lower power level due to power constraints.


In a power-constrained environment, the 1800W version may actually be more suitable for mass production. It sacrifices some of the extreme configurations but helps reduce power delivery and thermal pressures, bringing it closer to a form factor that customers can deploy more quickly.


HBM Still a Bottleneck, Downsizing Not a Saturation of Demand


The tight supply of HBM is another main driver behind the reduction of memory to 192GB.


According to SemiAnalysis, HBM is relatively tighter in capacity compared to TSMC's frontend manufacturing. By reducing the amount of HBM per GPU, NVIDIA can allocate limited supply to more products and hedge against cost pressures from future HBM price increases. This assessment also needs to be seen in the context of the current tightness in the AI hardware supply chain.


This has a direct impact on the industry chain. If the mainstream GPU single-card HBM capacity is lower than the previous market expectations, HBM suppliers' unit demand estimates, cloud vendor procurement models, and server BOM will all change. However, this does not mean that HBM is not tight. On the contrary, it is precisely because of the tight HBM supply that NVIDIA is motivated to adjust the mainstream solution to be more HBM efficient and easier to mass produce.


The Rubin Ultra may not necessarily be left with only one version. According to SemiAnalysis, the high-power Max-P version may still be targeted at specific customers and extreme scenarios, just not necessarily becoming the most widely deployed mainstream SKU.


This disclosure appears more like a product of AI computing power expansion entering the engineering constraint phase. NVIDIA has not stopped pushing for larger system scales, but whether the next-generation GPU can be smoothly mass-produced increasingly depends on HBM supply, data center power, and customer deployment capabilities, rather than just chip-level parameters.



Welcome to join the official BlockBeats community:

Telegram Subscription Group: https://t.me/theblockbeats

Telegram Discussion Group: https://t.me/BlockBeats_App

Official Twitter Account: https://twitter.com/BlockBeatsAsia

举报 Correction/Report
Choose Library
Add Library
Cancel
Finish
Add Library
Visible to myself only
Public
Save
Correction/Report
Submit