BlockBeats news, September 15 — TMT Breakout, citing a SemiAnalysis report in its latest morning note, said Nvidia's Vera Rubin platform could achieve up to about 7x the tokens-per-unit-power throughput of Blackwell in pre-release testing. Even if Rubin's per-chip hourly rental cost is higher than Blackwell Ultra, under an inference scenario of 80 TPS and P90 latency, the token output corresponding to its three-year total rental cost could still be about 62% higher, and the advantage under high-interaction workloads may widen further.
However, alongside rising performance expectations, a new variable has emerged in the system-level delivery cadence of Rubin Ultra. The Kyber NVL144 rack, originally planned to carry 144 Rubin Ultra GPUs, may be delayed by more than 12 months from its original 2027 target due to PCB midplane manufacturing challenges, pushing it to 2028; the alternative NVL72x2 is also said to have been canceled.
This risk will mainly fall on high-density racks, NVLink interconnect, and large-scale system integration. Even if the Rubin chip itself can proceed on schedule, cloud providers' pace of obtaining 144-GPU-class scale-up capability in 2027 may still be constrained, and AI infrastructure investment will also rely more on alternatives such as the existing Oberon architecture.
Nvidia only responded to the related report by saying that "the product roadmap remains unchanged," and has not yet confirmed the Kyber delay or an overall Rubin Ultra timeline adjustment.

