TL;DR:
AI memory shortage may persist throughout the entire cycle. Morgan Stanley believes that the continuous growth in model size, context length, and inference concurrency will keep absorbing newly added storage supply.
Nvidia has begun offering a "reduced-spec" option for Rubin. Some HBM and LPDDR5 capacities may be lowered, but demand has not disappeared—it is shifting toward NAND, DRAM, and high-speed interconnects.
The AI inference architecture is being restructured. The separation of Prefill and Decode is expected to improve hardware utilization and create opportunities for technology solutions such as Cerebras and Nvidia Groq.
CXL is expected to become a new growth driver. Morgan Stanley estimates that the related semiconductor market will reach approximately $6 billion by 2030, with Astera Labs and Marvell as potential beneficiaries.
The key for memory stocks is the duration of the cycle, not the short-term magnitude of price increases. Morgan Stanley maintains overweight ratings on Micron and SanDisk, with the main risk being a slowdown in AI investment and data center construction.
Editor's note: AI's next bottleneck may no longer be the number of GPUs, but memory. As large models scale up, context windows grow, and AI Agents continue to consume more inference resources, supply pressure on HBM, DRAM, and NAND is transmitting across the entire data center. More notably, in the face of memory shortages, Nvidia is exploring product options with lower memory configurations for its next-generation Rubin platform to maintain system delivery and deployment pace.
This raises a seemingly contradictory question: if AI demand for storage is so strong, why are chipmakers instead beginning to consider reducing memory capacity? Does memory "downspeccing" mean demand is about to peak, or does it indicate that the supply bottleneck has become so severe that it is forcing the entire industry to adjust its technology roadmap?
In an October 5 research report titled "How Can the AI Ecosystem Work Around Memory Bottlenecks?" by Morgan Stanley semiconductor analyst Joseph Moore and others, the report proposed that the memory shortage may persist throughout the current AI cycle, but AI buildout will not wait for DRAM fabs to complete capacity expansion. The industry needs to bypass existing supply constraints by reducing configurations for some products, splitting inference tasks, and improving memory sharing efficiency.
This means that competition in AI hardware may shift further from simply increasing GPU and memory capacity to the resource utilization efficiency of the entire computing system. The memory shortage does not necessarily weaken long-term demand for memory manufacturers, but it could change the value distribution across the supply chain. In addition to memory companies such as Micron and SanDisk, manufacturers with specialized computing architectures and interconnect technologies, such as Cerebras, Astera Labs, and Marvell, may also gain new growth opportunities from this.
The following is a translation of the original text:
AI infrastructure construction is facing a problem that is becoming increasingly difficult to solve: computing power can be expanded by purchasing more GPUs, but memory supply cannot increase at the same pace.
Morgan Stanley believes that for the foreseeable future, AI demand may continue to absorb large amounts of new memory supply. Even if the degree of shortage fluctuates across different periods, DRAM capacity expansion will still find it difficult to quickly eliminate the supply-demand gap.
In conversations with companies in the computing industry, Morgan Stanley found that how to bypass the memory bottleneck has become an increasingly important topic. NVIDIA CEO Jensen Huang also discussed the need to alleviate supply constraints through the co-design of computing, networking, and storage systems.
Analysts believe this is not a warning that memory demand is weakening, but rather a reality the entire AI industry must face: when memory cannot be supplied as originally planned, companies must find new system architectures so that existing hardware can continue to support AI growth.
The most direct approach is to reduce the memory capacity used by a single server or a single GPU.
Morgan Stanley points out that tight DRAM and NAND supply and rising prices are forcing the AI supply chain to reconsider product configurations. For chip manufacturers, rather than sticking to original specifications and causing systems to fail to be delivered as planned, it is better to offer versions with different memory capacities so customers can choose according to actual needs.
The research report predicts that NVIDIA may offer multiple lower-memory configurations for the Rubin platform.
In terms of rack-level memory, Rubin's originally planned LPDDR5 capacity was approximately 54TB, while some new configurations may be reduced to 28TB, corresponding to SOCAMM2 memory module capacity dropping from 192GB to 96GB.
In terms of HBM, Rubin's single GPU was originally expected to be equipped with 288GB HBM4, using an 8-group 12-layer stacking design; Morgan Stanley expects that NVIDIA may add a product version equipped with 192GB HBM4, reducing the number of stacked layers to 8.
For Rubin Ultra, the research report argues that after adjusting to a dual compute chiplet design, 512GB is a more appropriate comparison baseline, and different capacity options ranging from approximately 192GB to 384GB may emerge in the future.
These are all Morgan Stanley's judgments on product configurations as of the research report's publication date, and not all specifications have been officially confirmed by Nvidia.
From a cost perspective, this strategy has its rationale. Reducing the number of HBM stacks can lower capacity, while theoretical bandwidth does not necessarily decline in tandem when conditions such as the number of interfaces and pin speeds remain unchanged.
However, capacity and bandwidth address two different problems. For applications with smaller models and shorter contexts, reducing capacity may have limited impact; but as models grow larger and inference tasks become more complex, insufficient memory capacity can still lead to more data needing to be transferred between different storage tiers, increasing latency or GPU usage.
More importantly, reducing HBM does not make the data that AI originally needs to process disappear—it only forces that data to find new places to reside.
For example, when GPU HBM capacity is insufficient, some data may need to be stored in main memory; when rack-level LPDDR5 capacity is reduced, some KV Cache (key-value cache) demands may shift to NAND flash.
KV Cache is a cache used during large model inference to store historical computation information. As contexts grow longer and the number of concurrently running requests increases, its capacity requirements rise accordingly.

Storage demand does not disappear—it shifts to other tiers.
Morgan Stanley notes that while Nvidia is reducing certain LPDDR5 configurations, the industry chain has already observed new demand coming from NAND. Reduced HBM capacity may also increase data transfers between GPUs, putting greater pressure on high-speed interconnect networks.
This means that Nvidia's configuration-reduction strategy may temporarily ease DRAM supply constraints, while simultaneously increasing demand for NAND, interconnect chips, and larger-scale GPU clusters.
This pressure shift has its long-term backdrop.
According to historical data from Epoch AI cited in the research report, the parameter scale of frontier models in the large language model era once showed a growth trend of doubling approximately every six months. Context windows have also undergone rapid expansion. At the same time, more AI applications are shifting from simple Q&A to coding, complex reasoning, and long-running Agent tasks, further driving up memory usage.
Morgan Stanley believes that model size, context length, and inference concurrency will continue to drive growth in storage demand. Therefore, reducing configurations for some products is more likely a transitional measure during periods of supply tightness, rather than a signal of a long-term decline in AI memory demand.
If reducing memory configurations is a short-term response, then changing the computing approach for AI inference may bring about deeper industry changes.
The second path Morgan Stanley focuses on is Disaggregated Inference, which means splitting different inference tasks that were originally concentrated in the same computing system and assigning them to hardware that is better suited for each.
Traditional large model inference mainly consists of two stages.
The first stage is Prefill, which processes the user's input prompt, files, or historical context and performs large-scale parallel computation. This stage typically relies more on computing performance.
The second stage is Decode, in which the model generates output tokens one by one. This process requires repeated access to model weights and KV Cache, and therefore relies more on memory bandwidth, capacity, and data access latency.
In the past, the same set of GPUs often handled both types of tasks simultaneously. But as inference demand expands, this configuration may not always be the most efficient choice.
Morgan Stanley believes a more reasonable direction may be to let compute-intensive hardware handle Prefill, while architectures optimized for low latency and high bandwidth take on Decode.
This approach can not only improve hardware utilization, but also allow data centers to expand the two types of computing resources separately, without having to configure exactly the same GPUs and HBM for all tasks.

The Technical Difference Between Prefill and Decode
One of the most direct potential beneficiaries is Cerebras.
Cerebras adopts a wafer-scale processor architecture, integrating a large number of compute units and SRAM onto the same chip. SRAM (Static Random Access Memory) typically has lower capacity density than DRAM, but features low latency and high bandwidth, making it suitable for certain inference tasks that require frequent data reads.
By bringing compute units closer to data, Cerebras can reduce the need for external memory access and improve processing speed for specific inference stages.
Morgan Stanley noted that Cerebras has partnered with AMD and AWS to explore combining different processors into a unified inference system.
In the joint AMD-Cerebras solution, AMD Helios handles the more compute-intensive Prefill and large-context processing, while the Cerebras wafer-scale engine handles Decode. The research report estimates that this solution will enter production in the fourth quarter of 2026.
AWS is taking a similar approach, using Trainium for Prefill and Cerebras for Decode, with the combined solution expected to enter Amazon Bedrock in the first quarter of 2027.
Based on performance data disclosed by Cerebras and AMD for specific solutions, the combination is expected to achieve up to approximately five times throughput improvement while maintaining Cerebras's inference speed. This does not mean all inference tasks can achieve the same level of improvement, but it indicates that configuring hardware for different tasks may significantly improve system economics.
For Cerebras, this shift means not only more chip sales opportunities. Since the company also operates its own inference cloud service, if the same hardware can generate more tokens, the cost per token could decline, thereby improving gross margins or providing room to lower customer prices.
Nvidia is also exploring a similar direction.
By integrating Groq-related technology, Nvidia is combining HBM-based GPUs with LPU architecture using high-speed SRAM. In the solution described in the research report, Rubin GPUs handle Prefill and some decoding computations requiring large-capacity cache, while the Groq architecture takes on operations better suited to its low-latency characteristics, with NVIDIA Dynamo software coordinating tasks across different processors.
It is worth distinguishing that Nvidia and Groq reached a technology licensing and talent recruitment arrangement in 2025, not a full acquisition of Groq.
Morgan Stanley believes that as inference spending grows faster than training, heterogeneous computing architectures optimized for different inference stages may gain greater market space. According to the research report's infrastructure spending forecast, by 2030, inference is expected to account for approximately 56% of AI infrastructure spending, with training accounting for approximately 44%.
This also brings a new competitive logic to the AI chip industry: in the future, the metrics for measuring the competitiveness of an AI system may not only include peak compute power, but also how many tokens are generated per second, and how much cost is required to generate each token.

In addition to changing the way inference tasks are allocated, Morgan Stanley also focuses on how to improve the efficiency of using existing memory resources. This is exactly where the opportunity for CXL technology lies.
CXL (Compute Express Link) is a high-speed interconnect protocol that enables processors to access external memory more flexibly, without always having to rely on directly attached local DRAM.
In traditional servers, memory resources are often tied to a specific CPU. Even if one server has a large amount of idle memory, it is difficult to make it directly available to another server. CXL improves the overall utilization of DRAM in data centers through memory expansion, sharing, and pooling. Among these, memory expansion can provide additional capacity for a single processor; memory sharing allows multiple processors to access the same memory resources; and memory pooling organizes scattered memory into a unified resource pool that is dynamically allocated according to workloads.
This technology was previously used more in traditional CPU computing environments, because CPU workloads are relatively more tolerant of the additional access latency brought by external memory. By contrast, AI GPUs have long relied on HBM to provide extremely high data bandwidth, and external DRAM connected via CXL is difficult to directly replace HBM.
However, as AI inference demands change, the value of this architecture is rising.
For example, long-context inference requires storing a large amount of KV Cache, but not all cached data must always remain in the fastest HBM. The system can keep performance-sensitive data in HBM and move some data that is accessed less frequently and has relatively looser latency requirements to external DRAM.
This can both ease the capacity pressure on expensive HBM and improve the utilization efficiency of existing DRAM.
Morgan Stanley believes that AI is driving CXL from traditional server memory expansion toward a broader market for accelerator memory connection and sharing.
In the past, Astera Labs estimated the potential market size of CXL memory controllers at more than $4 billion. Morgan Stanley currently expects that, with the growth in demand for AI inference, KV Cache offloading, and rack-level memory pooling, by 2030 the potential market size of CXL and related memory interconnect semiconductors is expected to reach about $6 billion.

It should be emphasized that this figure represents analysts' projection of the potential market space, not market revenue that has already been realized.
On specific companies, the research report focuses on Astera Labs (ALAB) and Marvell (MRVL).
Astera Labs' opportunity mainly comes from its Leo memory controller products. Morgan Stanley noted that the company expects standardized and customized Leo products to begin volume ramp-up in 2027 with two major U.S. cloud service provider customers, including a KV Cache offloading design for AI inference.
Marvell, through Structera X and Structera S, is positioning itself in memory expansion and rack-level memory pooling, respectively. The company previously projected that its CXL business could contribute more than $1 billion in revenue around 2028.
For both companies, the AI storage shortage could generate new commercial demand for CXL technology, which had previously been slow to gain traction.
However, Morgan Stanley also acknowledged that the actual size of the CXL market remains difficult to predict precisely. Different cloud providers may opt for non-CXL interconnect technologies, higher-capacity HBM, or flash-based caching solutions, and the ultimate pace of market adoption could differ significantly from current expectations.
Therefore, whether the CXL investment thesis can be realized still requires monitoring cloud providers' actual deployment progress, product shipment volumes, and related business revenue.
For the memory supply chain, Nvidia reducing certain memory configurations does not appear to be good news. If HBM capacity per GPU declines and LPDDR5 usage per rack decreases, memory makers may be able to sell fewer memory units than originally planned.
Morgan Stanley acknowledged that from the perspective of short-term profit maximization, this could indeed erode some of the demand and pricing opportunities that memory makers would otherwise have captured. But analysts believe this should not be simply interpreted as the memory cycle coming to an end. The key point is that the current downsizing is primarily driven by supply constraints, not because customers no longer need as much memory.
When AI companies want to procure more DRAM but cannot secure sufficient supply, reducing configurations can help them maintain system deliveries. Once supply increases in the future, companies will still have incentives to raise configuration levels again. This means that currently unmet demand could form a certain backlog of future purchasing space.
Based on this assessment, Morgan Stanley is more focused on the duration of the memory industry's upcycle rather than the peak magnitude of short-term price increases. If AI models continue to scale, Agent applications keep proliferating, and inference concurrency continues to grow, then even if DRAM manufacturers increase capacity, the additional supply could be rapidly absorbed.
As a result, Morgan Stanley maintains its Overweight ratings on Micron (MU) and SanDisk (SNDK), believing that memory supply tightness will not end quickly.
But this does not mean the memory industry can completely escape cyclical risks. The research report argues that the most fundamental risk still comes from AI demand itself. If the growth rate of model development, inference applications, or compute investment slows significantly, the entire computing and memory supply chain could be impacted.
Additionally, there is a more complex scenario: AI demand remains robust, but data center construction is delayed due to land, power, or infrastructure constraints. In this scenario, memory products may have already been manufactured but cannot be absorbed as expected because server deployment has been pushed back, resulting in a temporary supply-demand imbalance.
Morgan Stanley believes that such construction bottlenecks could create additional disruptions for memory manufacturers, potentially even causing their short-term performance to diverge from that of some computing chip companies. Therefore, to judge whether the memory cycle can continue going forward, one must not only monitor DRAM prices and HBM orders, but also track the actual deployment progress of AI infrastructure.
The ultimate impact of the AI memory shortage may not be that the industry reduces its use of memory, but rather that it forces the industry to reconsider: which data must remain in HBM, which can be moved to DRAM or NAND, and which computing tasks should be assigned to different types of chips.
For investors, the core variables to watch going forward are also becoming clearer: whether Rubin's actual shipment configurations are reduced, whether decoupled inference can achieve commercial-scale deployment, whether CXL products can ramp up as planned by 2027, and whether AI data center construction continues to absorb new memory supply.
If these technological adjustments can sustain AI system deployment while long-term memory demand continues to grow, then the memory and interconnect supply chains could benefit together; but if AI investment slows, or power and land constraints continue to hinder data center buildouts, the current supply-demand tightness could also face repricing.
This is precisely Morgan Stanley's most important judgment on this AI memory cycle: what the industry truly needs to solve is not how to wait for more memory, but how to continue expanding AI computing power while memory remains perpetually scarce.
Welcome to join the official BlockBeats community:
Telegram Subscription Group: https://t.me/theblockbeats
Telegram Discussion Group: https://t.me/BlockBeats_App
Official Twitter Account: https://twitter.com/BlockBeatsAsia