Press Enter to search · ESC to close

AI × Crypto

AI’s Memory Crunch: How the Industry Is Rewiring Itself Around a Structural Shortage

Memory shortages have become the most persistent structural constraint in the AI buildout, prompting the industry to cut memory specs, disaggregate inference workloads, and pool memory via CXL. The workarounds create new winners in CXL controllers and workload disaggregation while memory makers like Micron and SanDisk remain favored on a duration-over-amplitude thesis.

AI’s Memory Crunch: How the Industry Is Rewiring Itself Around a Structural Shortage

Memory shortages have become the most persistent structural constraint in the current AI buildout cycle, and the pace of AI compute expansion shows no sign of waiting for new fabs to come online. Facing that contradiction, the industry is pursuing three workarounds — cutting memory specifications, disaggregating inference workloads, and pooling memory through CXL — while creating new investment opportunities along the way.

Nvidia CEO Jensen Huang recently said the industry needs entirely new thinking to address the memory bottleneck, framing it not as a pessimistic signal but as a proactive acknowledgment of a shortage expected to persist for years. The intensity of the shortage may fluctuate, but AI is expected to absorb essentially all available supply for the foreseeable future.

Spec Reduction: A Stopgap That Moves Pressure, Not Eliminates It

With DRAM and NAND prices both climbing, the most direct response is to compress memory configurations per device or per server. Nvidia is offering multiple specification tiers for its Rubin platform: LPDDR5 capacity per rack has been compressed from a planned 54TB to 28TB, with modules shrinking from 192GB SOCAMM2 to 96GB. For HBM, Rubin was originally slated for 288GB per GPU (8-high 12hi HBM4), but a 192GB version (8-high 8hi HBM4) is now expected. Rubin Ultra specifications are also being scaled back, with the comparable baseline potentially falling to a 192GB to 384GB range.

Yet spec reduction is not a true solution. When memory capacity is cut at one layer, data pressure inevitably shifts downward — shrinking HBM and LPDDR creates incremental demand for NAND and network interconnect. The bottleneck moves rather than disappears. Three structural forces keep pushing AI memory demand higher: model sizes doubling every six months in the LLM era, context windows expanding roughly fivefold annually, and rising inference concurrency requiring independent KV cache storage per session. Spec reduction can only be temporary; once supply eases, the industry will quickly revert to higher specifications.

Workload Disaggregation: Dedicated Hardware for Each Inference Phase

Inference consists of two distinct phases: prefill, which processes input prompts in parallel and is compute-intensive, and decode, which generates output token by token and is more dependent on memory bandwidth and capacity. Separating the two and assigning each to the hardware best suited for it is the second path to improving memory efficiency.

This trend accelerated noticeably around late 2025. After acquiring Groq, Nvidia integrated its LPU — based on high-bandwidth on-chip SRAM — into the Vera Rubin platform, with the Rubin GPU handling prefill and decode requiring large KV cache capacity, Groq handling feed-forward and mixture-of-experts components, and Nvidia Dynamo coordinating activation transfers. AWS then announced a heterogeneous architecture using Trainium for prefill and Cerebras for decode, planned for Amazon Bedrock in Q1 2027; a similar Cerebras-AMD collaboration is expected to enter production in Q4 2026. Matrix’s Corsair has also entered mass production.

Cerebras stands out as the most direct pure-play beneficiary of workload disaggregation. Its wafer-scale processor integrates large amounts of SRAM with compute on the same silicon, significantly reducing data movement between processor and external memory — a decisive advantage in the decode phase. Disaggregation can also materially improve the economics of Cerebras Cloud: according to disclosures from Cerebras and AMD, the joint system can boost throughput by up to 5x while maintaining Cerebras inference speeds, generating more revenue per deployed unit and lowering cost per token.

CXL: From Memory Expansion to a New AI Inference Battleground

CXL (Compute Express Link) is a high-speed interconnect protocol built on the PCIe physical layer. Its core value is allowing multiple processors to access a shared memory resource pool, fundamentally breaking the traditional architecture in which memory is bound to a specific CPU.

Its three main applications are memory expansion (adding capacity beyond a single processor’s local memory channel limit), memory sharing (multiple processors accessing the same data), and memory pooling (dynamically allocating memory across processors or servers to reduce idle waste). Historically, CXL mainly served general-purpose CPU computing, and its penetration in AI was limited by latency disadvantages. But as inference workloads demand sharply more capacity for long context, agentic workloads, and KV cache, much data does not need to reside permanently in HBM — giving CXL-attached DRAM a foothold as a lower-cost, higher-capacity memory tier.

On market size, the addressable market forecast for CXL has been raised from an earlier CPU-based estimate of over $4 billion to roughly $6 billion by 2030, with the increment mainly from new AI server-side deployment scenarios. Astera Labs’ Leo memory controller series (including custom designs for KV cache offload, expected to enter mass production in 2027) and Marvell’s full product portfolio covering memory expansion (Structera X) and rack-level pooling (Structera S) are two core names favored in this space. Marvell management has previously positioned CXL as a revenue contributor of over $1 billion around 2028.

Implications for Memory Stocks: Duration, Not Amplitude

From a short-term profit-maximization perspective, these workaround strategies have some negative implications for DRAM makers — if the AI ecosystem stalled entirely due to memory shortages, near-term pricing could be more favorable. But that scenario never materialized. The current memory cycle should be analyzed with duration as the primary driver rather than price amplitude.

The core argument: spec reduction stems from supply constraints, not fading demand. Once supply improves, the industry will quickly revert to higher specifications, and pent-up demand will easily absorb new capacity. One risk factor to watch: if an AI slowdown stems not from declining demand but from infrastructure bottlenecks such as land, power, or fab capacity, such a pause could hit memory names differently than compute names.

Micron and SanDisk remain overweight-rated, with the view that the memory shortage cycle is far from over and their allocation value still holds.

Key Takeaways for Investors

  • Memory shortage is structural, not cyclical: Model scaling, context window growth, and inference concurrency are compounding demand faster than supply can respond.
  • Spec reduction is a pressure valve, not a fix: It shifts demand from HBM/LPDDR to NAND and networking rather than eliminating it.
  • Workload disaggregation creates pure-play winners: Cerebras is the most direct beneficiary as prefill and decode split onto dedicated hardware.
  • CXL is the emerging AI inference battleground: The addressable market has been revised up to ~$6 billion by 2030, with Astera Labs and Marvell as core enablers.
  • Anchor on duration, not amplitude: The memory cycle’s length matters more than price swings; Micron and SanDisk remain preferred exposures.

View original

Share
Risk notice This site provides news and information on the crypto, blockchain and Web3 industry for reference only and does not constitute investment advice or any promise of returns. Virtual currency-related activities are illegal financial activities in mainland China; digital asset prices are highly volatile; use at your own risk. This site does not provide trading, token issuance or related referral services.

Related Reading

Latest News

TREE NEWS share card
Long-press image above → Save to Photos / Share
Pitch us Feedback