The Memory Gap: AI's Most Honest Stress Gauge

Column Overview
Memory chips are getting more expensive again, and this time the reason has nothing to do with the usual boom-bust rhythm of the semiconductor industry. Laptops cost more to build, phone makers are quietly trimming RAM configurations, and PC builders are watching DDR5 prices climb month over month. It looks, on the surface, like just another memory shortage โ the kind the industry has lived through every few years since the smartphone era began. It isn't. What's happening to memory right now is a side effect of the AI capital race, and it turns out to be one of the more honest instruments we have for measuring that race's temperature.
A shortage that isn't shaped like the old ones
Three companies โ Samsung, SK Hynix, and Micron โ control roughly 95% of the global DRAM market. All three have spent the past two years redirecting new fabrication capacity toward High Bandwidth Memory, the specialized chip stacks that feed AI accelerators like Nvidia's GPUs. HBM is not a side project for them anymore; it's where the margins are, and margins dictate where capital goes.
The problem is that HBM is brutally expensive to produce relative to ordinary DRAM. Producing HBM3E at equivalent effective bit capacity consumes roughly 2.5 to 3 times the wafer resources of standard DDR memory; HBM4, the next generation now ramping, pushes that ratio closer to 3 to 4 times. Every wafer diverted into an HBM stack is a wafer that isn't making phone or laptop memory. Industry estimates put AI servers on track to consume 66% of total global DRAM output by 2026, leaving the remaining 34% to be split among phones, PCs, appliances, cars, and everything else that has ever used memory.
Layered on top of that is a demand-side squeeze few outside the industry noticed until prices moved. OpenAI and other frontier AI labs have reportedly locked in enormous multi-year wafer supply agreements with Samsung and SK Hynix โ deals said to add up to something close to 40% of global monthly DRAM output. That's capacity contracted years in advance, which means it never touches the spot market at all. Whatever ordinary manufacturers and consumers are competing for is what's left after both the HBM shift and these long-term reservations have taken their cut.
The timeline nobody wants to hear
The near-term outlook, as most analysts describe it, is not encouraging. Through 2026, the shortage is expected to remain structural rather than cyclical โ pricing pressure is likely to be sharpest in the first half of the year, with some easing possible in the second half as new HBM4 capacity comes online, though prices are expected to stay elevated regardless.
Nikkei Asia has reported that even with accelerated capacity expansion, global memory production by the end of 2027 is projected to satisfy only around 60% of market demand โ leaving a gap of roughly 40% still unresolved. The more widely cited relief window sits further out: DRAM shortages are broadly expected to ease sometime in 2027, with the possibility of oversupply emerging by 2028, as new fabs from Samsung, SK Hynix, and Micron finish construction and ramp production. Notably, when Micron and Taiwan's Nanya Technology addressed the topic in January 2026, both companies pointed to 2028 as the earliest realistic point of relief โ later than some of the more optimistic industry chatter had suggested.
At the pessimistic end of the range, SK Group's chairman has suggested the shortage could persist into 2030, while Japan's Kioxia expects tight NAND flash supply to continue at least through 2027. One variable that could meaningfully change the picture is China's CXMT (ChangXin Memory Technologies), which has laid out plans to reach 500,000 wafers of monthly capacity by 2028 โ enough additional supply to meaningfully cap how far the three incumbents can push prices.
Put together, the most defensible summary is this: the earliest signs of a turning point are likely by late 2027, with gradual easing through 2028, but prices are unlikely to fully retrace. This isn't a shortage that resolves โ it's a shortage that becomes the new baseline.
Why this cycle is structurally different
Every prior memory crunch โ the smartphone boom of the early 2010s, the 2017-2018 stretch when manufacturers coordinated production cuts into strong demand, the 2021 shortage driven by pandemic-era remote work and the automotive chip crisis โ was fundamentally the same story: total demand outran total capacity for a while, factories expanded, and prices came back down within a year or two. Standard DDR memory ran on shared production lines. A fab making memory for PCs could, with modest retooling, make memory for phones or servers. Capacity was fungible.
This time, capacity isn't fungible, and AI infrastructure sits at the front of the line by design. Advanced memory now depends on sub-15-nanometer DRAM process scaling, extreme ultraviolet lithography, through-silicon vias, and hybrid bonding for advanced packaging โ a set of capabilities that is globally scarce and notoriously hard to scale in yield. Producing a single HBM stack consumes wafer capacity equivalent to manufacturing dozens of conventional DDR chips. Meanwhile, older DDR4 production lines have already been decommissioned, and new fab construction realistically takes until 2028 to bear fruit. There's no quick retooling path back to the old equilibrium, because the old equilibrium's production lines no longer exist.
If the earlier shortages were seasonal colds, this one reads more like a change in body chemistry. AI hasn't just eaten into memory supply โ it has redefined the product mix, the customer hierarchy, and the underlying process technology of the entire industry.
The hardware gap is a projection of something bigger
Step back from the wafer math and the memory shortage starts to look like exactly what it is: the physical, measurable footprint of a much larger financial story. The technology driving all of this is genuinely real โ inference costs have fallen more than 99.7% over roughly two years, and that's not a marketing claim, it's a hard cost curve. But the pace of capital spending is running well ahead of the revenue that's supposed to eventually justify it. Capital expenditure among the five largest cloud providers is projected to hit roughly 33% of revenue in 2026 โ a historic high that exceeds even the 20% peak seen during the dot-com infrastructure buildout, and an investment intensity some analysts compare to the run-up before the 2000-2001 crash.
What makes this moment different from a simple bubble narrative is that the leading players are already diverging, and divergence is itself informative. In a pure hype cycle, everyone tells the same story and nobody's numbers are checked too closely. Once real performance starts separating winners from laggards, that's usually a sign the market has started grading on substance. In the US, OpenAI and Anthropic โ both selling frontier models, both burning enormous sums on compute โ are no longer running the same trajectory. OpenAI's losses have continued to widen even as user growth shows signs of plateauing. Anthropic, by contrast, has pushed gross margins above 70%, turned quarterly profitable, and now counts more than a thousand customers each spending over a million dollars annually. Two companies, same underlying technology wave, meaningfully different financial trajectories.
China is running a different playbook entirely
The picture outside the US looks even less like a monolithic AI boom and more like a genuine market bifurcation. Goldman Sachs recently raised its estimate for combined annualized recurring revenue across Chinese large-model companies in 2026 from $10 billion to $13 billion, with a projection that this figure could climb to $125 billion by 2030 โ a roughly 25-fold increase over five years. For context on scale, Anthropic's current annualized revenue alone, at roughly $44 billion, already approaches a third of that entire Chinese industry's present-day total.
Pricing tells its own story. API prices in China have fallen to 10-20% of 2024 levels, with DeepSeek offering rates as low as 0.001 yuan per thousand tokens โ among the cheapest of any mainstream model globally. The competitive axis has shifted from raw parameter count to efficiency and return on investment, and the business model is shifting in step, moving from "selling tokens" toward "selling outcomes." Goldman's framing labels 2025 the "DeepSeek moment," when mixture-of-experts architecture drove inference costs down sharply, and 2026 the "Zhipu moment," with the GLM model family demonstrating strong performance in coding and agentic tasks โ high-value categories where reliability commands a premium. Goldman projects China's large-model API and subscription revenue growing from roughly 35 billion yuan in 2026 to nearly 879 billion yuan by 2030.
One data point captures the shift particularly well: Zhipu raised prices by 83% on the strength of GLM-5.1's coding performance, and call volume rose 400% anyway. Professional users, it turns out, will pay more for reliability once they trust the output.
The more useful lens for understanding China's AI landscape, though, is the split between platform-affiliated and independent labs. Models built inside larger technology groups โ ByteDance's Doubao, Alibaba's Qwen, Tencent's Hunyuan, Baidu's Ernie โ don't need to raise independent capital or answer to outside investors about a standalone path to profitability. Their value is measured by the lift they give to a parent company's e-commerce, cloud, advertising, or productivity businesses. Independent labs โ DeepSeek, Moonshot's Kimi, MiniMax, Zhipu โ are a different animal entirely. They are the ones actually comparable to OpenAI and Anthropic: standalone model companies facing the full weight of price competition and the pressure to eventually turn a profit on their own. DeepSeek's now well-known achievement โ training DeepSeek-V3 on a cluster of 2,048 Nvidia H800 chips for roughly $5.58 million โ matters precisely because it's a rebuttal to the "just buy more compute" playbook that dominates the US market.
What the price tag on your next laptop is actually telling you
Here's the connective thread: memory prices spiking is, in its own blunt way, reassuring. It means the physical infrastructure buildout โ the fabs, the wafer contracts, the HBM production lines โ is still accelerating, not slowing down. Nobody has quietly stopped building. If the capital race were about to unwind, the first sign wouldn't be a headline; it would be underutilized fab capacity and softening long-term supply contracts. That isn't what's happening. The memory market, in that sense, is one of the most honest indicators available, precisely because it can't be spun โ wafer allocation is a physical fact, not a forecast on an earnings call.
What actually matters from here
The memory story is a symptom. The variables worth watching sit one layer up.
Whether model companies' unit economics are genuinely improving is the first and clearest signal โ Anthropic's margin trajectory reads as encouraging, OpenAI's user plateau does not. Second, and harder to see from outside: how much of the current capital buildout rests on circular financing arrangements between chipmakers, cloud providers, and model companies, and how much exposure that creates if any single link comes under stress โ this is arguably the single largest difference between this cycle and prior technology bubbles, and also the hardest to independently verify. Third, the interest rate environment matters more than almost anything else; historically, bubbles rarely burst because valuations got too high on their own โ they burst because money got more expensive. Fourth, whether China's independent labs can keep proving out a lower-cost training path gives the broader industry a fallback model for what survives if the higher-spend approach falters. Multiple industry analyses point to these as the load-bearing questions, even if none of them can be answered with precision today.
Zoom out further and the shape of this era starts to resemble a pattern that has played out before: the story turns out to be broadly correct, a number of well-funded companies don't survive, and the physical infrastructure outlives them. Nobody could have predicted in 2000 that Amazon would still be standing while Pets.com wouldn't โ but the fiber optic cable laid during that boom got resold cheap and quietly underwrote two decades of internet growth afterward. Data centers, HBM production lines, and GPU clusters built during this cycle are likely to follow the same arc โ repriced and reassigned, not torn out. Most industry forecasts place the memory market's shift from shortage to potential oversupply somewhere between late 2027 and 2029, with HBM likely to hold its value better, given its direct tie to AI chip roadmaps, while the standard DDR memory pushed out of the production queue today is both the first candidate for a rebound and the first candidate for a hard correction once new capacity finally lands.
The question of whether AI is useful has, for most practical purposes, already been answered โ it is. What remains genuinely unresolved is which of the companies racing to build it will still be standing by the time revenue actually catches up to spending.
As above, this reflects industry observation only and does not constitute investment advice.