Investment & M&A
The Bottleneck Moved From GPUs to Memory: SK Hynix Just Sold Out All of 2026
Published: 2026-06-24
SK Hynix has pre-sold its entire 2026 DRAM, HBM and NAND supply. Quarterly operating profit hit a record KRW 11.4 trillion ($8.02B), up about 62% year over year. The AI buildout’s bottleneck is shifting from GPUs to memory, and any AI startup planning capacity now faces compute cost and availability risk at once.
What Happened
SK Hynix said it has effectively pre-sold all of the DRAM, HBM and NAND it will produce this year, meaning its entire 2026 output is locked into contracts. Nvidia is the anchor customer, and a letter of intent with OpenAI for HBM supply has been reported. The earnings back the picture: quarterly operating profit hit a record KRW 11.4 trillion ($8.02 billion), up roughly 62% year over year, on revenue up about 39%. The industry is no longer calling this a cyclical upswing, it’s a “super-cycle.” Every AI accelerator stacks multiple layers of high-bandwidth memory (HBM), and as models grow and context windows lengthen, that demand rises steeply rather than linearly. Supply can’t catch up on a short horizon: HBM requires advanced packaging, so adding lines takes time, and as a result memory prices are climbing. For the past couple of years the AI compute bottleneck was framed as “can you get GPUs.” This signal says the bottleneck is moving one notch, from the GPU die itself to the memory bolted onto it.
What This Means for Founders
For any AI startup that buys or rents compute, this throws two risks at once. The first is cost: when memory prices rise, GPU instance pricing and self-hosting economics rise with them, and that volatility is a variable you don’t control. The second is availability: with 2026 output already committed to large customers, lower-priority buyers may not get what they want, when they want it. If your fundraising deck says “inference cost falls over time,” put an asterisk on that assumption. Per-token model pricing can drop while the unit cost and procurability of the physical compute that runs the model move the other way. This is where strategy splits. A design that runs ever-larger models at ever-larger volume is exposed directly to memory pricing. The opposite, efficient smaller models, quantization, and open-weight models you can host on your own infrastructure, shrinks that exposure. The ability to do the same work with less memory becomes both margin defense and a supply-risk hedge. The natural read-across: a near-frontier open-weight model you self-host reduces dependence on a single supplier’s pricing, but it doesn’t make the underlying memory cheaper, it just gives you more control over how much of it you burn.
What You Can Do Now
First, model compute cost as variable, not fixed. Put a memory-price-increase scenario into your 12-to-18-month plan and check that unit economics still hold under it. Second, promote efficiency to a product decision: are you routing tasks a small model could handle to a large one, and can quantization hit the same quality on less memory? Third, evaluate open-weight models seriously, a near-frontier model you can run on your own infrastructure cuts your exposure to a vendor’s price moves. Fourth, don’t bet compute on a single source. Lock into one cloud and one instance type and you have nowhere to run when prices climb or allocation tightens. The teams that treat memory as a planned constraint, not a surprise, are the ones who keep shipping when the super-cycle squeezes.
Sources
Related Content