Compute Is Now AI’s Scarcest Layer: Amazon Expects to Fall Short of Demand Through 2027

Amazon plans to spend about $220 billion on capital expenditures this year, and CEO Andy Jassy still told investors the company will not have enough capacity to meet all of its 2026 demand. He expects the same in 2027. When the largest cloud provider cannot spend its way out of a shortage, the scarce resource in AI is the capacity to run models, and the money moving through the sector is following it.

The Shortage Shows Up in Cloud Earnings

Jassy said on Amazon’s second-quarter call that the company will fall short of demand in 2026 even at that level of spending. “I believe this dynamic will also be true in 2027, too,” he said. Amazon reported an AWS backlog of $496 billion, growing at triple-digit rates year over year, and AWS revenue rose 37% to $42.2 billion.

The other large providers describe a similar squeeze. Microsoft’s Azure revenue grew 43% in its fiscal fourth quarter, and CFO Amy Hood said demand continued to exceed available capacity. Google Cloud revenue rose 82% to $24.8 billion, and its backlog reached $514 billion. Meta guides to $130 billion to $145 billion in capital spending for 2026.

Company 2026 capital spending guidance Cloud demand signal
Amazon About $220 billion (cash capex) AWS backlog $496 billion; AWS revenue up 37% to $42.2 billion
Alphabet $195 billion to $205 billion Google Cloud backlog $514 billion; revenue up 82% to $24.8 billion
Microsoft About $175 billion (fiscal 2027) Azure revenue up 43%; demand exceeded available capacity

Figures come from each company’s July 2026 earnings reports and calls. Microsoft’s fiscal 2027 began in July 2026, so its figure is not a calendar-year number.

By their own guidance, Amazon, Alphabet and Meta alone plan to spend roughly $545 billion to $570 billion in 2026. A backlog counts signed commitments that have not yet become revenue. Set beside Amazon’s and Microsoft’s capacity warnings, these figures describe buyers who want more compute than sellers can deliver today.

A $14.5 Billion Valuation Built on Backlog

The Wall Street Journal reported on October 6 that Lambda, an Nvidia-backed AI cloud provider, is raising up to $4 billion in a final private round before an IPO it is targeting for 2027. Blackstone and Coatue Management are leading the round, which values the company at $14.5 billion not counting the new money. The round has not closed, and Lambda declined to comment.

The number that explains the valuation is the backlog. According to the Journal, Lambda’s backlog grew from $15 billion in June to $50 billion in September. A backlog is not revenue. It converts only as Lambda delivers the capacity it has sold, which makes delivery the main risk behind the price.

The reading that fits the evidence is that investors are valuing Lambda on contracted demand for scarce capacity. Lambda sells access to the resource Amazon and Microsoft say they cannot supply fast enough, and its backlog suggests customers are signing ahead of delivery. The caution is just as plain: a valuation built on backlog depends on Lambda delivering the capacity it has sold.

Model Developers Are Now Customers of GPU Clouds

GMI Cloud announced a $668 million financing on October 1. It combines $223 million of Series B equity, led by San Francisco investment firm ARCHIV with NVIDIA participating, and a $445 million credit facility led by CTBC. The company did not disclose a valuation.

GMI says its contracted annual recurring revenue now stands at more than nine times its year-end 2025 level, while live ARR in production has grown more than 4.5 times over the same period. The platform processes about four trillion tokens a week. The company says the money will support capacity expansion in the United States, Taiwan and the rest of Asia-Pacific, along with inference services and hiring.

Named customers include Fireworks AI, OpenRouter and Nous Research, companies that build, serve or route AI models. Fireworks co-founder Chenyu Zhao called GMI “one of our strongest and most reliable providers across NVIDIA GB200 and GB300 NVL72 systems.” The praise centers on reliability, which is the quality a capacity shortage puts to the test.

The financing structure adds a signal of its own. About two-thirds of the total, $445 million of $668 million, is debt. The gap between the two ARR figures points the same way, since the multiple on contracted revenue is double the multiple on live revenue. Signed demand is running ahead of the capacity already serving customers, and most of the new money is borrowed.

Inference Grows While Orbit Stays a Test

Deloitte predicted in November 2025 that inference would account for roughly two-thirds of all compute in 2026, as AI moves from training models to answering enterprise and consumer questions. It also expected most of that work to stay on expensive, power-hungry AI chips worth $200 billion or more, sitting mainly in large data centers rather than at the edge. That is a forecast, but GMI’s four trillion weekly tokens shows what the inference side looks like at one provider.

Orbit is the far end of the search for capacity. Google launched a prototype Project Suncatcher satellite carrying four of its TPUs on October 1 aboard a SpaceX rocket, built with Planet Labs. The mission tests how the chips handle spaceflight, radiation and temperature extremes. Project lead Travis Beals said: “I don’t see this being something where it’s cheaper to do this in the next five years. I think it will take longer than that.”

SpaceX has gone further on paper. In January it asked the FCC for approval to operate up to one million satellites as orbital data centers, and its filing claims that launching one million tonnes of satellites a year could add 100 gigawatts of AI compute annually. Orbital compute matters as a signal of how far buyers will look for power and space. On the Suncatcher lead’s own timeline it will not beat ground-based compute on cost for more than five years, so the capacity that matters to buyers today sits in data centers on the ground.

Governments Treat Capacity as a Strategic Asset

Epoch AI’s cluster tracking puts about three-quarters of global GPU cluster performance in the United States and 15% in China, based on data from May 2025 that the group notes includes planned and uncertain entries. Epoch also estimates that Amazon, Google, Meta, Microsoft and Oracle together held 71% of the world’s cumulative AI compute as of the fourth quarter of 2025, up from 63% in the first quarter of 2024. A handful of companies own most of the capacity everyone else is trying to rent.

Washington regulates the flow of chips directly. A Commerce Department rule effective January 15, 2026 moved licensing for H200-class chips to China and Macau from a presumption of denial to case-by-case review, with conditions that include third-party testing and a cap on aggregate exports at 50% of U.S. shipments. Beijing is funding capacity from its side: National Development and Reform Commission spokesperson Jiang Yi said on July 31 that computing power networks are expected to receive 4 trillion yuan, about $589.2 billion, in new direct investment between 2026 and 2030, and that the build-out relies mainly on enterprise investment.

Enforcement is part of the picture. Epoch estimates that between 290,000 and 1.6 million H100-equivalents were smuggled to China through 2025, with a median estimate of 660,000, roughly a third of China’s total compute. On October 1 the Justice Department announced the arrest of Greg Lui of Earthmade Computer, who is charged with conspiring to smuggle more than $300 million in export-controlled GPU servers to China through Malaysia and Singapore. An indictment is an allegation, and Lui is presumed innocent.

The implication is that governments now license, finance and police compute directly. A provider’s capacity plan depends partly on those rules, alongside chips and power.

What the Shortage Means for Enterprise Buyers

The implication for enterprise buyers is practical. Most companies will rent capacity rather than build it, and nothing in these earnings reports suggests that renting stops working. The risk lies in who supplies it. Amazon and Microsoft say demand exceeds what they can deliver, and GMI Cloud’s contracted revenue is growing at twice the multiple of its live revenue.

My take: capacity belongs in the procurement conversation next to model choice. A team planning production agents should ask each provider what capacity is reserved for it, from what date, and what happens if delivery slips. An agent that runs all day draws on inference all day, so a late delivery becomes a delayed product.

Two numbers will test this account over the next year. One is whether Lambda converts a $50 billion backlog into delivered capacity on its way to a 2027 listing. The other is whether Amazon’s forecast of a shortage in 2027 holds or supply catches up sooner. If the shortage persists, valuations tied to capacity will look justified. If it eases, they will face scrutiny.

The post Compute Is Now AI’s Scarcest Layer: Amazon Expects to Fall Short of Demand Through 2027 appeared first on DataFLOQ.

Leave a Reply

Your email address will not be published. Required fields are marked *

Subscribe to our Newsletter