Hyperscaler debt, the depreciation fight, and how an industrial can think about AI.

AI use : 25% — argument, structure and most sources provided by the author; drafting and data sourcing assisted by AI, full human review. All figures as of August 2026. This is not financial advice.

Rain, Steam and Speed — The Great Western Railway. 1844, oil on canvas

Every industrial company I talk to is signing AI contracts right now. Multi-year commitments, agentic systems wired into planning, quality, maintenance, procurement. Sensible moves. What almost none of them have examined is the balance sheet of the supply chain underneath those contracts.

That supply chain is currently the largest debt-financed capital expenditure cycle in corporate history. Its core asset has a lifespan nobody agrees on, disputed by a factor of two to three. And its unit economics get repriced roughly every quarter by open-weight labs, most of them Chinese. You cannot predict how this resolves. You can, however, write contracts that survive every resolution. That is the argument of this essay, and it ends with four procurement rules we apply across our own portfolio of twenty-two industrial software companies and actively recommend to industrials.

The buildout has become macroeconomics

Start with the scale, because the scale is the story. The five largest US hyperscalers spent $405 billion on capital expenditure in 2025. Estimates for 2026 sit between $600 and $750 billion depending on which bank you ask; Goldman Sachs pencils in close to $1.2 trillion for 2027. Computing infrastructure has roughly doubled its share of US GDP since 2022, to about 1.5 percent, with AI the driver (Epoch AI, June 2026).

In the first quarter of 2026, investment in computers and software contributed about 1.09 percentage points to the 2.0 percent annualized real GDP growth of the United States. All of household consumption contributed 1.08. A category worth under four percent of GDP matched a category worth sixty-eight percent of it. S&P Global estimates that data center and related investment accounted for roughly 80 percent of the growth in US private domestic demand in the first half of 2025. Strip the buildout out and the American economy is close to flat.

Now the financing. In 2025 roughly a quarter of hyperscaler capex was funded with investment-grade debt, about $108 billion of issuance. Goldman expects that share to reach a third in 2026, around $250 billion, and 35 percent in 2027, around $400 billion. Morgan Stanley counts nearly $570 billion of AI-related debt issuance across the ecosystem in 2026 alone, and a $1.5 trillion gap between planned data center capex through 2028 and what operating cash flows can cover. Private credit lending to AI companies went from about $3 billion in 2010 to over $40 billion in 2025 (BIS). Moody’s counts roughly $660 billion of hyperscaler data center lease commitments sitting off balance sheet, more than these same companies carry as adjusted on-balance-sheet debt. Add the special purpose vehicles, the neoclouds, and the vendor financing loops between chipmakers, labs, and cloud providers, and the honest description is: a lot of debt, in a lot of places, with the visibility worst exactly where the leverage is highest. The market has noticed. Oracle’s credit default swaps roughly tripled after September 2025. Alphabet raised $84.75 billion of equity in June 2026, an odd move for a company famous for buying its shares back, not selling new ones.

What does this do to GDP over time? The candid answer comes from a Federal Reserve staff paper that modeled data center investment’s contribution to 2027 growth and produced a range from plus 1.2 percentage points to slightly negative, depending on the scenario. When the referee’s own models disagree on the sign, the summary is that nobody knows. Solow’s old paradox, computers everywhere except in the productivity statistics, has returned as a live question with twelve zeros attached. The investment shows up in GDP today because spending is counted when it happens. The productivity is a forecast.

Why should a plant director or an industrial CFO care about any of this? Because your AI suppliers sit on top of this capital structure, and leverage does not stay upstream. It travels down supply chains the way it always has: as price volatility, as contract rewrites, as vendor mortality.

The depreciation fight is really a fight about your future prices

In November 2025, Michael Burry made a niche accounting question famous. His claim: hyperscalers depreciate GPU fleets over five to six years while the real economic life of the hardware is closer to two or three, given that Nvidia ships a new generation, two to three times better, on a roughly annual cadence. By his math, the industry will understate depreciation by about $176 billion between 2026 and 2028, flattering reported profits at some firms by more than 20 percent. Nvidia’s counter is that customers observe four to six years of productive use.

Interestingly, the operators themselves diverged. Amazon shortened the useful life of a subset of its servers from six years to five, citing the pace of AI hardware development, and absorbed roughly $700 million of operating income to do it. Meta lengthened its estimates in the same period. CoreWeave had already stretched from four years to six back in 2023. Groq’s founder argued the honest number is one year. And Satya Nadella, whose company is spending over $100 billion a year on this hardware, said the quiet part: “the biggest competitor for any new Nvidia chip is its predecessor,” adding that Microsoft deliberately avoids loading up on a single chip generation. Same physics, same chips, five different clocks.

Both camps also flatten a real subtlety: a GPU has more than one life. A chip retired from frontier training cascades down to inference, then to batch and internal workloads. The cascade is real. But its economics depend on power prices and on demand for cheap tokens, which loops us straight into the demand question below.

Industrials are well placed to read this argument, because you depreciate machines for a living. You know that a wrong depreciation assumption does not change reality; it only changes when reality arrives. If the true life of the fleet is closer to three years than six, then today’s token prices under-recover capital, and the correction shows up later in one of three forms: higher prices, fire-sold capacity, or dead vendors. Capex-heavy industries usually deliver all three at once. The practical translation: every AI price you are quoted today embeds someone’s guess about a depreciation schedule, and nobody agrees on the guess.

Efficiency is a weapon, and it cuts in two directions

While American balance sheets absorb the buildout, Chinese labs turned efficiency itself into the product. DeepSeek made frugality a strategy: its V4 line runs a mixture-of-experts architecture that decides per task how hard to think, with a budget tier priced around twenty cents per million input tokens where Western frontier models charge dollars. Alibaba’s Qwen 3.6 delivers frontier-grade agentic coding from a model that activates about three billion parameters per token, under an Apache license, on hardware a mid-sized company can own. Moonshot published the full weights of a 2.8-trillion-parameter model in July 2026. The market voted: OpenRouter’s routing data shows Chinese open-weight models going from a rounding error to a majority of routed tokens in about eighteen months.

The result is deflation without precedent in industrial inputs. For constant capability, inference cost has fallen roughly tenfold per year since 2021; GPT-4-class performance that cost tens of dollars per million tokens in early 2023 costs well under a dollar today. One academic estimate has constant-quality inference cost halving every 2.6 months. Moore’s law, the previous record holder, needed two years to do that.

Falling prices for the input. What happens to total demand? Two honest unknowns.

First, nobody knows whether the compute required per unit of delivered intelligence goes up or down from here. Sparsity, distillation, and better data curation push it down. Test-time reasoning, where models spend more compute thinking rather than being bigger, pushes it up. Serious people hold each view.

Second, there is Jevons. In 1865, William Stanley Jevons observed that Watt’s fuel-efficient steam engine multiplied England’s coal consumption rather than reducing it, because cheaper use invents new uses. Commentators tend to merge this with Solow’s paradox; they are different animals. Solow asks whether the investment shows up in productivity. Jevons asks whether efficiency shrinks or multiplies demand. On the second question the early data is loud. OpenRouter’s analysis of more than 100 trillion routed tokens found average prompt size per request growing fourfold in thirteen months, with reasoning models passing half of all token consumption by mid-2025. An agent that plans, calls ten tools, writes and tests code, fails, and retries burns orders of magnitude more tokens than the chat turn it replaced. Inference has overtaken training as the majority of AI compute. Enterprise spending on AI grew several-fold in 2025 in the middle of a price collapse. Tech has run this movie many times: storage, bandwidth, compute cycles. Cheaper always meant more, in aggregate.

But note the asymmetry inside the deflation. Two-year-old capability is approaching free; frontier reasoning remains scarce and rationed. Whether Jevons holds at the frontier, not just at the commodity floor, is what decides whether $700 billion a year of capex earns a return. Nobody knows that either. Which is why a supplier who claims certainty about it is telling you about their marketing, not their economics.

A procurement doctrine for industrials

We deploy AI into more than 3,800 factory sites through our portfolio. What follows is not theory; it is the pattern of the invoices.

Treat every price as a policy, not a cost

The application layer sits between you and the leveraged, deflating machine described above, and the price it shows you is a strategic choice. Often it is a subsidy: ICONIQ’s January 2026 survey puts the average AI-native gross margin at 52 percent, with inference alone eating 23 percent of revenue; flat-rate plans go negative on the heaviest users; Replit disclosed outright negative gross margins; a wave of repricing followed across coding and agent tools through 2026. Sometimes it is the opposite, a markup on inference whose underlying cost is collapsing. The same sticker can be below cost this quarter and above cost the next, and either way it will move. So ask three questions. Who pays for inference, on which models, at what realized cost per completed task? What happens to my price when your model mix changes? Can you show me the unit economics of my workflow? Suppliers who answer plainly are structurally lower risk, whatever the answer is. The ones who dodge are telling you the price is a policy and you are the variable.

Plan for model churn, not model loyalty

The model your pilot ran on will not be the model your production runs on in two years, maybe even two months. Capability rankings reorder every quarter, and it is now normal for the best model for scheduling, for document extraction, and for shop-floor copilots to come from three different labs. A buyer locked to a single provider pays frontier prices forever and inherits none of the tenfold annual deflation; a buyer whose stack can swap models inherits every price drop the moment it lands. The well-run suppliers already work this way internally, routing most calls to cheap models and reserving frontier ones for the hard residue. Contract for portability and per-workflow evaluations, not for a logo.

Unbundle the three layers

There are three distinct businesses in your AI bill: the agentic application, the model, and the compute underneath it. They have different economics. The application layer carries the workflow lock-in and the margin. The model layer deflates tenfold a year and churns. The compute layer is where the debt and the disputed depreciation live. A bundled contract hides which layer you are paying for and lets the vendor arbitrage between them at your expense. Unbundling does not require three contracts; it requires pricing each layer separately even inside one contract, with the right to re-source each. This is more feasible than it has ever been. Open weights make the model layer genuinely portable, including into your own cloud tenant when the data must not travel, a point that matters for European industrials. And compute is closer to a spot commodity than at any moment since 2022.

Refuse token-plus

The worst pricing model in this market is the supplier who meters your tokens and adds a margin. It looks transparent. It is cost-plus on an input the seller controls, which means the seller earns more when the agent is verbose, when calls are redundant, when nothing is cached and nothing is routed. The incentive divergence starts on day one and compounds with usage. The best players we know pay for their own inference and sell you a fixed price or an outcome. The behavioral difference shows up within weeks: routing, caching, distillation, prompt discipline, because every wasted token now comes out of their margin instead of yours. Outcome pricing also puts the cost of a failed agent run on the vendor’s P&L, which is where it belongs. And a supplier who owns their inference bill inherits the deflation curve on your behalf; competition eventually hands it to you.

The precedent, and where it breaks

We have watched a debt-financed infrastructure race end before. Around 2000, carriers laid hundreds of billions of dollars of fiber on borrowed money. Most of them died. The glass survived, bandwidth prices fell more than 90 percent, and the lasting winners were the buyers who had refused twenty-year capacity contracts at 1999 prices and bought transit as it deflated. The analogy breaks in exactly one place, and the break makes today’s cycle sharper: fiber in the ground stayed competitive for decades, while a GPU ages in a few years and its replacement is two to three times better. The asset depreciates while its substitute improves. Both the boom and the correction run faster.

You do not need to know who ends up holding the debt. You need to be structured so that either resolution pays you. If Jevons wins and demand explodes, an unbundled, multi-model stack scales without renegotiation. If the depreciation clock wins and the debt sours, you are not contractually married to a casualty, and you get to buy the fire sale. That is the whole doctrine. Refuse to make the macro bet. Collect the deflation. And choose suppliers whose incentives point at your output rather than your token count.

Sources