AI compute is set to become the most expensive noun in business history. Here is what the thing actually is, what you are really paying for, and why the company that makes the chips will only backstop part of what they turn out to be worth.

Ask any executive what AI compute is and you will usually get a wave at a warehouse of chips. That loose definition now has a $500 billion financing target riding on it.

So, plainly. AI compute is the whole technology stack that trains and runs AI models: the chips, plus the memory, networking, power, cooling and software that make them usable.

Note the “AI” in front. Compute on its own means something much wider. Amazon still sells it as a general-purpose category covering servers, containers and serverless code. The gap between the two isn’t academic. It’s how someone says the word confidently and still doesn’t know what they’re buying.

The chips are the part that gets pictured. They aren’t the part that decides whether the money makes sense. In everyday business use, compute has gone from a verb to a very expensive noun in about twenty years.

It used to be a verb. To compute meant to calculate. Computing was something a machine did rather than a line item you managed. You bought a computer, and computing was what happened inside it.

Then the cloud turned it into something you order. Once Amazon began selling processing capacity by the hour in 2006, compute became a unit you buy, like electricity. It still meant any processing power at all: your phone has compute, a weather model has compute, the payroll system has compute. Ordinary, and unambiguous.

Then AI narrowed it. AI training models got enormous, one company came to supply more than 90% of the world’s data center GPUs, and in a lot of business conversation “compute” now arrives meaning that specific hardware unless someone says otherwise. The older sense hasn’t gone anywhere; your cloud bill still meters ordinary compute by the hour. But the two now travel under one word. Same word. Two jobs. Plenty of people using it have only met one.

Then it became an asset. Not just metered capacity on someone else’s bill, but something firms buy outright, finance, depreciate and carry on a balance sheet. In August, Nvidia and six of the largest firms on Wall Street signed memorandums to mobilize more than $500 billion of outside capital around it. The agreements remain subject to execution.

That last shift gets explained least, and it’s the one that decides whether the money works.

What You’re Actually Buying: Chips, Power And CUDA

Take the list above and put weights on it. A GPU alone does very little: it needs fast memory beside it, networking to act as one machine with thousands of others, processors feeding it data, and storage underneath. Then the building, the power and the cooling. Fill a site with that and you have what Nvidia calls an AI factory: a data center built for one purpose, wrapped around a very expensive hardware core.

There’s also software, and it’s the part buyers underestimate. There’s a software layer too, called CUDA. Your engineers never touch the chips directly. Their code goes through CUDA to get there. That’s not a technical detail, it’s a switching cost. You aren’t only buying silicon, you’re buying into the toolchain your engineers already know and your code already assumes.

None of it ages at the same speed. The building stands for decades. The cooling runs for years. The chips are the question mark, and semiconductors and their related components run to roughly two-thirds of what an AI data center costs. It’s why counting transformers turned out to be a useful way to read this buildout.

A long-lived shell wrapped around a fast-moving core is the most important fact about compute as an asset, and nearly every argument about the AI buildout is downstream of it.

Training And Inference Are Two Different Clocks

Compute does two jobs with completely different shapes. Training is building the model, and any single run is bounded: enormous, expensive, and it ends. Inference is the trained model producing answers, and for a live service it doesn’t stop. Each answer is small, but multiply it across millions of users all day and it becomes a meter that keeps turning.

That explains where old chips go. Training wants peak performance. Inference will settle for less, and batch work for less still. So a chip that has aged out of the first job is often perfectly employable in the second.

How AI Compute Is Priced And Sold

There are three common ways to pay, and underneath the pricing each one answers the same question: who absorbs the risk that this machine is worth less next year than it is today?

Buy it. You own the asset and you own the problem.

Rent it by the hour. You pay a premium and hand the aging problem to somebody else.

Skip the hardware and pay per token through an API, a software connection that lets you buy model output without owning machines. You own none of the depreciation. Whether it costs more depends on volume: cheap when demand is spiky, expensive when it’s constant.

That trade rarely makes it into the pitch. The mistake I’ve watched companies make, more than once, is to buy their own hardware for the satisfaction of owning their AI. Then they run it at a fraction of what it could do, while it quietly loses value on the shelf.

The public numbers point the same way, though they aren’t all measuring the same thing. One 2026 model says owning only pulls ahead above roughly 70% sustained use, and renting wins below 30%. Where your own line sits depends on what you paid and how long you assume the hardware lasts. Against that, Cast AI’s telemetry across tens of thousands of enterprise clusters found average GPU use at 5% of what had been provisioned, measured across a full day. Then ask companies to estimate their own, and 53% put it between 51% and 70%.

Where instruments are attached, they read in single digits. Where people are asked, they answer in the fifties. That distance is the part worth worrying about.

Why Wall Street Wants To Finance $500 Billion Of AI Compute

The pitch is clean and it isn’t stupid. A data center full of GPUs throws off cash, and cash under contract can be lent against. I wrote about the mechanics when the platform was announced. The problem is the collateral.

A toll road still collects tolls in year forty. A leased plane trades in a market with decades of sales behind it, so lenders know what a used one fetches. Compute has neither. A secondary market only opened in July, and it posts estimates rather than completed sales. There’s no public benchmark for what a whole cluster fetches when the seller has no choice.

Bernie Margulies sells insurance against this hardware losing value, so he has every reason to talk the market up. Here’s what he told the equipment-finance trade press: “The disagreement is whether the right number is 10% or 60%.” Fifty points of spread on the same asset. That’s the difference between a sound loan and a hole.

Computing hardware has been here before. When the mainframe market turned, Margulies writes, it “repriced immediately,” and lessors who had booked aggressive residuals “went bankrupt within months.” His summary is six words: “Technology obsolescence is sudden, not gradual.”

The counter-argument is real. Old chips keep working and keep earning. CoreWeave told investors in August it had signed a contract for A100 capacity, a 2020 architecture, running into 2029. Silicon migrates down the ladder rather than dying.

How Much Of The Risk Is Nvidia Taking?

The most revealing number isn’t the $500 billion. Announcing the platforms, Jensen Huang wrote that “in some cases, NVIDIA may provide a residual-value support mechanism for up to 25% of an opportunity, assessed carefully on a project-by-project basis.” That sentence isn’t in the press release. TechCrunch and Bloomberg picked it up separately.

Read the qualifiers. May. Up to. Of an opportunity. Assessed one project at a time. The company that designs the chips, and decides when they’re replaced, has capped how much of that question it will answer for you.

So Should You Own, Rent, Or Buy By The Token?

Before signing anything, write down the useful life you’re assuming, the years you expect the hardware to earn, and run it a year shorter. Amazon cut the assumed life of some servers and networking gear from six years to five, adding $1.4 billion to one year’s depreciation. If your deal only works at six, you’ve found the soft spot before you stepped on it.

Then find out what your real use is, measured rather than estimated. If it sits below the break-even for owning, you’re not in the compute business. You’re in the storage business, and the thing you’re storing is losing value.

Compute is the raw material of this economy now. But raw materials usually trade in markets that discover their value in the open, and this one is priced on confidence that the machines keep earning long enough to pay for themselves.

Half a trillion dollars measures how badly the market wants that to be true. It isn’t yet a measure of whether AI compute will earn it.

Share.
Leave A Reply

Exit mobile version