Key Takeaways
- Training a model is a one-time capital bet with an uncertain payoff, while inference is a recurring, margin-driven business — collapsing the two hides where the real financial risk actually sits.
- Chip scarcity isn't just about fabrication capacity; power availability, high-bandwidth memory supply, and data center buildout timelines each impose their own multi-year constraints on the same supply chain.
- A capex-heavy foundation lab and a thin-margin API reseller both get called 'AI companies' in the same headline, but they carry entirely different cost structures, risk profiles, and paths to profitability.
Ask five people why AI is expensive and four of them will say some version of "GPUs are the bottleneck." It's not wrong. It's also not an explanation — it's the kind of thing people repeat because it sounds informed, the way "it's the economy" explains an election result without telling you anything you could actually act on. The real answer has layers, and once you separate them, a lot of confusing headlines about AI spending stop being confusing.
We think the confusion comes from treating "AI compute" as one expense line, when it's really at least three different businesses stacked on top of each other, each with its own economics, its own risk profile, and its own reason for being expensive. Knowing which one you're looking at is most of the battle.
Training Is a Bet, Inference Is a Utility Bill
Training a frontier model means spending an enormous amount up front, before anyone knows whether the result will justify it. That's R&D spending in the classic sense: most of the cost lands whether or not the outcome is good, and you only find out if it worked after the money is already gone. It's why people compare training runs to drilling an exploratory well — the cost structure looks like probabilistic research, not like manufacturing a known product with a known margin.
Inference is different in kind, not just in size. Once a model exists, serving it to users is a recurring, marginal-cost activity: every query consumes some amount of compute, and the total scales roughly with usage. That behaves much more like a utility — forecastable, priceable per unit, and improvable over time through optimization, smaller distilled variants, better serving infrastructure, and caching. A company can be losing enormous sums on the training side while running a perfectly sound, even profitable, business on the inference side of that exact same model. That isn't a contradiction. Training and inference are different economic activities that happen to share a product.
This distinction matters because a lot of public argument about whether AI "makes money" quietly conflates the two. Someone points to a company's overall losses as proof the business doesn't work, without separating how much of that loss is exploratory capital spending on the next model versus how the current model is actually performing against live traffic. Ask "unprofitable at what — the training bet, or the inference business built on top of the last bet that already landed?" and a surprising number of arguments answer themselves.
The Scarcity Isn't the Chip, It's Everything Around the Chip
People say "there aren't enough GPUs" as if the constraint were a single factory that just needs to make more units. The real constraint is a chain, and every link in it has its own lead time, and the tightest link sets the pace for everyone downstream of it. Leading-edge chip fabrication is concentrated among a small number of facilities capable of producing at the process nodes modern accelerators require, and building new fabrication capacity is measured in years, not quarters — you can't expand it the way you'd expand a warehouse by leasing another building.
Even once chips come off the line, they need to be packaged together with high-bandwidth memory, and that packaging and memory supply is its own separate capacity constraint, run by a different set of specialized suppliers with their own backlog. The finished accelerators then need servers, and the servers need racks and networking built to move enormous data volumes between chips efficiently, and the whole assembly needs a building with enough power and cooling to run continuously at extreme density. Grid capacity and the permitting that comes with it are themselves multi-year processes in most places, regardless of how much capital anyone is willing to throw at the problem.
Add money to any single link in that chain and you don't get proportionally more finished compute — you just move the bottleneck to whichever link is next. That's why "just spend more" doesn't fix this the way it would in a normal manufacturing scale-up. At every stage there's a different, mostly fixed near-term ceiling, and the compute available to the entire industry in a given year is set by whichever ceiling is lowest, not by how much capital anyone is willing to commit against it.
Two Businesses Wearing the Same 'AI Company' Badge
On one side of the industry are organizations that own or commit to massive, long-term compute infrastructure, train their own frontier models, and fund that spending largely with outside capital while the underlying business is still finding its footing. That looks and behaves like a capital-intensive industrial company — closer to a utility or an airline than to a typical software business, because the fixed costs are enormous relative to near-term revenue. On the other side are companies that build products on top of somebody else's model, paying for access by the token and charging their own customers a markup. That's a thin-margin distribution business, much closer to logistics or resale, where the core skill is packaging and go-to-market rather than owning infrastructure at all.
Both get called "AI companies" in the same breath and judged by the same headlines about spending and revenue, which produces bad analysis in both directions. Point at the capex lab's losses and call the whole industry unprofitable, and you're mistaking an infrastructure buildout for a failing business model — plenty of capital-intensive industries look terrible on a five-year view and perfectly fine on a fifteen-year one. Point at the reseller's thin margins and call it a weak business, and you're ignoring that thin margins on high volume with low fixed costs can be a perfectly healthy model — it's just a different one, with different rules for what success looks like. Neither company is lying about being "in AI." They're just occupying different parts of a stack that non-experts read as a single industry.
Idle Silicon Is the Real Enemy
There's a second-order cost that rarely comes up: once you've secured compute, keeping it busy is its own problem, and an idle accelerator isn't a neutral expense. It's actively destructive, because you're paying depreciation, power, and opportunity cost on hardware that's doing nothing productive while the clock runs. That creates a strong incentive to commit to long-term usage contracts and pre-purchase capacity years in advance — not necessarily because demand is proven that far out, but because the alternative is worse: showing up to buy in a shortage, with no guaranteed supply, at whatever price happens to be available then.
That's why some of the largest reported infrastructure commitments in this industry look less like confident demand forecasts and more like insurance policies against a scarce resource getting scarcer. Locking in a seat at the table is a hedge, not necessarily a prediction that every seat locked in will be filled on schedule. It also explains why public commentary about compute so often sounds contradictory: a company can genuinely be short on capacity for its highest-priority workloads while simultaneously sitting on committed capacity elsewhere that isn't yet fully utilized. Both things are true at once, because utilization is a scheduling and allocation problem layered on top of a raw supply problem, and getting the two aligned takes real time even after the hardware physically exists and is plugged in.
Cheaper Per Token Doesn't Mean Cheaper Overall
Every year, the cost of running a given amount of inference tends to drop — smaller models get more capable, serving infrastructure gets more efficient, hardware improves per dollar spent. The intuitive assumption is that this should make AI cheaper in aggregate over time. It usually doesn't, and the reason is worth sitting with: when the cost per unit of something drops, people rarely hold usage constant and pocket the savings. They use more of it. Cheaper inference doesn't mean the same tasks get done for less money; it means longer context windows become normal, multi-step workflows that call a model dozens of times to complete one task become economically viable, and use cases that were previously too expensive to justify get built in the first place.
That's a very old pattern in industrial economics, and it's worth knowing by name because it predicts something counterintuitive: the more efficient AI compute gets, the larger the aggregate compute market tends to grow, not shrink — at least for as long as there's unmet appetite for doing more of a given task rather than a fixed amount of it to do more cheaply. If you're waiting for AI to get so efficient that total industry compute spending starts falling on its own, efficiency alone probably won't get you there. It will just change what a dollar buys. It won't change how many dollars get spent chasing the next thing that efficiency just made viable.
None of this makes AI compute cheap, and none of it should. But it does mean the question "why is this so expensive" has the wrong shape when asked as if there's one answer. It's not one expense — it's a stack of different expenses with different logic: a capital bet, a utility bill, a supply chain with a dozen separate bottlenecks, a scheduling problem, and an efficiency trap that grows demand instead of shrinking it. The next time someone tells you GPUs are the bottleneck, ask which part of the stack they actually mean. If they can't answer, they're repeating a headline, not explaining an economy — and the entire point of understanding this stuff is that you stop having to take anyone's word for it.