How AI works

What is the difference between training and inference?

Short answerTraining finds a model's weights, once, over weeks or months. Inference uses those weights to answer a request, every time anyone asks, for as long as the model is deployed. Over a model's life, inference is the larger and the permanent load.

TrainingInference
What it doesFinds the weightsUses them to answer
How oftenOnce, over weeks or monthsEvery request, continuously
HardwareTens of thousands of processors, tightly coupledSpread out, running around the clock
Cost shapeA large one-time capital costAn ongoing operating cost

Training is the process of finding a model's weights from data. For a frontier model it means tens of thousands of processors working as one machine for months. That is the figure usually quoted, and it is a real one. Then it ends.

Inference is using the finished model. Each answer is cheap, and the answers never stop: every request, every day, for as long as the model is in service.

Power drawn over time. Training is a tall block that ends after a few months. Inference is a lower band that starts at deployment and continues, so its total area passes training's and keeps growing.
Training is tall and finishes. Inference is lower and keeps going. The claim is the area under each curve, which is why an accurate training figure can still describe the wrong load.

Why this decides data-center demand

Google metered its own machine-learning energy at about three-fifths inference to two-fifths training, holding across three consecutive years, with machine learning at 10 to 15% of the company's total energy in each of them. Measured.

The exact ratio is contested and moving. The shape is settled. Training is bounded and ends; inference begins at deployment and continues. So demand grows with use, and what gets run, and how often, stays a lever for as long as a system is deployed.

Two questions get collapsed here. How large a facility has to be is set by training, a single job that needs a great many machines at once. How much that facility draws across its life is set by use. An argument that answers one while claiming to settle the other is the most common error in this debate.

The same rack doing two jobs, drawn at true relative scale. About 1,000 racks, roughly 120 megawatts, train one frontier-class model over two to three months. One rack serves about 10,000 people at ten-second responses, continuously.
Same hardware, two jobs. The building is sized by the training run and filled by use. Rack power and GPU counts are published by the manufacturer; rack counts and the people one rack serves are our arithmetic, labeled on the figure.

What a query actually costs

Per-query energy is smaller than most headlines suggest. A 2026 peer-reviewed study, using production serving assumptions, puts a frontier-model query at a median of 0.31 watt-hours and finds widely cited estimates overstate it by 4 to 20 times. The same work finds long reasoning queries raise that median roughly thirteenfold, so the pressure runs both ways at once.

A number to stop usingYou will see it said that inference is 80 to 90% of AI energy. We could not find a primary source for it. It is not in the International Energy Agency's Energy and AI report, and it circulates through trade press citing itself. Google's metered three-fifths is the figure that holds. More in AI numbers to stop using.

What this means for an organization using AI

Because inference is the permanent cost, the choices that matter most are made after deployment: which model handles which task, how often it is called, and whether a smaller model would do. Routing each task to the smallest model that does it well is one of the largest levers on both cost and energy.

Sources

  1. Patterson, Gonzalez, Hölzle, et al., The Carbon Footprint of Machine Learning Training Will Plateau, Then Shrink, IEEE Computer, 2022.
  2. Oviedo, Kazhamiaka, Choukse, et al., Energy Use of AI Inference, Efficiency Pathways, and Test-Time Scaling, Joule, 2026.
  3. International Energy Agency, Energy and AI, April 2025. Data centers at about 415 TWh in 2024, roughly 1.5% of global electricity, projected to about 945 TWh by 2030.
  4. NVIDIA, GB200 and GB300 NVL72 documentation. About 120 kW per 72-GPU rack.

Reviewed