How AI works

Does all AI need a hyperscale data center?

Short answerNo. Training frontier models does. Most deployed AI runs on ordinary hardware, and much of what runs on frontier models does not need to: one study estimates that routing each task to the smallest model that does it well would cut global AI energy use by 27.8%.

Predictive systems such as fraud detection, demand forecasting, and diagnostic risk scoring run on modest infrastructure, often on an organization's own machines. They have been in production for years without hyperscale anything.

Even within generative AI, matching the model to the task is a large lever. Many production tasks do not need the largest available model. They need a smaller one, tuned, which is cheaper, faster, and can run privately.

The smallest model that does the job

A 2025 study estimated that routing each task to the smallest model that does it well could reduce global AI energy consumption by 27.8%, about 31.9 TWh in 2025, roughly the annual output of five nuclear reactors. Savings by task ranged from 1% to 98%. Modeled, by the study's authors from stated assumptions.

So model choice is a cost decision and a privacy decision at once. A smaller model running where the organization controls it also reduces how much data leaves the organization at all.

Why size stopped predicting cost

Three techniques have broken the link between how large a model is and what it takes to run:

  • Lower precision. The same model stored in fewer bits per number. At 16-bit precision a model takes about two bytes per parameter, so 1.5 trillion parameters is roughly three terabytes; at four bits it is a quarter of that, on a fraction of the hardware. A rule of thumb, which moves with the format and the implementation.
  • Adapters. Leave the model alone and train a small layer beside it. Orders of magnitude less to train and to store, and adapters for different jobs can be swapped onto one base model.
  • Sparse activation. A model can be enormous and use a slice of itself on any given request. One with well over a trillion parameters may put only tens of billions to work on each token.

One limit. These reduce what it costs to run a model. They leave untouched what it cost to train it.

Why a compute threshold needs maintenance

Compute is how several regimes decide which models carry the heaviest obligations. The EU AI Act presumes a general-purpose model carries systemic risk once its cumulative training compute passes 10²⁵ floating-point operations, and the same article lets the European Commission amend that figure for algorithmic improvements or more efficient hardware.

That is sound design, with a consequence: a proxy that needs regular maintenance is still a proxy, and someone has to own maintaining it. Efficiency gains move the line every year without any model getting smaller.

What this means for an organization

  • Test the smaller model first. Before defaulting to the largest model for every task, check whether a smaller one does the job. A short list of approved models matched to task types costs less than one model for everything.
  • Read size claims carefully. A parameter count says little about running cost. Ask about precision, adapters, and how much of the model is active per request.
  • Check thresholds against the current text. If the EU AI Act's general-purpose model rules reach a model you use or supply, confirm the threshold in force. The Commission can change it.

Sources

  1. da Silva Barros, Giroire, Aparicio-Pardo, and Moulierac, Small is Sufficient: Reducing the World AI Energy Consumption Through Model Selection, arXiv:2510.01889, October 2025.
  2. European Union, Artificial Intelligence Act, provisions on general-purpose AI models with systemic risk: the 10²⁵ floating-point operation presumption and the Commission's power to amend it.

Reviewed