How AI works

Is it predicting something, or making something?

Short answerPredictive AI scores something that already exists, and the real answer eventually arrives to check it against. Generative AI produces something new, where there is usually no answer to check. That one difference sets how each is tested, what it runs on, and where its return can be measured.

PredictiveGenerative
What it doesAssigns a number or a label to something that existsProduces new text, images, code, audio, or video
ExamplesFraud, credit risk, demand, machine failureDrafts, summaries, code, images
Ground truthArrives eventually, so accuracy is measurableOften absent, so accuracy is often undefined
HardwareModest, often an organization's own machinesFrontier models need hyperscale facilities

Ground truth is the whole difference

A fraud model flags a transaction. Weeks later the record shows whether it was fraud, and the model has been scored. That loop is what makes a predictive system testable: you can measure its accuracy, watch it drift, and compare it with whatever it replaced.

A generative model drafts a paragraph. No record will later say what the right paragraph was. "Accuracy" is frequently not even well defined, so evaluation becomes judgment: did a reviewer accept it, did it state anything false, did it break a rule.

Why the infrastructure differs

Predictive systems run on modest hardware and have been in production in banking and insurance for years, under existing regulation. Frontier generative models are what require hyperscale data centers, and most public argument about AI's energy and cost concerns them, usually without saying so. More in Does all AI need a hyperscale data center?

Where the measurable return sits

Most AI workloads running in production today are predictive: fraud detection, forecasting, underwriting, maintenance. That is where most of the measurable return sits. The systems generating the headlines and the systems generating the returns are, for now, largely different systems.

Where returns are measurable, they concentrate in predictive systems doing bounded, repetitive work with a ground truth to check against. Generative deployment is real and growing, and its returns are harder to attribute. That is a measurement problem, and it is no evidence of failure.

About "95% of AI pilots fail"The figure comes from one report, never peer reviewed, built on 52 executive interviews, which scored any pilot without measurable return within about six months as a failure. The definition does most of the work: a coding assistant that quietly saves engineering hours across a company is real value, and it would count as a failure. On that evidence the figure supports no conclusion. More in AI numbers to stop using.

What this means for an organization

  • Ask which kind a proposal is. A predictive project should arrive with a baseline and an accuracy target, because the answer will come in to check it. A generative one needs a different measure, such as time saved or the share of drafts accepted.
  • Expect different evidence from vendors. For a predictive system, ask for accuracy measured against outcomes. For a generative one, ask how outputs are reviewed and what gets recorded.
  • Give generative work a fair clock. Gains spread thinly across many people are real and slow to show up in a six-month payback.

Sources

  1. MIT Project NANDA, The GenAI Divide: State of AI in Business 2025. The source of the "95% of pilots fail" figure: 52 executive interviews, not peer reviewed, with a six-month definition of success.
  2. Codewave, The State of AI Enterprise Adoption in 2026. Production workload mix. A vendor compilation of other surveys, so the proportion is indicative.

Reviewed