Control and its limits

What can and cannot be guaranteed about an AI system?

Short answerYou can guarantee the space of allowed actions: a system can be built so it is capable only of what a defined set permits. You can never prove that a perception is correct. That holds for every system that acts on perception, people included.

These are two kinds of claim. The first is about a system's permissions, and permissions are definable, bounded, and checkable. A system can be built so that it is capable only of what a defined set allows and incapable of everything else.

The second is about the world: whether an image is real, whether a document is what it claims to be, whether a situation is what it appears to be. These are judgments about external reality, made from incomplete information. No amount of engineering converts them into proofs.

Left: the space of allowed actions is a closed set of operations such as read record, draft reply, summarize, and flag for review; an attempt to send funds is refused at the boundary. Right: a perception maps an open world of images, documents, and unseen situations onto a claim such as "this document is authentic," and no boundary can be drawn around it.
The permitted actions can be listed, so a boundary exists and an attempt can be refused at it. The world cannot be listed, so there is nothing to draw a boundary around.

This is a structural feature of every system that acts on perception, including human beings. A better model, more data, or more money changes how often a perception is right. None of them makes it provable.

Two consequences

  • Requirements that depend on perfect perception cannot be met. "Detect every deepfake" and "never misclassify" are requirements to solve an unsolved problem. A policy, contract, or vendor promise built on one is a commitment nobody can keep.
  • Verification lives at runtime. Action can be bounded and behavior can be recorded, so observing a system while it runs is where the available verification is. It is the part of assurance anyone can check.

The clearest example: deepfakes

Deciding whether a recording is genuine is a perception problem, so it inherits the limit. Three approaches are in use, and each has its own ceiling:

  • Detection models trained to spot synthetic media generalize poorly to generation methods they were not trained on. Reported performance falls sharply outside test conditions.
  • Provenance metadata, credentials attached when content is created, is usually stripped in transit. Recompression, screenshots, and re-sharing remove it, so the content reaching the widest audience is the least likely to carry any.
  • Watermarking is probabilistic. It produces false positives and false negatives, and a watermark that can be detected can eventually be forged.

What can be enforced sits on the provable side:

  • Bound the action. Where content is generated or distributed inside a governed system, the system can be prevented from producing or passing certain categories of content at all.
  • Observe the output. What a deployed system produced is a matter of record, if something is recording it.
  • Assign accountability. Obligations can attach to the platform, the model provider, or the person who prompted it. That is a policy choice, and it holds whether or not detection works.

The TAKE IT DOWN Act, signed in May 2025, is an enacted example. It criminalizes non-consensual intimate imagery, including synthetic material, and requires platforms to remove it within 48 hours. It places a duty on a named party and asks nobody to detect anything perfectly.

Sources

  1. Research on deepfake detection generalization, C2PA provenance metadata survival, and watermark robustness.
  2. TAKE IT DOWN Act, signed May 2025. Criminalizes non-consensual intimate imagery, including synthetic material, and requires platforms to remove it within 48 hours.

Reviewed