Control and its limits
What can and cannot be guaranteed about an AI system?
Short answerYou can guarantee the space of allowed actions: a system can be built so it is capable only of what a defined set permits. You can never prove that a perception is correct. That holds for every system that acts on perception, people included.
These are two kinds of claim. The first is about a system's permissions, and permissions are definable, bounded, and checkable. A system can be built so that it is capable only of what a defined set allows and incapable of everything else.
The second is about the world: whether an image is real, whether a document is what it claims to be, whether a situation is what it appears to be. These are judgments about external reality, made from incomplete information. No amount of engineering converts them into proofs.
This is a structural feature of every system that acts on perception, including human beings. A better model, more data, or more money changes how often a perception is right. None of them makes it provable.
Two consequences
- Requirements that depend on perfect perception cannot be met. "Detect every deepfake" and "never misclassify" are requirements to solve an unsolved problem. A policy, contract, or vendor promise built on one is a commitment nobody can keep.
- Verification lives at runtime. Action can be bounded and behavior can be recorded, so observing a system while it runs is where the available verification is. It is the part of assurance anyone can check.
The clearest example: deepfakes
Deciding whether a recording is genuine is a perception problem, so it inherits the limit. Three approaches are in use, and each has its own ceiling:
- Detection models trained to spot synthetic media generalize poorly to generation methods they were not trained on. Reported performance falls sharply outside test conditions.
- Provenance metadata, credentials attached when content is created, is usually stripped in transit. Recompression, screenshots, and re-sharing remove it, so the content reaching the widest audience is the least likely to carry any.
- Watermarking is probabilistic. It produces false positives and false negatives, and a watermark that can be detected can eventually be forged.
What can be enforced sits on the provable side:
- Bound the action. Where content is generated or distributed inside a governed system, the system can be prevented from producing or passing certain categories of content at all.
- Observe the output. What a deployed system produced is a matter of record, if something is recording it.
- Assign accountability. Obligations can attach to the platform, the model provider, or the person who prompted it. That is a policy choice, and it holds whether or not detection works.
The TAKE IT DOWN Act, signed in May 2025, is an enacted example. It criminalizes non-consensual intimate imagery, including synthetic material, and requires platforms to remove it within 48 hours. It places a duty on a named party and asks nobody to detect anything perfectly.
Sources
- Research on deepfake detection generalization, C2PA provenance metadata survival, and watermark robustness.
- TAKE IT DOWN Act, signed May 2025. Criminalizes non-consensual intimate imagery, including synthetic material, and requires platforms to remove it within 48 hours.
Reviewed