AI systems • 004 | 7 October 2026 | Provenance

A signal is not a verdict.

What does an AI watermark actually prove?

Maddipalli Gopalakrishna · AI / ML Engineer

45 seconds · music only · examples on screen are illustrative

AI-generated text is getting invisible watermarks. But a watermark is not a truth detector.

What OpenAI announced

On 5 October 2026 OpenAI described textGrain, a text watermark that adds “an invisible statistical signal to the model’s word choices”, which a detector then looks for. It is opt-in through the API for select models, off by default, and is coming to ChatGPT and Codex output in the EU over the next few weeks. Detector access starts with approved researchers and expert organizations rather than the public, and OpenAI says it plans to release the technology as open source.

I'm less interested in the product than in the abstraction people will build on top of it. OpenAI is unusually direct about what a detection can and can't tell you:

OpenAI says a watermark…Because
Does not verify accuracyIt doesn't tell you whether a passage is true, misleading or harmful.
Absence is not proof of human authorshipText may be too short, edited or translated, come from an unsupported model, predate watermarking, or come from another company's tools.
Does not measure human contributionIt can indicate an OpenAI system generated or processed part of a passage, not how much human judgment went into it.
Does not identify the user or establish ownershipNo person, account, prompt or conversation is associated with the text.

Four different questions

These get blurred together in conversations about AI detection. They aren't the same question.

Provenance
Where did this content come from? A watermark is one signal that can speak to this.
Authenticity
Has the content stayed intact, or been modified? Knowing the origin doesn't answer this.
Attribution
Which system or actor produced it? Provenance signals can contribute, but don't prove every attribution claim.
Truthfulness
Is the information correct? That needs separate verification.

A detected watermark signal speaks to provenance. It leaves authenticity, attribution and truthfulness unresolved. Two documents can carry the same signal and differ completely on factual accuracy.

The dangerous shortcut

watermark detected  →  AI
no watermark        →  human

That turns a probabilistic signal into an oracle. The second line is the riskier one: a document with no detected signal is better described as origin: insufficient evidence than as human-written, for exactly the reasons OpenAI lists. A signal is not a verdict.

Detection as evidence

A sturdier design treats the watermark as one input. Combine it with metadata, signed credentials and generation records where they exist (they won't always), keep the result as evidence rather than a label, and let a policy decide what happens next: automation for low-impact, high-confidence cases and human review when the decision is high-impact or the evidence is uncertain. That isn't a universal policy; it's a way to keep a human in the loop where a wrong call is costly.

Content is ingested, then a watermark detector, metadata analysis and signature or credential check feed an evidence store and a provenance engine. Low-risk cases go to automation, high-impact or uncertain cases to human review, and everything lands in an audit log.
Conceptual architecture · not a product design

Store the context around a detection, not just its outcome. A conceptual record might hold:

content_id, content_hash
detector, detector_version
signal            signal_detected | no_signal_detected | error
confidence        (as reported, and how it is defined)
transformations   none | edited | translated | unknown
supporting_metadata, credentials
policy_version, decision, review_status

This is my own sketch, not an OpenAI schema. The point is that detectors, thresholds, policies and models change, so a stored result has to stay interpretable later. Never keep only “AI = true”.

Evaluate the detector

Treat the detector like any other production component and test it. OpenAI's own figures show why: in its evaluations on 400-token English passages, replacing 10% of words with synonyms reduced detection from about 92% to 66%, and replacing 25% reduced it to 17%. Its technical report is also careful that a theoretical false-positive guarantee holds under stated assumptions and that a deployed key needs empirical calibration checks.

DimensionCases
OutcomesTrue positives, false positives, true negatives, false negatives.
LengthShort and long text. OpenAI reports detection is lower for shorter and more constrained text.
TransformationsCopy and paste, light editing, paraphrasing, translation, partial extraction, formatting changes, mixed human and AI editing.
VariationDomain, model and decoding settings, where they are meaningful for the detector.
MetricsPrecision, recall, false-positive rate, false-negative rate, calibration, and robustness per transformation.

Report results per condition rather than as one headline number, and re-run them whenever the detector, key or threshold changes.

AI provenance should be engineered as an evidence system, not a binary AI detector.

How would you design provenance checks if the decision could affect a student, employee, customer or publication?