Home / Several agents and adversarial checking
An unverified premise in a prompt comes back as a conclusion
The framing text an orchestrator writes into each agent's prompt stays invisible to adversarial review, which receives the produced artefact and the success criteria, never the intent or the history of whoever built the task: a false claim slipped into that framing therefore travels through the whole chain untested, and a fan-out that sends the same framing to several agents multiplies this exposure instead of diluting it.
An adversarial verifier receives the artefact to be judged and the success criteria it must test, never the history of the task nor the reasoning of the person who built it, according to the rationale Anthropic documents for this role. This deliberate limit, which keeps the verifier focused on what is in front of it, leaves one precise blind spot: it says nothing about what the orchestrator had already assumed to be true before even writing the first prompt.
The framing travels without ever being tested
The introductory text of a prompt often states a factual claim presented as already established rather than as a conclusion to be checked, for example that supplied data is already clean or that a scope has already been validated by a third party. The agent that receives this prompt treats that sentence as a premise, not as a fact to challenge, and reflects it back in its own output without ever restating it as a hypothesis. The adversarial verifier, for its part, only reviews that output: the original premise appears nowhere in a form it can put to the test.
Prompt sent to each agent in the fan-out:
"The catalogue provided contains only active references.
Write the product sheet for each of the listed references."
The adversarial verifier receives only the sheets produced,
never the sentence about active references.
The fan-out multiplies the exposure instead of diluting it
A fan-out sends the same framing text to every agent launched in parallel. A false premise slipped into that text is therefore not diluted by the number of agents, it is repeated in each of their outputs. Several independent agents that separately arrive at the same claim can then give the impression of corroboration, when in fact they all share the same source, the same starting prompt, and the same blind spot. Assigning a reviewer to specifically challenge the framing, not just the conclusions built on top of it, gives adversarial review a grip on this sentence that it normally never sees. The previous lesson shows what an adversarial verifier receives to judge a piece of work; here, the problem comes precisely from what it never receives.
The invisible framing becomes a fact cited in the deliverable
Sentence in the prompt: the catalogue provided contains only active references. Never tested, never shown to the verifier.
Product sheet: active reference, in stock. Cited as an established fact, without ever being verified.
A project manager writes, at the top of the prompt sent to six agents, that the customer database has already been cleaned of duplicates, then launches the six agents to each write a summary sheet from this database.
Write in one sentence what this situation establishes, and in one sentence what it does not establish.
What this establishes: The situation establishes that the claim about the duplicate cleanup was written into the prompt's framing and not produced as a conclusion verified by any of the six agents.
What this does not establish: It does not establish that the database actually contains duplicates or that it is genuinely free of them, no check of this fact took place in the situation described.
The three most common miscalibrations
- Too broad This situation proves that the six summary sheets produced will necessarily contain errors linked to undetected duplicates.
- Too narrow This situation allows no claim at all about the customer database, since writing a sentence in a prompt has no bearing on the actual state of a database.
- Off target This situation shows that the team has six agents reliable enough to write customer summary sheets autonomously.
- An adversarial verifier receives the produced artefact and the success criteria to test, never the history or intentions of the person who built the task, which leaves the prompt's framing outside its reach.
- A claim stated as a premise in a prompt's text generally does not appear in the final artefact in a form the verifier can challenge, so it passes through the review untested.
- A fan-out that sends the same framing text to several agents does not dilute a false premise, it reproduces it in every output, which can create a false impression of corroboration between independent agents.
- Rereading your own prompt before launching a fan-out lets you spot a factual claim slipped into the framing, before it becomes a conclusion cited as established in the deliverable.
Before your next fan-out, reread your own prompt and underline every factual claim it contains, as if it were already validated data, then write for each one the check that would test it before launching the agents.
These points depend on an interface or a rule that may have changed since this was written. Check them on your own screen before relying on them.
- If you already use an adversarial verifier, check its current prompt to see whether it receives the framing text sent to the agents or only their outputs: this difference determines whether it can spot a false premise.
Every datable claim in this lesson links here to the public text behind it. A source that does not open proves nothing.
- Anthropic, building multi-agent systems, when and how to use them consultée le 2026-01-23