Home / Checking what it gives back
Three agents agree, one single source
Several agents that return the same conclusion while citing the same page confirm nothing between the three of them: agreement is counted in independent sources, and a different address is not yet a different source.
Take a typical case. Three agents questioned separately on the same question return the same conclusion, and the team that launched them sees three verifications converging. If all three cite the same page as their basis, the team holds one source, cited three times: agreement between opinions is worth what the distinct sources behind it are worth.
What a text deduplication counts
The first step is to collect the address each agent says it consulted, then count what remains once the duplicates are removed. This count has a limit that needs naming: sort -u removes lines that are identical character for character, so it counts strings of text. Two spellings of the same page, with and without an anchor after the # sign, pass for two; a mirror that copies the page under another domain does too.
The following demonstration creates its own folder and files, then counts in three ways:
mkdir -p demo-agreement && cd demo-agreement
printf 'conclusion: option A compliant\nsource: https://example-doc.test/limits\n' > opinion-1.txt
printf 'conclusion: option A compliant\nsource: https://example-doc.test/limits#quota\n' > opinion-2.txt
printf 'conclusion: option A compliant\nsource: https://mirror-doc.test/copy/limits\n' > opinion-3.txt
echo "agreeing opinions: $(grep -l 'option A compliant' opinion-*.txt | wc -l | tr -d ' ')"
echo "distinct address strings: $(grep -h '^source:' opinion-*.txt | sort -u | wc -l | tr -d ' ')"
echo "distinct addresses without anchor: $(grep -h '^source:' opinion-*.txt | sed 's/#.*//' | sort -u | wc -l | tr -d ' ')"
printf 'The quota resets to zero every minute.\n' > extracted-text-example-doc.txt
printf 'The quota resets to zero every minute.\n' > extracted-text-mirror-doc.txt
cksum extracted-text-example-doc.txt extracted-text-mirror-doc.txt
agreeing opinions: 3
distinct address strings: 3
distinct addresses without anchor: 2
2178574758 39 extracted-text-example-doc.txt
2178574758 39 extracted-text-mirror-doc.txt
Three opinions, three strings, two pages once the anchor is removed. The tr -d ' ' removes the spaces that wc -l puts in front of its result on macOS, so that the output is the same on every system. The last two lines come from cksum, which computes a checksum, a number derived from the contents of a file: two files identical byte for byte return the same one.
Normalise, then trace back to the origin
Normalise first: remove the anchor, the tracking parameters and the trailing slash, so that two spellings of the same page merge. The anchor is the clearest case, since the browser does not send it to the server: two addresses that differ only by it point to the same resource.
Then trace back to the origin. Two domains can carry the same text, copied or syndicated. Two real copies of a web page nonetheless differ in their template, menus or ads, so a checksum computed on the raw pages returns two different numbers. It is only useful on the extracted and cleaned text, as in the two files of the demonstration; failing that, you compare the quoted passages, or read the attribution that the copy carries. An identical checksum signals a copy, and two different checksums do not prove two independent investigations, since one page can reword another.
The lesson on counting the whole population asks you to name a measure for what it captures: here, deduplication captures strings, normalisation captures pages, and comparing the extracted text approaches the origin. Fan-out search, in track 6, takes the problem up from the collection side: a claim is only cited there after it has withstood a real attempt to refute it.
Three measures on the same three opinions
| Measure | What it captures | Result |
|---|---|---|
| sort -u on the addresses | Strings of text identical character for character | 3 |
| Anchor removed, then sort -u | Pages, whatever their anchor | 2 |
| cksum on the extracted text | Cleaned texts identical byte for byte, whatever domain publishes them | 1 |
From the number of voices to the number of origins
A team gives the same question to three checking agents. Each returns a favourable conclusion and attaches the address it says it consulted. The three attached addresses are https://doc.exemple.test/limites, https://doc.exemple.test/limites#quota and https://doc.exemple.test/limites#calcul.
Write in one sentence what this situation establishes, and in one sentence what it does not establish.
What this establishes: It establishes that the three addresses differ only by their anchor and therefore point to the same page, so the three opinions rest on one cited source at most.
What this does not establish: It does not establish that the conclusion is wrong, nor that the agents read this page and drew their conclusion from it: the situation reports what they cited, not what they did.
The three most common miscalibrations
- Too broad Three agents citing the same page confirm that this page is accurate, since three separate readings found the same thing in it.
- Too narrow The situation establishes that the three addresses start the same way; knowing whether they point to one page or several would require opening them.
- Beside the point The situation establishes that the three agents received the same search instruction, which led them to the same page.
- Agreement between several opinions is measured against the number of independent sources that support it, not the number of agents that returned it.
- A sort -u deduplication counts distinct strings of text, so two spellings of the same page count double there.
- Removing the anchor and parameters from an address brings the count closer to the number of pages actually cited.
- A copy published under another domain is spotted on the extracted text of the pages or on its attribution, because two raw pages differ in their template even when they carry the same text.
Take the latest set of opinions returned by several agents on one question, collect the address each one cites, remove anchors and parameters, open the remaining pages, and note how many carry a distinct origin text before concluding there is agreement.