Home / Several agents and adversarial checking
Succeeded, failed, interrupted, no trace
When a cutoff stops a queue of tasks, a log kept outside the process, written before and after each task, separates four states: succeeded, failed, interrupted midway, and no trace of an attempt; confusing them leads you to rerun what may already have acted, or to declare failed what nobody has examined.
Take a typical case. A chain of checks processes a queue of tasks one after another. A usage cap hits in the middle of the queue and stops the supervisor process along with the agents it was driving. On restart, the supervisor marks as failed every task with no result, including those that no agent had started.
Four states, because a cutoff creates two
A system that only knows succeeded and failed files under failed two cases that the cutoff produces. The interrupted task has started: it may have acted in part, and it gets inspected before being rerun. The task with no trace is rerun, once it has been checked that no other process picked it up. Writing the "running" line before starting is what separates them: without it, an interrupted task looks like a task never touched.
The absence of a line establishes one bounded thing, as the lesson on what a rigorous check does not see reminds us: this log carries no trace of an attempt, and says nothing about another process that may have taken the task.
The demonstration
The two scripts are created in a throwaway folder. The first writes six tasks into queue.json, then simulates a cutoff during the fourth. The second sets this starting list against the log: it is the list that reveals the tasks with no trace.
mkdir -p demo-queue && cd demo-queue
cat > process-queue.js << 'EOF'
// Builds the starting queue, then processes it, logging before and after each task.
const fs = require('fs');
const queue = ['task-1', 'task-2', 'task-3', 'task-4', 'task-5', 'task-6'];
fs.writeFileSync('queue.json', JSON.stringify(queue));
fs.writeFileSync('log.jsonl', '');
const results = { 'task-1': 'succeeded', 'task-2': 'failed', 'task-3': 'succeeded' };
for (const id of queue) {
fs.appendFileSync('log.jsonl', JSON.stringify({ id, state: 'running' }) + '\n');
if (id === 'task-4') process.exit(1); // simulated cutoff during the fourth task
fs.appendFileSync('log.jsonl', JSON.stringify({ id, state: results[id] }) + '\n');
}
EOF
cat > read-log.js << 'EOF'
// Rereads the starting queue and the log, and sorts each task into four states.
const fs = require('fs');
const queue = JSON.parse(fs.readFileSync('queue.json', 'utf8'));
const last = {};
for (const line of fs.readFileSync('log.jsonl', 'utf8').split('\n')) {
if (line) { const e = JSON.parse(line); last[e.id] = e.state; }
}
const classes = { succeeded: [], failed: [], interrupted: [], 'no trace': [] };
for (const id of queue) {
const state = last[id];
if (state === undefined) classes['no trace'].push(id);
else if (state === 'running') classes.interrupted.push(id);
else classes[state].push(id);
}
for (const [name, ids] of Object.entries(classes)) console.log(name + ': ' + (ids.join(', ') || '-'));
EOF
node process-queue.js
echo "exit code: $?"
cat log.jsonl
node read-log.js
exit code: 1
{"id":"task-1","state":"running"}
{"id":"task-1","state":"succeeded"}
{"id":"task-2","state":"running"}
{"id":"task-2","state":"failed"}
{"id":"task-3","state":"running"}
{"id":"task-3","state":"succeeded"}
{"id":"task-4","state":"running"}
succeeded: task-1, task-3
failed: task-2
interrupted: task-4
no trace: task-5, task-6
Exit code 1 signals the abnormal stop. Task 4 carries "running" with no result: interrupted, not failed.
A trace that outlives the process
A tally kept in the supervisor's memory disappears with it. The log therefore lives outside the process: a file as here, a database table or an external message queue would also do. appendFileSync finishes writing the line before moving on: if the process is killed, the lines already written stay in the file. A power cut is another case, because the system may hold the last lines in memory before putting them on disk; an explicit sync, fs.fsyncSync in Node, forces that write.
The lesson on the three failure states separates, among the tasks that returned a result, the recoverable failure that will be retried and the final failure that goes up to a person; this one deals with the tasks the cutoff left without a result.
What the log records, task by task
Four states, four actions
| State | Trace in the log | Next action |
|---|---|---|
| Succeeded | "running" then "succeeded" | None |
| Failed | "running" then "failed" | Handle according to the type of failure |
| Interrupted | "running" and nothing after | Inspect the partial effects, then rerun |
| No trace | No line, task present in the starting list | Check that no other process picked it up, then rerun |
A process runs through a queue of six tasks listed in a starting file. It writes to a log on disk a "running" line before each task and a result line after. A usage cap stops the process. Reread after the stop, the log contains for task 1 "running" then "succeeded", for task 2 "running" then "failed", for task 3 "running" then "succeeded", and for task 4 "running". Tasks 5 and 6 appear in the starting file.
Write in one sentence what this situation establishes, and in one sentence what it does not establish.
What this establishes: It establishes that tasks 1 and 3 succeeded, that task 2 failed, that task 4 started and has no recorded result, and that the log carries no trace of an attempt for tasks 5 and 6.
What this does not establish: It does not establish that task 4 had no effect, since it may have acted in part before the stop, nor that no other process handled tasks 5 and 6: the log only describes this process.
The three most common miscalibrations
- Too broad Tasks 4, 5 and 6 failed because of the usage cap and all three are rerun in the same way.
- Too narrow The log establishes the state of tasks 1 to 3; for tasks 4, 5 and 6, it lets you distinguish nothing.
- Beside the point The situation establishes that the usage cap was set too low for the size of this queue.
- An interrupted task has started and may have acted in part, so it gets inspected before being rerun.
- A task with no trace is recognised by setting the starting list against the log, not by the log alone.
- The line written before starting is what separates an interrupted task from a task the processing never reached.
- A useful log lives outside the process that writes it: file, database or external message queue.
Take a script that processes a list of items, make it write a "running" line before each item and a result line after, stop it by hand in the middle, then check that you find the four states by setting its starting list against the log.