Read a saved agent trace and say what the run actually did -- including when it did not finish.
- Repository: edilec/agent-loop-trace-reader
- Area: Prompt & Agent Workflows
- License: MIT
Given a saved trace, it produces a report and a structured summary covering:
- steps, with the calls and waits belonging to each, and whether each one ended;
- tool calls, each with its id, its step, a digest of its arguments, its
duration and its result --
ok,error, ornullfor a call the trace never answered; - retries, linked to the call they retry;
- waits, totalled and attributed to the step that waited;
- repeated identical calls, grouped by tool and argument digest, with every call id preserved;
- loops, both the same call repeated and two consecutive steps doing identical work;
- budget consumption, compared with the budget the trace declares;
- the terminal outcome, and whether there was one at all.
Nothing is fetched, no model is called, and the only clock is the one the command line injects for this tool's own analysis budget.
A trace is read after something went wrong, which is the worst moment to be given a reassuring answer. Three properties follow from that:
- An unanswered call is not a successful one. A
tool-callwith notool-resultmeans the trace holds the question and not the answer. It is reported asnull, never asok, and it makes the runincomplete. - An interrupted trace stays incomplete. An outcome of
interrupted, no outcome at all, a step that never ends, or events after the outcome: each one means the run has no result to report, so this tool abstains. Exit2, never exit0. - Grouping never loses an id. A group that collapsed its members into a
count would tell you something repeated and not which calls to go and look
at. Every group carries every call id, in order, at its full declared length,
and every call also remains listed individually. The one place an id list is
bounded is a finding's
evidencefield, which the report contract bounds -- so it names whole ids, says how many it left out, and points at the artifact that carries them all.
And one property that is about the trace rather than the run: argument values are never read out. Calls are compared by a SHA-256 digest of their canonical form. A trace is the most likely place in a pipeline for a credential to be sitting inside a tool argument, and grouping identical calls is exactly the operation that tempts a tool to print them.
# A run that finished cleanly
node bin/agent-loop-trace-reader.mjs --trace examples/completed-trace.json --human
echo "exit: $?" # 0
# A run that looped and overran its budget -- a verdict, so exit 1
node bin/agent-loop-trace-reader.mjs --trace examples/looping-trace.json --human
echo "exit: $?" # 1
# A run that was interrupted -- no verdict is possible, so exit 2
node bin/agent-loop-trace-reader.mjs --trace examples/interrupted-trace.json --human
echo "exit: $?" # 2
# The machine-readable report, which is what stdout carries by default
node bin/agent-loop-trace-reader.mjs --trace examples/looping-trace.json | jq .summary
# The structured summary, written whatever the status
node bin/agent-loop-trace-reader.mjs \
--trace examples/interrupted-trace.json \
--summary build/interrupted-summary.json
npm run check # lint, tests, all three examples, and a packaging dry run{
"schemaVersion": "1",
"trace": {
"id": "invoice-export-2026-09-14",
"budget": { "maxSteps": 8, "maxToolCalls": 20, "maxWaitMillis": 30000, "maxTokens": 40000 },
"events": [
{ "seq": 1, "type": "step-start", "step": 1 },
{ "seq": 2, "type": "tool-call", "id": "call_01", "step": 1, "tool": "list_invoices", "arguments": { "month": "2026-08" } },
{ "seq": 3, "type": "tool-result", "callId": "call_01", "result": "ok", "durationMillis": 310, "tokens": 740 },
{ "seq": 4, "type": "step-end", "step": 1, "tokens": 1200 },
{ "seq": 5, "type": "outcome", "outcome": "completed" }
]
}
}Six event types: step-start, step-end, tool-call, tool-result, wait,
outcome. Four outcomes: completed, failed, interrupted,
budget-exhausted. seq is the ordering key and must strictly increase -- this
tool never reads a clock to decide what happened first.
Field-by-field details are in docs/trace-format.md.
An event field this tool does not read is reported as event-key-unrecognized
and ignored, so a trace carrying extra runtime metadata still works. An event
type it does not know is a different matter: it could be a call, so the run is
incomplete.
Forty-nine rules, each with a fixed severity, listed in full in
docs/trace-rules.md. Three classes:
| Class | Severity | Status | Exit |
|---|---|---|---|
| evidence the trace does not contain -- unread document, unreadable event, unanswered call, interrupted run, undeclared budget | error, and four warnings |
incomplete |
2 |
| a fact the trace establishes -- a loop, a failed run, a budget overrun | error |
fail |
1 |
| worth seeing, decides nothing -- a retry, one failed call, an unread field | info |
pass |
0 |
Forty-one of the forty-nine make a run incomplete. Four of those are only
warning severity -- budget-undeclared, call-outside-step,
no-events-recorded and step-unterminated -- so for them that membership is
the only thing standing between the trace and exit 0, and
test/incompleteness.test.mjs drives each one through a real entry point.
Severity is not defended by the table alone.
test/severity-behaviour.test.mjs drives every one of the forty-nine rules
through the real command line over a real trace and asserts the exit code,
because three declarations agreeing with each other can be edited together and
an exit code cannot.
stdout carries the JSON report and nothing else, so it can be piped straight
into a parser. --human writes the readable summary there instead. Diagnostics
go to stderr.
{
"schemaVersion": "1",
"tool": "agent-loop-trace-reader",
"status": "fail",
"summary": { "checked": 22, "errors": 5, "warnings": 0, "toolCalls": 5, "repeatedGroups": 1 },
"findings": []
}Findings are ordered by (location.pointer, ruleId, message), compared by
UTF-16 code unit -- never by localeCompare or Intl.Collator, whose ICU data
differs between Node builds and would let two correct machines disagree about
the same output. Pointers therefore sort as text: /trace/events/10 comes
before /trace/events/2. That is documented rather than papered over, because
making the sort numeric would mean emitting a pointer that is not the JSON
Pointer it claims to be; the summary carries seq for every event when you want
chronological order.
Two runs over the same trace produce byte-identical stdout.
--summary FILE writes the structured summary: steps, calls, groups, waits,
consumption and outcome. It is written whatever the status -- describing an
interrupted trace is the job -- and it carries "complete": false and the
status inside it, so a consumer reading only the artifact cannot mistake an
unfinished run for a finished one.
The destination is checked before anything is opened, and a refused one is a
configuration error (exit 2, empty stdout). A symbolic link at the
destination is refused on sight rather than resolved -- resolving it is what
writes through it -- as is a destination that is not a regular file, and any
spelling of the trace itself including a hard link to it, which is caught by
device and inode because it shares no path text with the trace. No root is
declared, because the summary may legitimately go anywhere you can write: a
symbolically linked parent directory is therefore followed unless it leads
back to the trace, and the guard's root check is exercised by
test/destination.test.mjs for a library caller that does declare one. The
destination's directory is created if it is missing.
| Code | Meaning |
|---|---|
0 |
the trace was read whole, it ended, and nothing in it failed |
1 |
the trace establishes a failure: a loop, a failed run, a budget overrun |
2 |
invalid usage or configuration (stdout empty), or a trace that was unreadable, unfinished, interrupted or bounded out (an incomplete report on stdout) |
The two shapes of exit 2 are deliberate. A configuration error means the run never had a subject, so there is nothing to report about. An unreadable trace means the run had a subject and failed to obtain evidence about it, and a consumer needs the report to know which input was not read.
Every limit is enforced, wired to a flag and to --config, and reported by
name. Nothing is truncated to fit: a trace over a limit is refused whole and the
run is incomplete, because the half that is missing is exactly where the loop
or the unanswered call would have been.
| Flag | Default | Effect |
|---|---|---|
--max-bytes |
8388608 | refuses the trace |
--max-events |
50000 | refuses the trace |
--max-depth |
32 | refuses the trace; the highest value accepted is 1000 |
--max-millis |
10000 | refuses the analysis; no partial model is reported |
--loop-threshold |
3 | identical calls that count as a loop (minimum 2) |
An unknown limit name, an unknown option, an unknown configuration key and a repeated flag are all refused rather than ignored: a typo that falls back to a default is a real failure reported as a green run.
--max-depth also has a ceiling of 1000, because the argument digest recurses.
Raised past what the stack can carry, the process died with a bare Maximum call stack size exceeded and nothing on stdout -- an input failure wearing the shape
of a configuration failure. Asking for more than 1000 is now a configuration
error you can read, and a trace deeper than the budget is trace-too-deep: an
incomplete report naming the limit.
This tool does not:
- produce traces, instrument an agent, or integrate with any runtime. It reads a file;
- speak OpenTelemetry, OTLP, or any vendor's trace format. It reads the schema
in
docs/trace-format.md, and refuses a document it does not recognise rather than guessing; - read or reproduce tool arguments, tool results or model text. Arguments are
digested; results are recorded as
ok,errorornulland nothing more; - explain why a loop happened, or decide whether a repeat was justified. It reports that identical calls were made and names them;
- measure real time.
seqorders the events and the durations come from the trace; the only clock is this tool's own analysis budget; - judge whether the agent did the right thing. A
completedrun with no findings is a run this tool has nothing to say about.
Trace content -- tool names, wait reasons, outcome reasons -- is data. Text in a trace asking for different behaviour is text in a document, not an instruction to this tool.
src/-- implementation (text,digest,destination,document,analysis,index)bin/-- the command line interfacetest/--node:test, no dependenciesexamples/-- a completed trace, a looping one, an interrupted onedocs/-- the trace format and the rule catalog
Zero runtime dependencies and zero development dependencies. Node 22 or later.
npm run checkMIT. See LICENSE.