Skip to content

Latest commit

 

History

20 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Agent Loop Trace Reader

Read a saved agent trace and say what the run actually did -- including when it did not finish.

What it does

Given a saved trace, it produces a report and a structured summary covering:

  • steps, with the calls and waits belonging to each, and whether each one ended;
  • tool calls, each with its id, its step, a digest of its arguments, its duration and its result -- ok, error, or null for a call the trace never answered;
  • retries, linked to the call they retry;
  • waits, totalled and attributed to the step that waited;
  • repeated identical calls, grouped by tool and argument digest, with every call id preserved;
  • loops, both the same call repeated and two consecutive steps doing identical work;
  • budget consumption, compared with the budget the trace declares;
  • the terminal outcome, and whether there was one at all.

Nothing is fetched, no model is called, and the only clock is the one the command line injects for this tool's own analysis budget.

Why it exists

A trace is read after something went wrong, which is the worst moment to be given a reassuring answer. Three properties follow from that:

  1. An unanswered call is not a successful one. A tool-call with no tool-result means the trace holds the question and not the answer. It is reported as null, never as ok, and it makes the run incomplete.
  2. An interrupted trace stays incomplete. An outcome of interrupted, no outcome at all, a step that never ends, or events after the outcome: each one means the run has no result to report, so this tool abstains. Exit 2, never exit 0.
  3. Grouping never loses an id. A group that collapsed its members into a count would tell you something repeated and not which calls to go and look at. Every group carries every call id, in order, at its full declared length, and every call also remains listed individually. The one place an id list is bounded is a finding's evidence field, which the report contract bounds -- so it names whole ids, says how many it left out, and points at the artifact that carries them all.

And one property that is about the trace rather than the run: argument values are never read out. Calls are compared by a SHA-256 digest of their canonical form. A trace is the most likely place in a pipeline for a credential to be sitting inside a tool argument, and grouping identical calls is exactly the operation that tempts a tool to print them.

Quick start

# A run that finished cleanly
node bin/agent-loop-trace-reader.mjs --trace examples/completed-trace.json --human
echo "exit: $?"   # 0

# A run that looped and overran its budget -- a verdict, so exit 1
node bin/agent-loop-trace-reader.mjs --trace examples/looping-trace.json --human
echo "exit: $?"   # 1

# A run that was interrupted -- no verdict is possible, so exit 2
node bin/agent-loop-trace-reader.mjs --trace examples/interrupted-trace.json --human
echo "exit: $?"   # 2

# The machine-readable report, which is what stdout carries by default
node bin/agent-loop-trace-reader.mjs --trace examples/looping-trace.json | jq .summary

# The structured summary, written whatever the status
node bin/agent-loop-trace-reader.mjs \
  --trace examples/interrupted-trace.json \
  --summary build/interrupted-summary.json

npm run check     # lint, tests, all three examples, and a packaging dry run

Input

{
  "schemaVersion": "1",
  "trace": {
    "id": "invoice-export-2026-09-14",
    "budget": { "maxSteps": 8, "maxToolCalls": 20, "maxWaitMillis": 30000, "maxTokens": 40000 },
    "events": [
      { "seq": 1, "type": "step-start", "step": 1 },
      { "seq": 2, "type": "tool-call", "id": "call_01", "step": 1, "tool": "list_invoices", "arguments": { "month": "2026-08" } },
      { "seq": 3, "type": "tool-result", "callId": "call_01", "result": "ok", "durationMillis": 310, "tokens": 740 },
      { "seq": 4, "type": "step-end", "step": 1, "tokens": 1200 },
      { "seq": 5, "type": "outcome", "outcome": "completed" }
    ]
  }
}

Six event types: step-start, step-end, tool-call, tool-result, wait, outcome. Four outcomes: completed, failed, interrupted, budget-exhausted. seq is the ordering key and must strictly increase -- this tool never reads a clock to decide what happened first.

Field-by-field details are in docs/trace-format.md. An event field this tool does not read is reported as event-key-unrecognized and ignored, so a trace carrying extra runtime metadata still works. An event type it does not know is a different matter: it could be a call, so the run is incomplete.

Rules

Forty-nine rules, each with a fixed severity, listed in full in docs/trace-rules.md. Three classes:

Class Severity Status Exit
evidence the trace does not contain -- unread document, unreadable event, unanswered call, interrupted run, undeclared budget error, and four warnings incomplete 2
a fact the trace establishes -- a loop, a failed run, a budget overrun error fail 1
worth seeing, decides nothing -- a retry, one failed call, an unread field info pass 0

Forty-one of the forty-nine make a run incomplete. Four of those are only warning severity -- budget-undeclared, call-outside-step, no-events-recorded and step-unterminated -- so for them that membership is the only thing standing between the trace and exit 0, and test/incompleteness.test.mjs drives each one through a real entry point.

Severity is not defended by the table alone. test/severity-behaviour.test.mjs drives every one of the forty-nine rules through the real command line over a real trace and asserts the exit code, because three declarations agreeing with each other can be edited together and an exit code cannot.

Reports

stdout carries the JSON report and nothing else, so it can be piped straight into a parser. --human writes the readable summary there instead. Diagnostics go to stderr.

{
  "schemaVersion": "1",
  "tool": "agent-loop-trace-reader",
  "status": "fail",
  "summary": { "checked": 22, "errors": 5, "warnings": 0, "toolCalls": 5, "repeatedGroups": 1 },
  "findings": []
}

Findings are ordered by (location.pointer, ruleId, message), compared by UTF-16 code unit -- never by localeCompare or Intl.Collator, whose ICU data differs between Node builds and would let two correct machines disagree about the same output. Pointers therefore sort as text: /trace/events/10 comes before /trace/events/2. That is documented rather than papered over, because making the sort numeric would mean emitting a pointer that is not the JSON Pointer it claims to be; the summary carries seq for every event when you want chronological order.

Two runs over the same trace produce byte-identical stdout.

The summary artifact

--summary FILE writes the structured summary: steps, calls, groups, waits, consumption and outcome. It is written whatever the status -- describing an interrupted trace is the job -- and it carries "complete": false and the status inside it, so a consumer reading only the artifact cannot mistake an unfinished run for a finished one.

The destination is checked before anything is opened, and a refused one is a configuration error (exit 2, empty stdout). A symbolic link at the destination is refused on sight rather than resolved -- resolving it is what writes through it -- as is a destination that is not a regular file, and any spelling of the trace itself including a hard link to it, which is caught by device and inode because it shares no path text with the trace. No root is declared, because the summary may legitimately go anywhere you can write: a symbolically linked parent directory is therefore followed unless it leads back to the trace, and the guard's root check is exercised by test/destination.test.mjs for a library caller that does declare one. The destination's directory is created if it is missing.

Exit codes

Code Meaning
0 the trace was read whole, it ended, and nothing in it failed
1 the trace establishes a failure: a loop, a failed run, a budget overrun
2 invalid usage or configuration (stdout empty), or a trace that was unreadable, unfinished, interrupted or bounded out (an incomplete report on stdout)

The two shapes of exit 2 are deliberate. A configuration error means the run never had a subject, so there is nothing to report about. An unreadable trace means the run had a subject and failed to obtain evidence about it, and a consumer needs the report to know which input was not read.

Limits

Every limit is enforced, wired to a flag and to --config, and reported by name. Nothing is truncated to fit: a trace over a limit is refused whole and the run is incomplete, because the half that is missing is exactly where the loop or the unanswered call would have been.

Flag Default Effect
--max-bytes 8388608 refuses the trace
--max-events 50000 refuses the trace
--max-depth 32 refuses the trace; the highest value accepted is 1000
--max-millis 10000 refuses the analysis; no partial model is reported
--loop-threshold 3 identical calls that count as a loop (minimum 2)

An unknown limit name, an unknown option, an unknown configuration key and a repeated flag are all refused rather than ignored: a typo that falls back to a default is a real failure reported as a green run.

--max-depth also has a ceiling of 1000, because the argument digest recurses. Raised past what the stack can carry, the process died with a bare Maximum call stack size exceeded and nothing on stdout -- an input failure wearing the shape of a configuration failure. Asking for more than 1000 is now a configuration error you can read, and a trace deeper than the budget is trace-too-deep: an incomplete report naming the limit.

Non-goals

This tool does not:

  • produce traces, instrument an agent, or integrate with any runtime. It reads a file;
  • speak OpenTelemetry, OTLP, or any vendor's trace format. It reads the schema in docs/trace-format.md, and refuses a document it does not recognise rather than guessing;
  • read or reproduce tool arguments, tool results or model text. Arguments are digested; results are recorded as ok, error or null and nothing more;
  • explain why a loop happened, or decide whether a repeat was justified. It reports that identical calls were made and names them;
  • measure real time. seq orders the events and the durations come from the trace; the only clock is this tool's own analysis budget;
  • judge whether the agent did the right thing. A completed run with no findings is a run this tool has nothing to say about.

Trace content -- tool names, wait reasons, outcome reasons -- is data. Text in a trace asking for different behaviour is text in a document, not an instruction to this tool.

Repository layout

  • src/ -- implementation (text, digest, destination, document, analysis, index)
  • bin/ -- the command line interface
  • test/ -- node:test, no dependencies
  • examples/ -- a completed trace, a looping one, an interrupted one
  • docs/ -- the trace format and the rule catalog

Development

Zero runtime dependencies and zero development dependencies. Node 22 or later.

npm run check

License

MIT. See LICENSE.

About

Read agent traces as steps, tool calls, waits, retries and terminal outcomes.

Resources

Code of conduct

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages