Posted by GitHub Copilot on behalf of @heiskr.
Summary
The safeoutputs CLI transport has a failure mode that is indistinguishable from success, both to the agent and to the run logs. An agent that makes one small shell mistake will believe it emitted a safe output, finish its turn, and the run ends with zero safe outputs.
I diagnosed one run end to end. The agent did the analysis correctly and correctly decided no issue was warranted. It failed only at signaling.
The mechanism
The documented invocation pipes JSON into the CLI with . as the stdin sentinel:
printf '{"message":"..."}' | safeoutputs noop .
The agent emitted the same command without the pipe:
printf '{"message":"..."}' safeoutputs noop .
printf takes a format string plus arguments, so it consumed safeoutputs, noop, and . as format arguments, printed the JSON, and exited 0. The CLI never ran.
The two forms are byte-identical in stdout and exit code:
$ printf '{"message":"no action needed"}' safeoutputs noop .
{"message":"no action needed"}
$ echo $?
0
$ printf '{"message":"no action needed"}' | safeoutputs noop .
{"message":"no action needed"}
$ echo $?
0
Worth noting: the first command succeeds identically on a machine where the safeoutputs binary does not exist at all. I ran the transcript above on a laptop with no gh-aw installed. Nothing in the result gives the agent a signal.
The agent then wrote that it had called noop, and stopped. tools/call count against the safeoutputs MCP server for the whole run was 0, and agent_output.json was {"items":[],"errors":[]}.
Why this is a design issue, not an agent bug
The sentinel design puts the binary name in the middle of the command line. Every other shape of this mistake also fails silently, because printf and echo accept arbitrary trailing arguments and ignore them.
By contrast, an inline-argument form would put the binary first, so any malformed variant still executes the binary and can be diagnosed or rejected.
The shipped prompt calls the CLI "an optional equivalent transport." It is not equivalent. The MCP tool path validates against the tool's inputSchema and returns an error the agent can see. The CLI path can skip execution entirely.
The prompt already handles a neighbouring stdin bug loudly:
Piping cat file | safeoutputs ... does not populate body and will be rejected.
That one can be rejected because the binary runs. The no-stdin-at-all case never reaches the binary, so nothing is in a position to reject it.
Related: the same prompt section warns agents not to do the thing that triggered this.
do NOT build the JSON payload with printf/echo embedding raw newlines or many escaped characters directly in the command line
The failing payload was a long single-line printf with escaped characters. The warning is there, and the agent still walked into it, which suggests prose guidance alone is not carrying the weight.
Scope
[aw] ... produced no safe outputs issues in this repo: 15 open, 100+ closed, across at least 11 distinct workflows. Six were opened on a single day.
To be precise about what I can and cannot support: I have proven this mechanism for one run. A genuine "the agent forgot to call any tool" is a separate and real cause, and the prompt itself calls that the #1 cause of workflow failures. I cannot tell you how many of the 115 are which, and neither can anyone else right now, which is the point of the third proposal below.
Every one of these I have looked at was closed as "transient" with no diagnosis. This one is not transient. It is deterministic.
Proposals
Listed cheapest first. They are independent.
-
Enrich the "produced no safe outputs" report. When a run ends with 0 safe outputs, scan the agent transcript for shell commands containing safeoutputs and include them in the issue body. Non-breaking, and it makes every future occurrence self-diagnosing instead of "transient." This alone would separate the two root causes across the ~115 existing reports.
-
Accept an inline JSON argument, for example safeoutputs noop '{"message":"..."}', and keep stdin as an option. This removes the mid-command binary placement that makes the mistake silent. Piping remains available for large bodies.
-
Drop the CLI transport for safeoutputs. The prompt already tells agents to prefer direct tool calls, and the MCP path validates payloads. The github CLI can stay. This is the largest change and may be more than you want.
A second, independent bug in the same run
The payload used a reason field. The noop schema requires message. Even piped correctly it would have failed. The CLI --help is schema-derived and would have shown this, but nothing forced the agent to look.
If noop is the designated escape hatch and the thing agents are told to call when they have nothing else to say, it may be worth accepting any single free-text field, or accepting no field at all, rather than hard-requiring message.
Posted by GitHub Copilot on behalf of @heiskr.
Summary
The
safeoutputsCLI transport has a failure mode that is indistinguishable from success, both to the agent and to the run logs. An agent that makes one small shell mistake will believe it emitted a safe output, finish its turn, and the run ends with zero safe outputs.I diagnosed one run end to end. The agent did the analysis correctly and correctly decided no issue was warranted. It failed only at signaling.
The mechanism
The documented invocation pipes JSON into the CLI with
.as the stdin sentinel:The agent emitted the same command without the pipe:
printftakes a format string plus arguments, so it consumedsafeoutputs,noop, and.as format arguments, printed the JSON, and exited 0. The CLI never ran.The two forms are byte-identical in stdout and exit code:
Worth noting: the first command succeeds identically on a machine where the
safeoutputsbinary does not exist at all. I ran the transcript above on a laptop with no gh-aw installed. Nothing in the result gives the agent a signal.The agent then wrote that it had called
noop, and stopped.tools/callcount against the safeoutputs MCP server for the whole run was 0, andagent_output.jsonwas{"items":[],"errors":[]}.Why this is a design issue, not an agent bug
The sentinel design puts the binary name in the middle of the command line. Every other shape of this mistake also fails silently, because
printfandechoaccept arbitrary trailing arguments and ignore them.By contrast, an inline-argument form would put the binary first, so any malformed variant still executes the binary and can be diagnosed or rejected.
The shipped prompt calls the CLI "an optional equivalent transport." It is not equivalent. The MCP tool path validates against the tool's
inputSchemaand returns an error the agent can see. The CLI path can skip execution entirely.The prompt already handles a neighbouring stdin bug loudly:
That one can be rejected because the binary runs. The no-stdin-at-all case never reaches the binary, so nothing is in a position to reject it.
Related: the same prompt section warns agents not to do the thing that triggered this.
The failing payload was a long single-line
printfwith escaped characters. The warning is there, and the agent still walked into it, which suggests prose guidance alone is not carrying the weight.Scope
[aw] ... produced no safe outputsissues in this repo: 15 open, 100+ closed, across at least 11 distinct workflows. Six were opened on a single day.To be precise about what I can and cannot support: I have proven this mechanism for one run. A genuine "the agent forgot to call any tool" is a separate and real cause, and the prompt itself calls that the #1 cause of workflow failures. I cannot tell you how many of the 115 are which, and neither can anyone else right now, which is the point of the third proposal below.
Every one of these I have looked at was closed as "transient" with no diagnosis. This one is not transient. It is deterministic.
Proposals
Listed cheapest first. They are independent.
Enrich the "produced no safe outputs" report. When a run ends with 0 safe outputs, scan the agent transcript for shell commands containing
safeoutputsand include them in the issue body. Non-breaking, and it makes every future occurrence self-diagnosing instead of "transient." This alone would separate the two root causes across the ~115 existing reports.Accept an inline JSON argument, for example
safeoutputs noop '{"message":"..."}', and keep stdin as an option. This removes the mid-command binary placement that makes the mistake silent. Piping remains available for large bodies.Drop the CLI transport for
safeoutputs. The prompt already tells agents to prefer direct tool calls, and the MCP path validates payloads. ThegithubCLI can stay. This is the largest change and may be more than you want.A second, independent bug in the same run
The payload used a
reasonfield. Thenoopschema requiresmessage. Even piped correctly it would have failed. The CLI--helpis schema-derived and would have shown this, but nothing forced the agent to look.If
noopis the designated escape hatch and the thing agents are told to call when they have nothing else to say, it may be worth accepting any single free-text field, or accepting no field at all, rather than hard-requiringmessage.