Description
agenticworkflows logs (with artifacts:["all"]) exports only grader-inferred failure_kind/classification fields (e.g. agent_logic, driver_exit) for job-level failures — never the raw stderr/console text or the safe-output-errors.json artifact itself. This has blocked confident root-causing in roughly 5 of the last 8 Safe Output Health audits (most recently #65507, 2026-10-04), each of which had to infer a root cause from the failure shape alone rather than the actual error message.
Expected Impact
Converts a recurring "root cause unconfirmed, inferred from grader shape only" caveat into a confirmable diagnosis for every future fleet-health analysis (DeepReport, Safe Output Health, Agent Job Health, etc.), reducing wasted investigation cycles.
Suggested Action
Include the raw safe-outputs job console log tail (or at minimum the safe-output-errors.json artifact referenced in multiple historical audits) in the bundle returned by the logs tool/CLI when a job-level failure is present, alongside the existing structured failure_kind classification.
Suggested Agent: Owner of pkg/cli logs/export tooling
Estimated Effort: Medium (1-4 hours)
Data Source: DeepReport analysis of discussion #65507 and recurring pattern across prior Safe Output Health audits
Generated by 🔬 Deep Report · claude · agent · 310.1 AIC · ⌖ 9.15 AIC · ⊞ 7.1K · ◷
Description
agenticworkflows logs(withartifacts:["all"]) exports only grader-inferredfailure_kind/classificationfields (e.g.agent_logic,driver_exit) for job-level failures — never the raw stderr/console text or thesafe-output-errors.jsonartifact itself. This has blocked confident root-causing in roughly 5 of the last 8 Safe Output Health audits (most recently #65507, 2026-10-04), each of which had to infer a root cause from the failure shape alone rather than the actual error message.Expected Impact
Converts a recurring "root cause unconfirmed, inferred from grader shape only" caveat into a confirmable diagnosis for every future fleet-health analysis (DeepReport, Safe Output Health, Agent Job Health, etc.), reducing wasted investigation cycles.
Suggested Action
Include the raw safe-outputs job console log tail (or at minimum the
safe-output-errors.jsonartifact referenced in multiple historical audits) in the bundle returned by thelogstool/CLI when a job-level failure is present, alongside the existing structuredfailure_kindclassification.Suggested Agent: Owner of
pkg/clilogs/export toolingEstimated Effort: Medium (1-4 hours)
Data Source: DeepReport analysis of discussion #65507 and recurring pattern across prior Safe Output Health audits