Skip to content

[deep-report] Daily Max Ai Credits Test failures aren't flagged intentional_failure, skewing success-rate rollups #66401

Description

@github-actions

Description

This week's fleet-log sample shows "Daily Max Ai Credits Test" had 7 runs with 3 failures, all driver_exit-type consistent with its per-run max-ai-credits guardrail firing as designed — but none of the 3 failed runs were flagged intentional_failure: true in the logs tool's metadata, unlike its sibling "Daily Credit Limit Test" (14/14 flagged correctly). This means fleet-wide success-rate rollups that exclude intentional_failure runs currently miss this workflow and undercount its expected failures as real regressions.

Expected Impact

Keeps weekly/fleet success-rate reporting accurate by correctly classifying this guardrail test's expected failures, preventing a recurring false signal in health reports.

Suggested Agent

General Go agent working on the agenticworkflows/logs-classification tooling (wherever intentional_failure is derived/tagged)

Estimated Effort

Quick (< 1 hour — likely a missing name/pattern match in the classification logic)

Data Source

DeepReport fleet-health log analysis, 2026-09-30 to 2026-10-07 (agenticworkflows logs sampling)

Generated by 🔬 Deep Report · claude · agent · 553.8 AIC · ⌖ 7.81 AIC · ⊞ 7.1K · ◷

  • expires on Oct 8, 2026, 5:41 PM UTC-08:00

Activity

  1. github-actions commented on Oct 9, 2026

    @github-actions
    ContributorAuthor

    This issue was automatically closed because it expired on 2026-10-09T01:41:33.303Z.

    Closed by Workflow

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions