Skip to content

[Feature]: evidence.md convention for speckit.bug.assess (pre-digested inputs such as reduced logs) #4804

Description

@alexcpn

Problem Statement

speckit.bug.assess ingests a bug report as pasted text or a URL. Real bug
reports often come with a log: a CI run, kubectl logs, a crash-looping
service. These run to thousands or millions of lines. Today the agent either
reads the raw log (blowing the context window and burying the one relevant
error under repeats) or skims it ad hoc. Whatever it looked at is not recorded
in .specify/bugs/<slug>/, so speckit.bug.fix and speckit.bug.test can't
rely on it, and a reviewer can't check what the agent actually saw.

Proposed Solution

In Execution → 1. Ingest the bug report of speckit.bug.assess, add:

If BUG_DIR/evidence.md exists (for example, written by a
before_bug_assess hook), read it as part of the report. Cite it under
Report and prefer it over re-reading any raw log it was derived from.

Also add evidence.md (optional) to the per-bug directory layout in the
README, alongside assessment.md, fix.md and test.md.

This names no tool. With no evidence.md, behavior is exactly as today. The
file can be written by hand or by any extension. Using the
before_bug_assess hook from #4799 is the natural route, but this convention
works without that hook.

Motivating producer: log intake

A community extension, speckit.logreduce.intake,
finds log files referenced in the report and runs
logreduce with a token budget.
logreduce is a single static binary that applies TF-IDF over masked templates
plus severity weighting. The extension writes BUG_DIR/evidence.md containing
the command it ran, the source path, the summary header and the reduced log.

  • On a 1M-line log, that is a ~99.9% token reduction.
  • On the LogDx CI-incident benchmark (35 cases), it kept 99% of the
    human-labelled critical lines at an 8k-token budget.

Assess, fix and test then all work from the same saved evidence file instead of
the raw log. I'll maintain that extension and submit it to the community
catalog. This issue only asks for the convention.

Alternatives Considered

  • Bake log reduction into speckit.bug.assess. That adds a tool-specific
    dependency to a bundled extension. A file convention keeps core neutral.
  • A standalone command the user runs first, with no convention. This works
    today, but assess doesn't know to read the result.

Component

Extensions: bundled bug extension (extensions/bug/)

AI Agent (if applicable)

All.

Use Cases

  1. A CI job fails with a 40k-line log. Intake writes about 7k tokens of reduced
    log to evidence.md, and the assessment cites the first error and the
    crash-loop pattern from it.
  2. speckit.bug.test fails and recommends re-running assess. The new failing
    log goes through the same intake, so evidence.md shows the before and
    after.

Acceptance Criteria

  • speckit.bug.assess reads BUG_DIR/evidence.md when present and cites it
    under Report.
  • evidence.md (optional) is listed in the per-bug directory layout in the
    bug extension README.
  • With no evidence.md, the command output is unchanged.

I'm happy to send the PR.

Additional Context

AI Disclosure

Drafted with Claude Code (Claude Opus 5.5) from my notes. I reviewed it.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    feature-assessRun the Spec Kit idea-assessment pipeline on this feature requesttriage-can-waitVerdict: valid and in-scope but deprioritized; held behind the evidence gate

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions