Skip to content

[Feature Request][RFC] Opt-in structured per-trial reporting for the autotuner #3375

Description

@WenzheWang

Required prerequisites

  • I searched existing issues/PRs for autotuning trial reporting, reports, history, feedback and callbacks.

Motivation

I'd like to work on an opt-in structured outcome stream for the autotuner, if this fits the maintainers' direction.

AutoTuner.run() aggregates compile and benchmark outcomes, but returns only the best AutotuneResult. A caller diagnosing a generated search space cannot directly associate every failed configuration with a stable index/outcome without parsing logs or wrapping internal methods. A small consumer could write per-run JSONL and extract only failed configs for a minimal reproduction. This is diagnostic infrastructure, not a claim of automatic self-improvement or faster tuning.

Proposed narrow scope

An optional synchronous observer on the aggregation/caller thread; illustrative API, not a fixed request:

with open("trials.jsonl", "w") as stream:
    result = tuner.run(
        on_trial=lambda record: stream.write(json.dumps(record) + "\n")
    )

The initial prototype records index, a detached JSON representation of config, status, latency_ms when measured, retained error text capped at 2048 characters, and reference-validation status. Stringifying an exception can still allocate its full message before truncation. It distinguishes compile failures, generic benchmark errors and timeouts without inferring correctness errors from traceback text. Skipped/absent reference checks are not reported as passed.

  • Preserve default scheduling, timeout/drain semantics, cache identity/format, best selection and AutotuneResult.
  • One record per config on normal search completion with a healthy sink and JSON-compatible payloads; aggregation order with explicit original config index, not strict physical completion order. Aborted runs or sink failures leave incomplete reports, without invented outcomes.
  • A cache hit produces a run-level marker, not reconstructed trial history; callbacks execute outside the shared cache lock.
  • No tuner-owned accumulated history or kernel/input-tensor payloads. Unsupported config serialization or ordinary sink failures warn and disable the sink for that run, preserving tuning; arbitrary Python config objects are outside the initial observer contract.
  • Initially exclude observer + early_stop, since current results do not identify full measurements versus early estimates.
  • No LLM/API dependencies, agent framework, persistent failure cache, search-space pruning, decorator forwarding or benchmark-worker lifecycle changes.

Prototype and validation plan

A local exploratory patch against main 994b44e and a JSONL failure-config consumer are prepared. It is not yet a published implementation. CPU mechanism checks extract the actual patched run() and worker methods and inject native boundaries; this is explicitly not package-import, compilation, GPU or performance validation.

Before an implementation PR: cover success/compile-unit failure/benchmark error/timeout/all-failed runs, stable indices and exact-once results, completion ordering, reporting on/off best-result equivalence, sink failure/config isolation, cache hits and disabled validation. Then run real package tests, the literal repository pre-commit wrapper, and a small single-GPU kernel tuning comparison with reporting on/off. A timeout fixture does not establish safe GPU-call drain.

Alternatives and overlap boundaries

Parsing existing logs avoids an API but couples callers to text and misses a uniform machine-readable outcome contract. A JSONL-only output option would be simpler if maintainers prefer a fixed artifact instead of callbacks. Extending the best-result cache with full history is outside scope.

This should remain separate from #2840 (timed-out call draining), #3337 (decorator profiler settings), #3338 (compile flags/cache identity), and #3329 (Ascend operator-generation/tuning agents). Some edits would touch the same aggregator in tuner.py, so I would rebase rather than include those PRs' changes.

Questions before expanding the patch

  1. Is structured per-config diagnostic feedback useful enough to maintain upstream?
  2. Would you prefer a callback, a JSONL output option, or an extension to an existing reporting interface?
  3. Is there ongoing work in this area or a preferred owner/interface to coordinate with?

I'll keep the exploratory scope small until the direction is clear. AI assistance was used in preparing the prototype/design; I will review code before publishing an implementation PR.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions