Skip to content

ci: demo an Allure test report published to GitHub Pages - #3006

Draft
Andy-Jost wants to merge 20 commits into
NVIDIA:mainfrom
Andy-Jost:ajost/allure-report-demo
Draft

Andy-Jost wants to merge 20 commits into
NVIDIA:mainfrom
Andy-Jost:ajost/allure-report-demo

Conversation

@Andy-Jost

@Andy-Jost Andy-Jost commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Summary

A demonstration, not a merge candidate. It renders every test job's JUnit XML as an Allure Report 3 report and publishes it to GitHub Pages next to the docs previews, following the suggestion in the discussion around #2963. The point is to see a real report on our own data and to put the costs on the table. The branch is stacked on #2963, which produces the JUnit artifacts, so only the top three commits belong to this PR.

What the demo does

  • ci/tools/allure_prepare.py (with tests) flattens the downloaded test-results-* artifacts into one directory and tags every <testsuite> with its CI configuration through the package attribute. Without the tag, Allure identifies a test by class name and test name only and folds the same test from every configuration into "retries" of one result. The script also writes an allurerc.mjs whose environments map the tag to one Allure environment per configuration, so each configuration is listed separately and a test's result shows side by side across configurations.
  • The final Check job status job downloads the artifacts, runs the script, generates the report with npx allure@3.20.0, uploads the report as a workflow artifact, publishes it to docs/test-report/pr-N/ on the gh-pages branch with the same deploy action the docs preview uses, and posts a sticky comment with the link. Every step is continue-on-error; the gate is unchanged.
  • Allure is a command-line tool from npm, run as a plain run: step, so the enterprise allowed-actions policy that blocked other reporters does not apply. It reads JUnit XML directly, so there is no new test dependency.
  • Two TEMPORARY commits make the run fast and give the report something to show: eight injected test failures, and a matrix trimmed to one row per platform.

Known blockers

Measured on a local run over the three-configuration demo data, about 20,000 test results:

  • Size. The report is 152 MB in 26,586 files on disk, about 12 MB compressed, which is close to what each deploy adds to the gh-pages history; a checkout of the branch grows by the full 152 MB. The full PR matrix is roughly twenty times the demo. Every run commits the whole report to the gh-pages branch, which already grows by about 0.5 GB a month (Investigate gh-pages impact on full clone size and time #2197). A per-PR report at this size is not viable. A nightly report overwritten in place might be.
  • Cleanup. ci/cleanup-pr-previews removes docs/pr-preview/pr-N/ for closed PRs and does not know about docs/test-report/. The folder this PR creates has to be removed by hand when the PR closes.
  • Captured output. pytest's JUnit files carry no stdout or stderr unless junit_logging is set. -o junit_logging=all -o junit_log_passing_tests=false would add the captured output of failed tests only, and Allure shows it as attachments. Not enabled here.
  • History. Allure's trend charts need the previous report's history file carried into the next run. Not wired here.
  • Deploy race. The docs preview and this report push to gh-pages in the same CI run, about an hour apart. The branch ruleset rejects non-fast-forward pushes and the deploy action retries, but this has not been exercised under load.

Measured in CI

The first run of this PR published the report at https://nvidia.github.io/cuda-python/test-report/pr-3006/ (three configurations, 19,757 results, 21 failed and 3 broken, one environment per configuration in the Environments menu).

  • The report is 152 MB in 26,738 files on disk, 12 MB compressed, now committed to gh-pages.
  • The GitHub Pages build that published it took 448 s against the 10-minute timeout. That is no slower than the docs-preview build before it (447 s) and the latest-docs build (547 s): the build time is set by the whole site, which is already near the cap.
  • In the two hours around this run, 6 of 12 Pages builds errored with "Page build failed", all docs-preview deploys. The branch-based Pages pipeline is already saturated; adding a deploy per CI run adds to that load.

Related Work

🤖 Generated with Claude Code

Andy-Jost and others added 20 commits September 29, 2026 09:08
A failed test job's log ended with about 700 lines of skip messages, the
pytest-run-parallel report, the durations table, and the warnings summary.
The traceback sat above all of that, and with -rxXs the short test summary
listed skips and xfails but not the failed test.

- run-tests: pytest runs with -rsxXfE, which orders the short test summary
  as skips, xfails, xpasses, failures, errors, so the FAILED/ERROR one-liners
  are the last lines before the final count. Each pytest call pipes stdout
  through ci/tools/fold_pytest_sections.py, which wraps the warnings summary,
  the durations table, the pytest-run-parallel report, and the skip list in
  GitHub log groups. The cuda.core calls pass --exclude-warning-annotations.
- cuda_core test group: add pytest-github-actions-annotate-failures, which
  emits one ::error annotation per failed test under GitHub Actions.

Only stdout goes through the fold script. The annotations go to stderr, and
the runner honors a workflow command only at the start of a line of its own
stream.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Revert this commit before the PR leaves draft.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
run-tests writes one JUnit XML file per suite under test-results/. Each
test job uploads that directory as a test-results-* artifact, and the
final "Check job status" job merges every job's files into two
informational checks:

- "Test Results (aggregate)" (EnricoMi/publish-unit-test-result-action):
  one entry per distinct test across the matrix, with the number of runs
  that failed, plus a PR comment when there are failures.
- "Test Results (by configuration)" (dorny/test-reporter): one section
  per test-results file, so each failure appears under the job
  configuration that produced it.

Both publish steps run with continue-on-error and never fail on test
failures, so the gate in the same job is unchanged. The two reporters
are published side by side for comparison; one of them is meant to be
kept.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Follow-ups from review:

- The nightly PyTorch interop, numba-cuda-mlir, and released cuda-core
  steps in the test workflows still ran pytest with -rxXs. They now use
  -rsxXfE, pipe through the fold script, and write JUnit XML into the
  same test-results/ directory, so every test log ends the same way and
  every suite reaches the "Test Results" checks.
- cuda.bindings gets pytest-github-actions-annotate-failures in its test
  group, so bindings failures are annotated like cuda.core failures. The
  released cuda-core step keeps warning annotations on because its test
  group comes from the released tag, which has no such plugin.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Two problem-matcher files under .github/matchers, each registered by one
step:

- compute-sanitizer.json, registered before the test steps in the Linux
  test workflow: memcheck findings ("Invalid __global__ read ...", the
  non-zero "ERROR SUMMARY" line) become annotations on the job page
  instead of staying buried in the pytest output of sanitizer jobs.
- compiler.json, registered after checkout in the wheel-build workflow:
  gcc/clang "file:line:col: error:", MSVC "file(line): error Cnnnn:",
  and Cython "file.pyx:line:col: message" lines become annotations.
  Errors only, to stay under the per-step annotation cap. Linux builds
  run inside a container, so their paths do not map onto the checkout
  and those annotations stay at the job level.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
One fast row per platform instead of the full matrix, so each CI round
takes about 25 minutes instead of an hour. Revert this commit before the
PR leaves draft.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…y search

dorny/test-reporter with use-actions-summary on writes only to the job
summary and creates no check run, so the "by configuration" report never
reached the PR. EnricoMi's publisher looked the PR up by branch and found
none, because CI runs on copy-pr-bot's pull-request/N mirror of a fork
PR; search_pull_requests finds it by commit and pull_request_build=commit
matches the pushed SHA.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…orters

Third informational publisher in the final job: worded status columns,
a plain-English preamble in the job summary, no annotations, and a PR
comment. Same continue-on-error posture as the other two; one of the
three is meant to be kept.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
EnricoMi: report_individual_runs off, so a test that fails on every
configuration gets one "All N runs failed" annotation instead of being
split by platform-specific traceback text.

dorny: fail-on-error on, so the per-configuration check is red when any
test failed. The step keeps continue-on-error, so the gate is unaffected.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…omment

Check runs created from inside a workflow are filed under the oldest
workflow that ran on the commit. On a copy-pr-bot mirror that is always
one of the fork PR's own workflows, so both "Test Results" checks showed
up under the assignee/label check. Neither reporter can choose the
suite, so neither creates a check run any more: EnricoMi writes its
counts to the job summary, dorny writes its per-configuration report
there too, and a short "how to read" paragraph precedes both.

The PR comment is now ours: ci/tools/junit_comment.py renders a plain
English summary from the same JUnit files (configurations, tests that
did not pass and where, execution totals, a link to the Summary page)
and marocchino/sticky-pull-request-comment posts it with recreate, so
it reappears at the bottom of the thread on every run.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ur own comment"

This reverts commit 42c80a7.

Back to check runs for both reporters. The Summary page has no link from
the PR without custom code, and the comment script was more code than the
report is worth. The next commit posts a one-line comment built from the
reporters' own step outputs instead.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Both reporters publish check runs again. The comment on the PR is now a
single line: dorny's passed/failed/skipped counts, and a link to each
check run (EnricoMi's check_url from its json output, dorny's url_html).
EnricoMi's own comment is off so there is one comment, not two.

Also drops list-files from the dorny step, which is not one of its inputs.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…layout"

This reverts commit 83db5a2.
The injected failures served their purpose on the demo runs.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This reverts commit 0248098.
The full PR matrix is back.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
# Conflicts:
#	.github/workflows/test-wheel-linux.yml
#	.github/workflows/test-wheel-windows.yml
The final job goes back to main's version: no EnricoMi or dorny check
runs and no test-results comment. The JUnit XML artifacts stay as the
data feed; the report experiments continue in a separate PR.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Revert this commit before the PR leaves draft.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
One fast row per platform. The aarch64 row uses the L4 pool, which had
runners when the A100 pool did not. Revert this commit before any merge.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The final job renders every test job's JUnit XML as an Allure Report 3
report and publishes it to docs/test-report/pr-N/ on the gh-pages branch,
the way the docs previews are published, with a sticky comment that links
to it. ci/tools/allure_prepare.py tags every suite with its CI
configuration and writes an allurerc.mjs that maps the tags to Allure
environments; without that, Allure folds the same test from every
configuration into retries of one result.

Allure is a CLI from npm run as a plain step, so the actions allowlist
does not apply, and it reads JUnit XML directly. Every step is
continue-on-error; the gate is unchanged.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@Andy-Jost Andy-Jost added this to the cuda.core 1.3.0 milestone Oct 2, 2026
@Andy-Jost Andy-Jost added enhancement Any code-related improvements P0 High priority - Must do! CI/CD CI/CD infrastructure labels Oct 2, 2026
@copy-pr-bot

copy-pr-bot Bot commented Oct 2, 2026

Copy link
Copy Markdown
Contributor

Auto-sync is disabled for draft pull requests in this repository. Workflows must be run manually.

Contributors can view more details about this message here.

@Andy-Jost Andy-Jost self-assigned this Oct 2, 2026
@Andy-Jost

Copy link
Copy Markdown
Contributor Author

/ok to test

@github-actions github-actions Bot added cuda.bindings Everything related to the cuda.bindings module cuda.core Everything related to the cuda.core module labels Oct 2, 2026
@github-actions

github-actions Bot commented Oct 2, 2026

Copy link
Copy Markdown
Contributor

@github-actions

github-actions Bot commented Oct 2, 2026

Copy link
Copy Markdown
Contributor

Allure test report for 9b5bb58: https://nvidia.github.io/cuda-python/test-report/pr-3006/

The report is 152M on disk in 26738 files (12 MB compressed), all committed to the gh-pages branch.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CI/CD CI/CD infrastructure cuda.bindings Everything related to the cuda.bindings module cuda.core Everything related to the cuda.core module enhancement Any code-related improvements P0 High priority - Must do!

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant