Conversation
A failed test job's log ended with about 700 lines of skip messages, the pytest-run-parallel report, the durations table, and the warnings summary. The traceback sat above all of that, and with -rxXs the short test summary listed skips and xfails but not the failed test. - run-tests: pytest runs with -rsxXfE, which orders the short test summary as skips, xfails, xpasses, failures, errors, so the FAILED/ERROR one-liners are the last lines before the final count. Each pytest call pipes stdout through ci/tools/fold_pytest_sections.py, which wraps the warnings summary, the durations table, the pytest-run-parallel report, and the skip list in GitHub log groups. The cuda.core calls pass --exclude-warning-annotations. - cuda_core test group: add pytest-github-actions-annotate-failures, which emits one ::error annotation per failed test under GitHub Actions. Only stdout goes through the fold script. The annotations go to stderr, and the runner honors a workflow command only at the start of a line of its own stream. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Revert this commit before the PR leaves draft. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
run-tests writes one JUnit XML file per suite under test-results/. Each test job uploads that directory as a test-results-* artifact, and the final "Check job status" job merges every job's files into two informational checks: - "Test Results (aggregate)" (EnricoMi/publish-unit-test-result-action): one entry per distinct test across the matrix, with the number of runs that failed, plus a PR comment when there are failures. - "Test Results (by configuration)" (dorny/test-reporter): one section per test-results file, so each failure appears under the job configuration that produced it. Both publish steps run with continue-on-error and never fail on test failures, so the gate in the same job is unchanged. The two reporters are published side by side for comparison; one of them is meant to be kept. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Follow-ups from review: - The nightly PyTorch interop, numba-cuda-mlir, and released cuda-core steps in the test workflows still ran pytest with -rxXs. They now use -rsxXfE, pipe through the fold script, and write JUnit XML into the same test-results/ directory, so every test log ends the same way and every suite reaches the "Test Results" checks. - cuda.bindings gets pytest-github-actions-annotate-failures in its test group, so bindings failures are annotated like cuda.core failures. The released cuda-core step keeps warning annotations on because its test group comes from the released tag, which has no such plugin. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Two problem-matcher files under .github/matchers, each registered by one
step:
- compute-sanitizer.json, registered before the test steps in the Linux
test workflow: memcheck findings ("Invalid __global__ read ...", the
non-zero "ERROR SUMMARY" line) become annotations on the job page
instead of staying buried in the pytest output of sanitizer jobs.
- compiler.json, registered after checkout in the wheel-build workflow:
gcc/clang "file:line:col: error:", MSVC "file(line): error Cnnnn:",
and Cython "file.pyx:line:col: message" lines become annotations.
Errors only, to stay under the per-step annotation cap. Linux builds
run inside a container, so their paths do not map onto the checkout
and those annotations stay at the job level.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
One fast row per platform instead of the full matrix, so each CI round takes about 25 minutes instead of an hour. Revert this commit before the PR leaves draft. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…y search dorny/test-reporter with use-actions-summary on writes only to the job summary and creates no check run, so the "by configuration" report never reached the PR. EnricoMi's publisher looked the PR up by branch and found none, because CI runs on copy-pr-bot's pull-request/N mirror of a fork PR; search_pull_requests finds it by commit and pull_request_build=commit matches the pushed SHA. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…orters Third informational publisher in the final job: worded status columns, a plain-English preamble in the job summary, no annotations, and a PR comment. Same continue-on-error posture as the other two; one of the three is meant to be kept. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… two reporters" This reverts commit 8483791.
EnricoMi: report_individual_runs off, so a test that fails on every configuration gets one "All N runs failed" annotation instead of being split by platform-specific traceback text. dorny: fail-on-error on, so the per-configuration check is red when any test failed. The step keeps continue-on-error, so the gate is unaffected. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…omment Check runs created from inside a workflow are filed under the oldest workflow that ran on the commit. On a copy-pr-bot mirror that is always one of the fork PR's own workflows, so both "Test Results" checks showed up under the assignee/label check. Neither reporter can choose the suite, so neither creates a check run any more: EnricoMi writes its counts to the job summary, dorny writes its per-configuration report there too, and a short "how to read" paragraph precedes both. The PR comment is now ours: ci/tools/junit_comment.py renders a plain English summary from the same JUnit files (configurations, tests that did not pass and where, execution totals, a link to the Summary page) and marocchino/sticky-pull-request-comment posts it with recreate, so it reappears at the bottom of the thread on every run. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ur own comment" This reverts commit 42c80a7. Back to check runs for both reporters. The Summary page has no link from the PR without custom code, and the comment script was more code than the report is worth. The next commit posts a one-line comment built from the reporters' own step outputs instead. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Both reporters publish check runs again. The comment on the PR is now a single line: dorny's passed/failed/skipped counts, and a link to each check run (EnricoMi's check_url from its json output, dorny's url_html). EnricoMi's own comment is off so there is one comment, not two. Also drops list-files from the dorny step, which is not one of its inputs. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…layout" This reverts commit 83db5a2. The injected failures served their purpose on the demo runs. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This reverts commit 0248098. The full PR matrix is back. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
# Conflicts: # .github/workflows/test-wheel-linux.yml # .github/workflows/test-wheel-windows.yml
The final job goes back to main's version: no EnricoMi or dorny check runs and no test-results comment. The JUnit XML artifacts stay as the data feed; the report experiments continue in a separate PR. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Revert this commit before the PR leaves draft. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
One fast row per platform. The aarch64 row uses the L4 pool, which had runners when the A100 pool did not. Revert this commit before any merge. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The final job renders every test job's JUnit XML as an Allure Report 3 report and publishes it to docs/test-report/pr-N/ on the gh-pages branch, the way the docs previews are published, with a sticky comment that links to it. ci/tools/allure_prepare.py tags every suite with its CI configuration and writes an allurerc.mjs that maps the tags to Allure environments; without that, Allure folds the same test from every configuration into retries of one result. Allure is a CLI from npm run as a plain step, so the actions allowlist does not apply, and it reads JUnit XML directly. Every step is continue-on-error; the gate is unchanged. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Contributor
|
Auto-sync is disabled for draft pull requests in this repository. Workflows must be run manually. Contributors can view more details about this message here. |
Contributor
Author
|
/ok to test |
Contributor
|
Contributor
|
Allure test report for 9b5bb58: https://nvidia.github.io/cuda-python/test-report/pr-3006/ The report is 152M on disk in 26738 files (12 MB compressed), all committed to the gh-pages branch. |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
A demonstration, not a merge candidate. It renders every test job's JUnit XML as an Allure Report 3 report and publishes it to GitHub Pages next to the docs previews, following the suggestion in the discussion around #2963. The point is to see a real report on our own data and to put the costs on the table. The branch is stacked on #2963, which produces the JUnit artifacts, so only the top three commits belong to this PR.
What the demo does
ci/tools/allure_prepare.py(with tests) flattens the downloadedtest-results-*artifacts into one directory and tags every<testsuite>with its CI configuration through thepackageattribute. Without the tag, Allure identifies a test by class name and test name only and folds the same test from every configuration into "retries" of one result. The script also writes anallurerc.mjswhose environments map the tag to one Allure environment per configuration, so each configuration is listed separately and a test's result shows side by side across configurations.Check job statusjob downloads the artifacts, runs the script, generates the report withnpx allure@3.20.0, uploads the report as a workflow artifact, publishes it todocs/test-report/pr-N/on the gh-pages branch with the same deploy action the docs preview uses, and posts a sticky comment with the link. Every step iscontinue-on-error; the gate is unchanged.run:step, so the enterprise allowed-actions policy that blocked other reporters does not apply. It reads JUnit XML directly, so there is no new test dependency.Known blockers
Measured on a local run over the three-configuration demo data, about 20,000 test results:
gh-pagesimpact on full clone size and time #2197). A per-PR report at this size is not viable. A nightly report overwritten in place might be.ci/cleanup-pr-previewsremovesdocs/pr-preview/pr-N/for closed PRs and does not know aboutdocs/test-report/. The folder this PR creates has to be removed by hand when the PR closes.junit_loggingis set.-o junit_logging=all -o junit_log_passing_tests=falsewould add the captured output of failed tests only, and Allure shows it as attachments. Not enabled here.Measured in CI
The first run of this PR published the report at https://nvidia.github.io/cuda-python/test-report/pr-3006/ (three configurations, 19,757 results, 21 failed and 3 broken, one environment per configuration in the Environments menu).
Related Work
🤖 Generated with Claude Code