Extracts test metrics from JUnit XML files and sends them to OTLP endpoints with minimal cardinality for efficient storage and querying.
- Reads JUnit XML files from a directory
- Parses test results and durations
- Generates low-cardinality OpenTelemetry metrics optimized for performance regression detection
- Ships metrics to OTLP-compatible backends (Prometheus, Mimir, Grafana Cloud, etc.)
- uses: redis-developer/cae-otel-ci-visibility@v4
with:
junit-xml-folder: './test-results'
otlp-endpoint: 'https://otlp.example.com/v1/metrics'
otlp-headers: 'authorization=Bearer ${{ secrets.OTLP_TOKEN }}'
# optional — set from your CI matrix when tests run against multiple
# server versions, so regressions are detected per version:
server-version: ${{ matrix.redis-version }}| Input | Required | Default | Description |
|---|---|---|---|
junit-xml-folder |
yes | - | Path to directory containing JUnit XML files |
otlp-endpoint |
yes | - | OTLP metrics endpoint URL |
otlp-headers |
no | - | OTLP headers (key=value,key2=value2 or JSON) |
branch-allowlist |
no | default branch | Branches to emit metrics for (comma-separated, * = all) |
server-version |
no | - | Version track of the system under test: unstable or major.minor like 8.4 — emitted as the server.version label. Anything else fails the build |
on-nondeterministic-ids |
no | warn |
What to do when test ids look nondeterministic: warn | skip | break (see below) |
Pass the version track you want a test compared against over time. Exactly two value shapes are accepted:
unstable— the server is built from master / under development.major.minor(e.g.3.4,8.10) — a stable server API version track. Release candidates, previews and GA builds collapse into the same track:8.10-rc2and8.10.0must both be passed as8.10. Each series then tells the story of that track's whole lifecycle, and a slowdown between two release candidates fires the same regression signal as one after GA — for a client library that's early warning either way.
Any other value fails the build. The server.version label multiplies
series per test, so its values must be stable and bounded — a patch version,
docker image tag, commit SHA or ${{ github.sha }} here would silently mint new
metric series every run. The failure message tells you what to pass instead
(8.4.0 → 8.4, 8.10-rc2 → 8.10).
The version set stays bounded on its own: new tracks enter the matrix a few times a year, old ones leave and their series age out. When the input is unset, all of a repo's runs share one series per test — fine for repos without a server under test, but set it whenever the CI matrix varies the server (the action reminds you in its log output when it's unset).
Branch names multiply metric cardinality, and short-lived branches (PRs) never
repeat — so by default metrics are emitted only when the workflow runs on the
repository default branch. Set branch-allowlist to a comma-separated list
(e.g. master,releases/v2) to emit for those branches instead, or * to emit
everywhere (not recommended).
Runs on the default branch are labelled vcs.repository.ref.name="default"
rather than with the branch name (since 4.5.0). Default branches are called
master, main, unstable, ... across repositories, and one literal lets a
dashboard select every repository's default branch with a single matcher.
Allowlisted extra branches keep their real name (releases/v2). When the
triggering event does not carry the repository's default branch, the real name
is kept — so dashboards should match master|main|default while older action
versions are still in use.
Generates one low-cardinality per-test metric optimized for performance regression detection, plus three small per-run/per-suite rollups (~5 series + one per top-level suite per repo) for headline stats and cheap alerting:
A gauge metric recording individual test execution duration. The cae namespace
and v16 schema version are hardcoded.
Labels:
| Label | Description | Cardinality |
|---|---|---|
test.id |
Unique test identifier (see below) | High but bounded |
vcs.repository.name |
Repository (e.g., owner/repo) |
Low |
vcs.repository.ref.name |
default on the default branch, else the branch name |
Low |
server.version |
System under test version (only when input set) | Low, bounded |
Total: up to 4 labels. Deliberately no per-run labels (run IDs, commit
SHAs): a label value that never repeats mints one new series per test on every
CI run, growing cardinality as tests × runs. With stable labels each run
appends samples to existing series and cardinality stays at
tests × repos × branches × server versions. Without server.version, a CI
matrix's jobs all write into one series per test, mixing distributions from
different environments — set it when the matrix varies the system under test.
Three additive gauges summarize each run without scanning thousands of per-test
series. They carry the same base labels as the per-test metric
(vcs.repository.name, vcs.repository.ref.name, and server.version when set
— no per-run labels), so the series-count impact is ~5 series + one per
top-level suite, per repo/branch/server-version.
Number of tests in the run, by result status — one data point per status.
| Label | Description | Cardinality |
|---|---|---|
test.result.status |
passed | failed | error | skipped |
4 |
vcs.repository.name |
Repository | Low |
vcs.repository.ref.name |
default on the default branch, else branch |
Low |
Cumulative duration of all tests in the run (sum of per-test times). One series
per repo/branch/server-version — the cheap target for "is this repo still
reporting?" alerts and per-repo trend panels. Labels: vcs.repository.name,
vcs.repository.ref.name, plus server.version when set.
Cumulative duration of all tests in each top-level test suite (nested suites roll up into their parent, so they get no point of their own). Catches "every test in the suite got a little slower" — invisible to the per-test regression gate.
| Label | Description | Cardinality |
|---|---|---|
suite.id |
Suite name, whitespace-normalized; over 256 chars truncated to head...tail___hash8; unnamed if empty |
One per suite |
vcs.repository.name |
Repository | Low |
vcs.repository.ref.name |
default on the default branch, else branch |
Low |
test.id is the full human-readable {class}.{test} name:
com.redis.lettucemod.RedisModulesClientTest.testTimeSeriesAdd
tests.unit.test_search.TestQueryBuilder.test_paging_offset
BF.ADD transformArguments
Rules:
- Suite names are dropped — they usually duplicate the class name (Surefire)
or are constant noise (
pytest). The suite is used as fallback context when the class name is empty, and folded back in only when two different tests would otherwise share an ID (same class + test name under different suites). - Repeats are collapsed — a test name that repeats the class as a dot- or
space-separated prefix appears once (
BF.ADD+BF.ADD transformArguments→BF.ADD transformArguments). - Names over 256 chars are truncated to
head...tail___hash8; the hash keeps truncated IDs unique (Mimir rejects label values over 2048 bytes). - Deterministic — the same test always generates the same ID.
Run-varying values in test names (UUIDs, timestamps, random ports, git SHAs,
memory addresses, temp paths) mint a new test_id series on every run — the
same churn removing per-run labels was meant to stop. The action scans generated
IDs for these patterns; every pattern was validated against the live fleet's
proven-stable test ids (zero churn measured) with zero false positives, so
stable names with fixed ports (localhost:6379), version ranges
([6] - [7.4.0]), argument lists or literal paths (/tmp/redis.sock) are not
flagged.
The on-nondeterministic-ids input decides what happens when the scan finds
offenders:
warn(default) — emit a workflow warning listing offenders and submit everything, exactly as before. The fix belongs in the test names (or the reporter's name template).skip— warn, drop the offenders' per-test data points, and submit the rest. Run and suite rollups still include the skipped tests: they carry notest.idlabel, so a churning name is no cardinality risk there, and dropping them would distort suite/run totals.break— list offenders, fail the build, upload nothing.
Since the detection is heuristic, a deliberately fixed value that merely looks
random (a hard-coded UUID or base64 literal in a test name) can be flagged; such
repos should stay on warn or rename the test.
Example Prometheus/Grafana queries for regression detection:
# Default-branch runs are labelled "default" (action >= 4.5.0); the master|main
# alternatives keep older action versions visible during the rollout.
# Baseline: average duration on default branch over 7 days
avg by (test_id, vcs_repository_name) (
avg_over_time(
cae_v16_test_duration_seconds{
vcs_repository_ref_name=~"master|main|default"
}[7d]
)
)
# Current: latest test duration
max by (test_id, vcs_repository_name) (
last_over_time(
cae_v16_test_duration_seconds{
vcs_repository_ref_name=~"master|main|default"
}[1h]
)
)
# Regression detection: current > 5x baseline
max by (test_id, vcs_repository_name) (
last_over_time(cae_v16_test_duration_seconds{vcs_repository_ref_name=~"master|main|default"}[1h])
)
> 5 * avg by (test_id, vcs_repository_name) (
avg_over_time(cae_v16_test_duration_seconds{vcs_repository_ref_name=~"master|main|default"}[7d])
)
# Cardinality churn: test ids first seen in the last day. Spikes after merges
# adding tests are normal; a persistently high value means test names are
# nondeterministic (the in-action warning should name the offenders).
count by (vcs_repository_name) (
last_over_time(cae_v16_test_duration_seconds[1d])
unless
last_over_time(cae_v16_test_duration_seconds[7d] offset 1d)
)
The action automatically extracts from GitHub context:
- Repository name (
owner/repo) - Branch name (labelled
defaultwhen it is the repository default branch) - Default branch (for branch gating and the
defaultlabel)
The commit SHA is logged in the action output for correlating a regression's timestamp with the commit that caused it, but is deliberately not a metric label (see cardinality note above).
No manual configuration needed for these values.
- JUnit XML files
- OTLP-compatible metrics backend
- Node.js 24+ runtime (provided by GitHub Actions)
v4 (v16) adds the optional server.version label so CI matrices that test
against multiple server versions get one clean series per test per version
instead of one mixed series per test:
| v3 (v15) | v4 (v16) |
|---|---|
| No matrix dimension | Optional server-version input → server.version |
cae_v15_* metric names |
cae_v16_* — new series start clean |
Update dashboards/alerts to query cae_v16_*. Duration baselines restart on
upgrade. If your CI runs a matrix against multiple server versions, pass the
version (e.g. server-version: ${{ matrix.redis-version }}) — otherwise the
matrix jobs keep writing into a single series per test.
v3 (v15) removes the per-run labels that churned one new series per test on
every CI run, and switches test_id to the full human-readable name:
| v2 (v13) | v3 (v15) |
|---|---|
ci_run_id label (new UUID / run) |
Removed — per-run values churn series |
vcs_repository_ref_revision label |
Removed — correlate commits via run timestamps / action logs |
test_id abbreviated + 6-char hash |
Full class.test name; truncated + hashed only over 256 chars |
| Emitted on every branch | Default branch only (branch-allowlist input to override) |
cae_v13_* metric name |
cae_v15_* — new series start clean |
All test_id values change on upgrade, so duration baselines restart from zero.
Update dashboards to query cae_v15_test_duration_seconds and drop any
ci_run_id / commit-SHA variables and matchers.
(A short-lived v14 schema shipped only in action v3.0.0: its test IDs kept XML
entities un-decoded and jest-junit duplication, and suites over 2,000 tests hit
the SDK cardinality limit. v3.0.1 fixed those and moved to v15 so the
malformed day-one series can be discarded wholesale.)
v2 uses a simplified, low-cardinality label set. Key changes:
| v1 (v12) | v2 (v13) |
|---|---|
service_name input required |
Auto-derived from repository |
service_namespace required |
Removed |
deployment_environment required |
Removed |
test_name label |
Folded into test_id |
test_class_name label |
Folded into test_id |
test_suite_name label |
Folded into test_id |
ci_run_id label |
Removed |
ci_job_id label |
Removed |
| 14+ labels | 4 labels |
Update your dashboard queries to use test_id instead of separate
name/class/suite labels.
- Processes all
.xmlfiles in the specified directory - Combines multiple XML files into a single report
- Handles malformed XML gracefully
- The OTel SDK cardinality limit is raised to 20,000 attribute sets (the SDK
default of 2,000 silently merges suites larger than 2k tests into a single
otel.metric.overflowdata point); the action warns if a report ever exceeds it - No outputs - metrics are the deliverable
Built for engineers who want observability without ceremony.