Skip to content

Let a newer Deploy Docs run cancel the one in progress - #3633

Merged
maxisbey merged 1 commit into
mainfrom
ci-docs-deploy-cancel-in-progress
Oct 2, 2026
Merged

maxisbey merged 1 commit into
mainfrom
ci-docs-deploy-cancel-in-progress

Conversation

@maxisbey

@maxisbey maxisbey commented Oct 2, 2026

Copy link
Copy Markdown
Contributor

On 2026-10-01 GitHub left one Deploy Docs run queued and never gave it a runner. It held the deploy-docs concurrency group for 25 hours, so each later run sat pending until the next merge replaced it, and the site missed nearly five hours of merges until the stuck run was cancelled by hand.

This sets cancel-in-progress: true on that group, so a new run cancels the run holding the group instead of waiting behind it. Nothing else in the workflow changes.

Cancelling a deploy part-way is safe in this workflow:

  • scripts/build-docs.sh fetches and builds the tips of main and v1.x, not the triggering commit, so a newer run publishes everything the cancelled one would have.
  • The site goes to Pages as one artifact in one deployment request, so a cancelled run leaves the live site as it was. deploy-pages also cancels its own pending deployment when the job is cancelled.
  • No other step leaves anything behind. The artifact belongs to its own run, configure-pages only reads, and the uv cache is saved only after a successful job.
  • No other workflow uses the deploy-docs group or the github-pages environment.

The cost is a wasted build when two merges land within about two minutes of each other: the first run is cancelled and the second publishes both. If the second commit breaks the docs build, the first one's content waits for the fix as well, which the docs check on pull requests is there to catch.

One point hasn't been observed. GitHub's docs describe cancel-in-progress as cancelling a "currently running" job or workflow in the group. The stuck run held the group while queued with no runner, and it isn't confirmed that such a run is cancelled the same way. If it isn't, the behaviour is the same as today.

Checked with prettier, zizmor and actionlint. The workflow file is in its own paths filter, so merging this runs a deploy.

AI Disclaimer

A run that never gets a runner held the deploy-docs concurrency group, so
every later deploy waited behind it and the site went stale. Each run builds
the tips of main and v1.x, so the newest run publishes everything.

@cubic-dev-ai cubic-dev-ai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

1 issue found across 1 file

Prompt for AI agents (unresolved issues)

Check if these issues are valid — if so, understand the root cause of each and fix them. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. If appropriate, use sub-agents to investigate and fix each issue separately.


<file name=".github/workflows/deploy-docs.yml">

<violation number="1" location=".github/workflows/deploy-docs.yml:34">
P2: This cancels runs even after `Deploy to GitHub Pages` has started, but cancellation does not roll back a deployment request already submitted by `deploy-pages`; an older artifact can therefore become live after a newer run starts. Separate the cancelable build from deployment and serialize the deployment phase, or otherwise prevent cancellation once publication begins.</violation>
</file>

Reply with feedback, questions, or to request a fix.

Re-trigger cubic

# and Pages takes the site as one artifact, so a cancelled run leaves the live
# site as it was. With `false`, a run that never gets a runner holds the group
# and every later deploy waits behind it.
cancel-in-progress: true

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2: This cancels runs even after Deploy to GitHub Pages has started, but cancellation does not roll back a deployment request already submitted by deploy-pages; an older artifact can therefore become live after a newer run starts. Separate the cancelable build from deployment and serialize the deployment phase, or otherwise prevent cancellation once publication begins.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. At .github/workflows/deploy-docs.yml, line 34:

<comment>This cancels runs even after `Deploy to GitHub Pages` has started, but cancellation does not roll back a deployment request already submitted by `deploy-pages`; an older artifact can therefore become live after a newer run starts. Separate the cancelable build from deployment and serialize the deployment phase, or otherwise prevent cancellation once publication begins.</comment>

<file context>
@@ -26,7 +26,12 @@ on:
+  # and Pages takes the site as one artifact, so a cancelled run leaves the live
+  # site as it was. With `false`, a run that never gets a runner holds the group
+  # and every later deploy waits behind it.
+  cancel-in-progress: true
 
 jobs:
</file context>

@maxisbey
maxisbey merged commit 4d29994 into main Oct 2, 2026
36 checks passed
@maxisbey
maxisbey deleted the ci-docs-deploy-cancel-in-progress branch October 2, 2026 17:51

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nothing blocking. The comments below are optional suggestions. There is no need to push a fix for them before merging.

Beyond the inline finding, I also checked whether a cancelled run could drop the triggering commit's docs content: scripts/build-docs.sh fetches and builds origin/main and origin/v1.x tips rather than the checked-out ref, so the comment's claim holds and the next completed run carries everything a cancelled one would have.

Extended reasoning...

The change flips cancel-in-progress to true on the deploy-docs concurrency group in .github/workflows/deploy-docs.yml and adds a rationale comment; no code, tests, or security-sensitive surface is touched. The build script was read to confirm it builds branch tips from origin, so the stated safety assumption is accurate. One inline finding about stale-site windows during merge bursts remains for the author to weigh; no CODEOWNERS file exists in the repo.

# and Pages takes the site as one artifact, so a cancelled run leaves the live
# site as it was. With `false`, a run that never gets a runner holds the group
# and every later deploy waits behind it.
cancel-in-progress: true

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 (optional) Docs readers get a site that stays stale through every merge burst, where the base branch published after each run. .github/workflows/deploy-docs.yml:34 cancels the running deploy on every push to main matching the paths filter, which includes src/mcp/**. The build runs scripts/build-docs.sh: two branch builds plus 12 language sites, so a run lasts longer than the gap between merges in this repo (git log shows merges at 13:45, 13:50, 13:51 and 14:17, 14:24, 14:32 on 2026-10-02). Fix: keep cancel-in-progress: false and bound the stuck-queue case another way (a job-level timeout-minutes, or cancel only runs still queued), so a deploy that has started always finishes.

Why this was flagged

git log on this checkout shows 17 merges on 2026-10-02, with clusters three merges in six minutes (13:45:32, 13:50:29, 13:51:05) and three in fifteen minutes (14:17, 14:24, 14:32); nearly all touch src/mcp/** or docs/, which are in the push paths filter at .github/workflows/deploy-docs.yml:9-24. Each run executes scripts/build-docs.sh, which fetches and builds main via scripts/docs/build.sh (uv sync, zensical strict build, then one zensical build per language for 12 languages per i18n/languages.yml) and then v1.x via a second uv sync and mkdocs build. That is well over the five-minute merge gaps observed. With .github/workflows/deploy-docs.yml:34 set to true, every merge in such a cluster cancels the run in progress, so no deployment completes until the cluster ends; on the base branch each queued run waited and the site advanced after every build. Every reader of the published docs sees content lagging by the length of the burst, and each cancelled run burns a full runner build. Remedy: leave the in-progress run alone and address the stuck-queued-runner incident with a job timeout or by cancelling only queued runs.

Verification: cancel-in-progress: true (.github/workflows/deploy-docs.yml:34) makes each new push cancel the in-progress run before its deploy-pages step, so nothing publishes until the last run of the burst completes; on the base (cancel-in-progress: false) the in-progress run finishes and publishes. git log shows merges at 2026-10-02 13:45:32, 13:50:29, 13:51:05.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant