Conversation
Cluster deletion held the lifecycle write lock until postRemoveCluster finished. On dev builds the mutex watchdog SIGABRTs when an exclusive lock is held longer than 10s, which killed Central during TestClusterDeletion. BeginDeletion still waits for in-flight writers, then drops the lock. The deleting flag keeps new writes out until cleanup finishes. Request: scratch the previous branch contents and attempt a proper fix. Partially generated by AI.
|
Skipping CI for Draft Pull Request. |
|
/retest-times 5 ocp-4-12-nongroovy-e2e-tests |
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Central YAML (base), Organization UI (inherited) Review profile: CHILL Plan: Enterprise Run ID: 📒 Files selected for processing (2)
Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 11 remain after this review. 📝 SummarySummary by CodeRabbit
Walkthrough
ChangesDeletion gate
Priority: ➖ Normal Estimated code review effort: 2 (Simple) | ~10 minutes Change: Bug fix Suggested reviewers: Merge Risk: ⚪ Minimal · up to The change releases the lifecycle lock before slow cluster cleanup. This avoids the dev-build 10-second lock watchdog abort and keeps new writes blocked during deletion. No concrete merge-blocking risk remains. 🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
✨ Finishing Touches 💡 1📝 Generate docstrings 💡
🧪 Generate unit tests (beta)
Comment |
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## master #23210 +/- ##
==========================================
- Coverage 51.96% 51.93% -0.04%
==========================================
Files 2904 2904
Lines 182959 182958 -1
==========================================
- Hits 95081 95014 -67
- Misses 79564 79613 +49
- Partials 8314 8331 +17
Flags with carried forward coverage won't be shown. Click here to find out more. ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
🚀 Build Images ReadyImages are ready for commit e8f1a31. To use with deploy scripts: export MAIN_IMAGE_TAG=5.1.x-155-ge8f1a31708 |
|
/test ocp-4-12-nongroovy-e2e-tests |
2 similar comments
|
/test ocp-4-12-nongroovy-e2e-tests |
|
/test ocp-4-12-nongroovy-e2e-tests |
Description
Cluster deletion takes a lifecycle write lock so in-flight sensor writes finish before cleanup starts.
BeginDeletionheld that lock untilpostRemoveClusterreturned.Dev builds use a mutex watchdog (
pkg/sync,//go:build !release). If an exclusive lock is held longer than 10 seconds, unlock aborts the process with SIGABRT. Deleting a cluster's deployments, nodes, secrets, and the rest takes longer than that, so Central dies in the middle ofTestClusterDeletion. The test's next API call fails with connection refused. Release builds use the standard library mutex and do not abort, which is why this shows up on CI dev images. On this run the Central log isRWMutex.UnlockinBeginDeletiontaking more than 10s, thenSIGABRT, about 14 seconds after delete.Enteralready refuses a new lease oncedeletingis set, and it checks that flag before taking a read lock. The write lock only has to wait for leases that are already held.BeginDeletionnow waits for those, then drops the lock before returning. The function it returns still clearsdeletingwhen cleanup finishes, so new writes stay out for the whole cleanup without holding the lock across it.User-facing documentation
Testing and quality
Automated testing
How I validated my change
ocp-4-12-nongroovy-e2e-testsAI-Assisted: cursor, gate change and unit test drafted by the agent, logic reviewed in chat.