Conversation
a0c1437 to
755e967
Compare
|
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## master #6855 +/- ##
==========================================
+ Coverage 48.04% 48.09% +0.05%
==========================================
Files 427 427
Lines 53591 53612 +21
Branches 7800 7803 +3
==========================================
+ Hits 25749 25786 +37
+ Misses 25986 25969 -17
- Partials 1856 1857 +1
Continue to review full report in Codecov by Harness.
🚀 New features to boost your workflow:
|
jyejare
left a comment
There was a problem hiding this comment.
The PR cleanly isolates Prometheus scrape aggregation in a dedicated process and preserves the existing in-process monitoring behavior. The documentation and unit tests cover the process construction and monitoring flags well, but the child shutdown path can deadlock because BaseServer.shutdown() is called from the same thread running serve_forever(). Startup failures and lifecycle management are also not surfaced to the parent, which can leave the service reporting metrics as started when the endpoint is unavailable.
ca670d8 to
b1dff9b
Compare
c3a6ca7 to
596d1ac
Compare
Signed-off-by: ntkathole <nikhilkathole2683@gmail.com>
596d1ac to
b4afb2d
Compare
Summary
Replace the daemon thread that serves the Prometheus
/metricsendpoint with amultiprocessing.Processso that scrape-time aggregation (MultiProcessCollectorreads + text serialization) runs with its own GIL, fully isolated from request-serving workers and the Gunicorn master.Problem
When Prometheus scrapes the
/metricsendpoint, the current daemon thread performs:.dbfiles inPROMETHEUS_MULTIPROCESS_DIRAll of this runs under the same GIL as the Gunicorn master process. With high metric cardinality (many label combinations), this scrape overhead can cause GIL contention that indirectly affects request latency.
Solution
Move the WSGI HTTP metrics server into a dedicated child process via
multiprocessing.Process(daemon=True):_run_metrics_server()— new top-level function that serves as the child process entry point. It re-importsprometheus_clientfor macOSspawnsafety and registersSIGTERM/SIGINThandlers for graceful shutdown.start_metrics_server()— now spawns aProcessinstead of aThread. The function signature and call sites are unchanged.Why this is safe
:8000in master process:8000in child processdaemon=TrueMultiProcessCollectorreads same mmap filesspawnsafety_run_metrics_serveris top-level, re-imports deps