fix(serve): log the inference failure the client is not allowed to see - #375
Merged
NandhaKishorM merged 1 commit intoSep 24, 2026
Merged
Conversation
A failed inference is reported to the client as a fixed `500 {"detail":"inference failed"}`
so that nothing about paths, weights or memory state leaks. That leaves the server log as
the only place the cause can appear, and there was nothing in it:
```
$ curl -s -X POST localhost:8000/v1/systemone -d @req.json
{"detail":"inference failed"}
# container log, before
POST /v1/systemone HTTP/1.1" 500 Internal Server Error
```
`NandhaKishorM#365` is what that costs. A `TORCH_INDEX=cuXXX` image has no C compiler, and triton — which
the CUDA torch wheel brings with it — JIT-compiles its driver module on the *first*
inference, so the container starts, loads weights, reports `{"status":"ok","device":"cuda"}`
on `/health`, and then fails every request. Diagnosing it meant reproducing `Router().predict`
in-process to get the actual error, because the running server never said it:
```
RuntimeError: Failed to find C compiler. Please specify via CC environment variable
```
Now the same failure is one line in the log, with the traceback:
```
ERROR laya.serve: inference failed for model=None
Traceback (most recent call last):
...
RuntimeError: Failed to find C compiler. Please specify via CC environment variable
```
and the client response is byte-for-byte what it was.
Deliberate choices tested:
- **`_log.exception`, not `_log.error`.** The traceback is the whole point: the failure in `NandhaKishorM#365` happens three frames down inside triton, so a message-only log would name the symptom and still hide the cause.
- **A module logger named `laya.serve`, with no handlers and no `basicConfig`.** Uvicorn configures the root logger, so this propagates there and picks up whatever the operator already set, including `LAYA_LOG_LEVEL`. A library that installs its own handler changes the application's logging without being asked.
- **Only the bare `except Exception` logs.** `ValueError` above it becomes a `422` with the message intact, because those messages name the question and what to fix — that is the caller's mistake, not a server fault, and logging it at ERROR would cry wolf on ordinary bad input. `HTTPException` is re-raised untouched. Both are asserted.
- **The client-facing text is unchanged.** That is the property the `except` exists to protect, so the test asserts the absence of `C compiler`, `triton`, `/opt/venv` and `site-packages` in the response body as well as the exact 500 payload.
- **A stdlib import at module scope is fine here.** `test_lazy_import.py` guards `import laya` and `import laya.serve` against pulling in torch; `logging` does not, and the suite still passes.
This is the observability half of `NandhaKishorM#365`, not the whole fix. The image still needs a compiler
for triton, which is the reporter's own suggested `gcc g++ libc6-dev` addition to the runtime
stage — that is a Dockerfile change I cannot build or test on this machine (no Docker, no
CUDA), so it is left alone rather than shipped unverified. What this does is make such a
failure self-diagnosing, which is the second of the three things `NandhaKishorM#365` asked for.
The 500 path had no test of any kind before this — `tests/test_serve.py` covered 400s, auth,
the body limit and the event-loop offload, but never a failing inference.
```
$ python -m pytest tests/test_serve.py -q 24 passed (22 before)
$ python tests/test_lazy_import.py all lazy-import tests passed
$ python tests/test_packaging.py all packaging tests passed
$ python -m ruff check laya/ --select=E9,F63,F7,F82,F401,F811 --line-length=120
All checks passed!
$ python -m compileall -q laya/ tests/
```
Known limitation: the log line goes wherever the root logger points, so a deployment that
configures no logging at all still sees nothing. That is deliberate — uvicorn sets one up by
default, and inventing a handler in a library is worse than inheriting the application's.
| # this the container logs show only the 500, so a deterministic failure such as a | ||
| # missing C compiler for triton's JIT (#365) is invisible from the running server | ||
| # and has to be reproduced in-process to be diagnosed at all. | ||
| _log.exception("inference failed for model=%s", model) |
Owner
|
Thank you, this is a small change with a big payoff for anyone running laya-serve. Failures like the missing C compiler in #365 now show up in the log with a full traceback, while the client response stays exactly as it was, and the tests check both the non-leak and that 422s are not logged. Merging. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
fix(serve): log the inference failure the client is not allowed to see
A failed inference is reported to the client as a fixed
500 {"detail":"inference failed"}so that nothing about paths, weights or memory state leaks. That leaves the server log as
the only place the cause can appear, and there was nothing in it:
#365is what that costs. ATORCH_INDEX=cuXXXimage has no C compiler, and triton — whichthe CUDA torch wheel brings with it — JIT-compiles its driver module on the first
inference, so the container starts, loads weights, reports
{"status":"ok","device":"cuda"}on
/health, and then fails every request. Diagnosing it meant reproducingRouter().predictin-process to get the actual error, because the running server never said it:
Now the same failure is one line in the log, with the traceback:
and the client response is byte-for-byte what it was.
Deliberate choices tested:
_log.exception, not_log.error. The traceback is the whole point: the failure in#365happens three frames down inside triton, so a message-only log would name the symptom and still hide the cause.laya.serve, with no handlers and nobasicConfig. Uvicorn configures the root logger, so this propagates there and picks up whatever the operator already set, includingLAYA_LOG_LEVEL. A library that installs its own handler changes the application's logging without being asked.except Exceptionlogs.ValueErrorabove it becomes a422with the message intact, because those messages name the question and what to fix — that is the caller's mistake, not a server fault, and logging it at ERROR would cry wolf on ordinary bad input.HTTPExceptionis re-raised untouched. Both are asserted.exceptexists to protect, so the test asserts the absence ofC compiler,triton,/opt/venvandsite-packagesin the response body as well as the exact 500 payload.test_lazy_import.pyguardsimport layaandimport laya.serveagainst pulling in torch;loggingdoes not, and the suite still passes.This is the observability half of
#365, not the whole fix. The image still needs a compilerfor triton, which is the reporter's own suggested
gcc g++ libc6-devaddition to the runtimestage — that is a Dockerfile change I cannot build or test on this machine (no Docker, no
CUDA), so it is left alone rather than shipped unverified. What this does is make such a
failure self-diagnosing, which is the second of the three things
#365asked for.The 500 path had no test of any kind before this —
tests/test_serve.pycovered 400s, auth,the body limit and the event-loop offload, but never a failing inference.
Known limitation: the log line goes wherever the root logger points, so a deployment that
configures no logging at all still sees nothing. That is deliberate — uvicorn sets one up by
default, and inventing a handler in a library is worse than inheriting the application's.