Skip to content

fix(serve): log the inference failure the client is not allowed to see - #375

Merged
NandhaKishorM merged 1 commit into
NandhaKishorM:mainfrom
PerryLink:fix/serve-log-inference-failure
Sep 24, 2026
Merged

NandhaKishorM merged 1 commit into
NandhaKishorM:mainfrom
PerryLink:fix/serve-log-inference-failure

Conversation

@PerryLink

Copy link
Copy Markdown
Contributor

fix(serve): log the inference failure the client is not allowed to see

A failed inference is reported to the client as a fixed 500 {"detail":"inference failed"}
so that nothing about paths, weights or memory state leaks. That leaves the server log as
the only place the cause can appear, and there was nothing in it:

$ curl -s -X POST localhost:8000/v1/systemone -d @req.json
{"detail":"inference failed"}

# container log, before
POST /v1/systemone HTTP/1.1" 500 Internal Server Error

#365 is what that costs. A TORCH_INDEX=cuXXX image has no C compiler, and triton — which
the CUDA torch wheel brings with it — JIT-compiles its driver module on the first
inference, so the container starts, loads weights, reports {"status":"ok","device":"cuda"}
on /health, and then fails every request. Diagnosing it meant reproducing Router().predict
in-process to get the actual error, because the running server never said it:

RuntimeError: Failed to find C compiler. Please specify via CC environment variable

Now the same failure is one line in the log, with the traceback:

ERROR laya.serve: inference failed for model=None
Traceback (most recent call last):
  ...
RuntimeError: Failed to find C compiler. Please specify via CC environment variable

and the client response is byte-for-byte what it was.

Deliberate choices tested:

  • _log.exception, not _log.error. The traceback is the whole point: the failure in #365 happens three frames down inside triton, so a message-only log would name the symptom and still hide the cause.
  • A module logger named laya.serve, with no handlers and no basicConfig. Uvicorn configures the root logger, so this propagates there and picks up whatever the operator already set, including LAYA_LOG_LEVEL. A library that installs its own handler changes the application's logging without being asked.
  • Only the bare except Exception logs. ValueError above it becomes a 422 with the message intact, because those messages name the question and what to fix — that is the caller's mistake, not a server fault, and logging it at ERROR would cry wolf on ordinary bad input. HTTPException is re-raised untouched. Both are asserted.
  • The client-facing text is unchanged. That is the property the except exists to protect, so the test asserts the absence of C compiler, triton, /opt/venv and site-packages in the response body as well as the exact 500 payload.
  • A stdlib import at module scope is fine here. test_lazy_import.py guards import laya and import laya.serve against pulling in torch; logging does not, and the suite still passes.

This is the observability half of #365, not the whole fix. The image still needs a compiler
for triton, which is the reporter's own suggested gcc g++ libc6-dev addition to the runtime
stage — that is a Dockerfile change I cannot build or test on this machine (no Docker, no
CUDA), so it is left alone rather than shipped unverified. What this does is make such a
failure self-diagnosing, which is the second of the three things #365 asked for.

The 500 path had no test of any kind before this — tests/test_serve.py covered 400s, auth,
the body limit and the event-loop offload, but never a failing inference.

$ python -m pytest tests/test_serve.py -q          24 passed   (22 before)
$ python tests/test_lazy_import.py                 all lazy-import tests passed
$ python tests/test_packaging.py                   all packaging tests passed
$ python -m ruff check laya/ --select=E9,F63,F7,F82,F401,F811 --line-length=120
All checks passed!
$ python -m compileall -q laya/ tests/

Known limitation: the log line goes wherever the root logger points, so a deployment that
configures no logging at all still sees nothing. That is deliberate — uvicorn sets one up by
default, and inventing a handler in a library is worse than inheriting the application's.

A failed inference is reported to the client as a fixed `500 {"detail":"inference failed"}`
so that nothing about paths, weights or memory state leaks. That leaves the server log as
the only place the cause can appear, and there was nothing in it:

```
$ curl -s -X POST localhost:8000/v1/systemone -d @req.json
{"detail":"inference failed"}

# container log, before
POST /v1/systemone HTTP/1.1" 500 Internal Server Error
```

`NandhaKishorM#365` is what that costs. A `TORCH_INDEX=cuXXX` image has no C compiler, and triton — which
the CUDA torch wheel brings with it — JIT-compiles its driver module on the *first*
inference, so the container starts, loads weights, reports `{"status":"ok","device":"cuda"}`
on `/health`, and then fails every request. Diagnosing it meant reproducing `Router().predict`
in-process to get the actual error, because the running server never said it:

```
RuntimeError: Failed to find C compiler. Please specify via CC environment variable
```

Now the same failure is one line in the log, with the traceback:

```
ERROR laya.serve: inference failed for model=None
Traceback (most recent call last):
  ...
RuntimeError: Failed to find C compiler. Please specify via CC environment variable
```

and the client response is byte-for-byte what it was.

Deliberate choices tested:

- **`_log.exception`, not `_log.error`.** The traceback is the whole point: the failure in `NandhaKishorM#365` happens three frames down inside triton, so a message-only log would name the symptom and still hide the cause.
- **A module logger named `laya.serve`, with no handlers and no `basicConfig`.** Uvicorn configures the root logger, so this propagates there and picks up whatever the operator already set, including `LAYA_LOG_LEVEL`. A library that installs its own handler changes the application's logging without being asked.
- **Only the bare `except Exception` logs.** `ValueError` above it becomes a `422` with the message intact, because those messages name the question and what to fix — that is the caller's mistake, not a server fault, and logging it at ERROR would cry wolf on ordinary bad input. `HTTPException` is re-raised untouched. Both are asserted.
- **The client-facing text is unchanged.** That is the property the `except` exists to protect, so the test asserts the absence of `C compiler`, `triton`, `/opt/venv` and `site-packages` in the response body as well as the exact 500 payload.
- **A stdlib import at module scope is fine here.** `test_lazy_import.py` guards `import laya` and `import laya.serve` against pulling in torch; `logging` does not, and the suite still passes.

This is the observability half of `NandhaKishorM#365`, not the whole fix. The image still needs a compiler
for triton, which is the reporter's own suggested `gcc g++ libc6-dev` addition to the runtime
stage — that is a Dockerfile change I cannot build or test on this machine (no Docker, no
CUDA), so it is left alone rather than shipped unverified. What this does is make such a
failure self-diagnosing, which is the second of the three things `NandhaKishorM#365` asked for.

The 500 path had no test of any kind before this — `tests/test_serve.py` covered 400s, auth,
the body limit and the event-loop offload, but never a failing inference.

```
$ python -m pytest tests/test_serve.py -q          24 passed   (22 before)
$ python tests/test_lazy_import.py                 all lazy-import tests passed
$ python tests/test_packaging.py                   all packaging tests passed
$ python -m ruff check laya/ --select=E9,F63,F7,F82,F401,F811 --line-length=120
All checks passed!
$ python -m compileall -q laya/ tests/
```

Known limitation: the log line goes wherever the root logger points, so a deployment that
configures no logging at all still sees nothing. That is deliberate — uvicorn sets one up by
default, and inventing a handler in a library is worse than inheriting the application's.
Comment thread laya/serve.py
# this the container logs show only the 500, so a deterministic failure such as a
# missing C compiler for triton's JIT (#365) is invisible from the running server
# and has to be reproduced in-process to be diagnosed at all.
_log.exception("inference failed for model=%s", model)
@NandhaKishorM

Copy link
Copy Markdown
Owner

Thank you, this is a small change with a big payoff for anyone running laya-serve. Failures like the missing C compiler in #365 now show up in the log with a full traceback, while the client response stays exactly as it was, and the tests check both the non-leak and that 422s are not logged. Merging.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants