Describe the bug
The LoLLMs binding reads response bodies without checking HTTP status. If /lollms_generate returns HTTP 401, 429 or 503, lollms_model_if_cache() returns the error body as a normal completion, or yields it as normal output with stream=True.
The embedding path has the same missing status check: an error JSON body becomes KeyError: 'vector', an HTML error becomes a JSON content-type error, and even a non-success response containing a vector field is accepted.
Reproduction
Reproduced on current main 6702eea912b12df5ba831f58609470ca9da34cfc (LightRAG 1.5.8), Windows, Python 3.13.12 and locked aiohttp 3.13.5. No model or API credentials are needed:
import asyncio
from aiohttp import web
from aiohttp.test_utils import TestServer
from lightrag.llm.lollms import lollms_model_if_cache
async def main():
async def unavailable(request):
await request.read()
return web.Response(status=503, text="upstream unavailable")
app = web.Application()
app.router.add_post("/lollms_generate", unavailable)
async with TestServer(app) as server:
base_url = str(server.make_url("/")).rstrip("/")
result = await lollms_model_if_cache("test-model", "hello", base_url=base_url)
print(repr(result))
stream = await lollms_model_if_cache(
"test-model", "hello", base_url=base_url, stream=True
)
print([chunk async for chunk in stream])
asyncio.run(main())
Actual output:
'upstream unavailable'
['upstream unavailable']
Expected: raise an HTTP response error carrying status 503, before returning or yielding the error body as model output.
Scope and proposed fix
The affected code is lightrag/llm/lollms.py, covering regular generation, streamed generation and embeddings. Enable aiohttp's raise_for_status on these sessions so HTTP errors propagate before response consumption.
Regression tests against a local HTTP server reproduce all three paths. Before the fix, 9 error cases fail and 3 success controls pass; with the fix, those 12 cases and the 7 existing LoLLMs tests pass.
This was found through source review and reproduced with a local HTTP fixture, not a live LoLLMs deployment.
Describe the bug
The LoLLMs binding reads response bodies without checking HTTP status. If
/lollms_generatereturns HTTP 401, 429 or 503,lollms_model_if_cache()returns the error body as a normal completion, or yields it as normal output withstream=True.The embedding path has the same missing status check: an error JSON body becomes
KeyError: 'vector', an HTML error becomes a JSON content-type error, and even a non-success response containing avectorfield is accepted.Reproduction
Reproduced on current main
6702eea912b12df5ba831f58609470ca9da34cfc(LightRAG 1.5.8), Windows, Python 3.13.12 and locked aiohttp 3.13.5. No model or API credentials are needed:Actual output:
Expected: raise an HTTP response error carrying status 503, before returning or yielding the error body as model output.
Scope and proposed fix
The affected code is
lightrag/llm/lollms.py, covering regular generation, streamed generation and embeddings. Enable aiohttp'sraise_for_statuson these sessions so HTTP errors propagate before response consumption.Regression tests against a local HTTP server reproduce all three paths. Before the fix, 9 error cases fail and 3 success controls pass; with the fix, those 12 cases and the 7 existing LoLLMs tests pass.
This was found through source review and reproduced with a local HTTP fixture, not a live LoLLMs deployment.