Skip to content

cgc api start: index jobs stay pending forever when the server is first created by a REST request #1744

Description

@CSchneiderHav

Summary

With cgc api start, POST /api/v1/index returns a job id but the job never runs.
check_job_status (via POST /api/v1/tools/call) reports pending indefinitely with
end_time: null and processed_files: 0; the process is idle. Other routes respond normally.

Cause

get_server() in api/router.py is a plain def used as a FastAPI dependency, so FastAPI runs
it in the threadpool. There is no running loop in that thread, so MCPServer.__init__ falls back
to asyncio.new_event_loop() (server.py L232-237), and nothing ever runs that loop.
add_code_to_graph then schedules the indexing coroutine on it with
asyncio.run_coroutine_threadsafe (tools/handlers/indexing_handlers.py L89; add_package_to_graph
L155 is the same), so the coroutine never starts.

This depends on which request builds the singleton first. If the first call comes through the
MCP-over-SSE handlers, which call get_server() on the loop, the server binds the right loop and
jobs run. Any REST route first → jobs never run.

Side effect: concurrent first REST requests each build their own MCPServer in separate worker
threads. In a local test, 8 parallel /api/v1/status calls created 8 instances, each with its own
DB driver and JobManager, so a job id from one can come back missing from another.

Environment

codegraphcontext 0.6.13 (router/server/handlers unchanged on main as of 2026-09-29),
Python 3.12, FastAPI 0.141. Reproduced with the Kùzu backend and with Neo4j.

Steps

  1. cgc api start
  2. POST /api/v1/index {"path": "<repo under cwd>"} → 200 with a job id
  3. POST /api/v1/tools/call {"name":"check_job_status","arguments":{"job_id":"…"}} → pending forever

Fix (PR)

Make get_server async def so it runs on the serving loop, and pass that loop explicitly.
The two direct callers in api/mcp_sse.py now await get_server(). This also removes the
first-request race. A regression test checks that the server is bound to the serving loop, and
that SSE list_tools still gets a server.

Known trade-off: MCPServer.__init__ (context resolution, DB driver creation) now runs once on
the loop during the first request. An alternative is building the server in a lifespan hook,
which also gives a place to call shutdown(), but then the gateway refuses to start while the
DB is down. Happy to switch if you prefer that.


Disclosure: the investigation and this report were prepared with AI assistance (Claude Code) and reviewed by a human before posting.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions