Issue
Title
OpenAI-compatible chat adapter ignores configured timeout and hard-codes 120 seconds
Description
CECLI appears to ignore the configured timeout value for OpenAI-compatible chat providers.
I am using a custom OpenAI-compatible provider with a long-running non-streaming model response.
My config includes:
timeout: 3000
model-providers:
very_slow_response_model:
api_base: "http://192.168.10.89:9000/v1"
requires_api_key: false
supports_stream: false
CECLI debug output confirms that the configured timeout is propagated into the request kwargs:
However, the request is abandoned at approximately 120 seconds, after which CECLI prints:
Retrying in 0.2 seconds...
The upstream proxy is still processing the request at that point and later completes successfully.
Example timing:
t = 0s
Request sent from CECLI
t ≈ 120s
CECLI:
Retrying in 0.2 seconds...
t ≈ 160s
Upstream proxy:
claude.completed duration_ms=160324 exit_code=0
chat_completion_response model=xxx
usage | xxx | in 32078 out 13033
The upstream proxy itself has a 300-second timeout, so it is not the source of the 120-second cutoff.
I traced this to:
cecli/helpers/llms/domains/chat.py
which contains:
and uses it directly in both the non-streaming and streaming paths:
async with make_client(timeout=DEFAULT_TIMEOUT, verify=VERIFY_SSL) as client:
resp = await client.post(...)
This means the higher-level configured timeout is effectively ignored by the OpenAI-compatible chat adapter.
Expected behavior
The adapter should honor the configured/request timeout, for example:
timeout = float(kwargs.get("timeout", DEFAULT_TIMEOUT))
async with make_client(timeout=timeout, verify=VERIFY_SSL) as client:
...
This should be applied to both:
chat_complete()
chat_stream()
Secondary retry issue
After the 120-second timeout triggers a retry, CECLI also crashes in the retry path with:
TypeError: cannot pickle '_thread.RLock' object
Traceback points to:
cecli/models.py
cecli/helpers/config_utils.py
where deep_merge() calls copy.deepcopy() on request kwargs:
kwargs = deep_merge(kwargs, {"allowed_openai_params": ["tools", "tool_choice"]})
and:
merged = copy.deepcopy(dict1)
So there appear to be two separate issues:
- OpenAI-compatible chat requests are hard-limited to 120 seconds despite a larger configured
timeout.
- The retry path can fail with
_thread.RLock during deepcopy().
Environment
CECLI version: 1.6.1
Python: virtual environment on Windows
Provider type: custom OpenAI-compatible
Streaming: disabled
Configured timeout: 3000 seconds
Observed client cutoff: ~120 seconds
Upstream successful completion: ~160 seconds
The same provider/model works correctly for shorter requests.
Version and model info
No response
Issue
Title
OpenAI-compatible chat adapter ignores configured
timeoutand hard-codes 120 secondsDescription
CECLI appears to ignore the configured
timeoutvalue for OpenAI-compatible chat providers.I am using a custom OpenAI-compatible provider with a long-running non-streaming model response.
My config includes:
CECLI debug output confirms that the configured timeout is propagated into the request kwargs:
However, the request is abandoned at approximately 120 seconds, after which CECLI prints:
The upstream proxy is still processing the request at that point and later completes successfully.
Example timing:
The upstream proxy itself has a 300-second timeout, so it is not the source of the 120-second cutoff.
I traced this to:
which contains:
and uses it directly in both the non-streaming and streaming paths:
This means the higher-level configured timeout is effectively ignored by the OpenAI-compatible chat adapter.
Expected behavior
The adapter should honor the configured/request timeout, for example:
This should be applied to both:
chat_complete()chat_stream()Secondary retry issue
After the 120-second timeout triggers a retry, CECLI also crashes in the retry path with:
Traceback points to:
where
deep_merge()callscopy.deepcopy()on request kwargs:and:
So there appear to be two separate issues:
timeout._thread.RLockduringdeepcopy().Environment
The same provider/model works correctly for shorter requests.
Version and model info
No response