Environment
- OS: Windows 11
- Python: 3.11
- LiteLLM: latest
- LLM Backend: third‑party OpenAI‑compatible proxy(tokenrhythm.studio/v1)
- Model:
deepseek‑v4‑flash‑0731 (thinking mode enabled by default)
- PentestAgent config (.env):
OPENAI_API_KEY=sk_tr_xxx
OPENAI_API_BASE=https://tokenrhythm.studio/v1/
PENTESTAGENT_MODEL=openai/deepseek‑v4‑flash‑0731
PENTESTAGENT_EMBEDDINGS=openai
Describe the bug
When running penetration‑testing tasks with multi‑turn tool calls, after the first round of model response, subsequent requests fail immediately:
litellm.BadRequestError: OpenAIException - The `reasoning_content` in the thinking mode must be passed back to the API.
Root cause analysis
- DeepSeek‑V4‑Flash thinking mode returns an extra
reasoning_content field inside assistant message, which must be preserved and sent back within assistant message on next request.
- Current PentestAgent implementation in
llm.py / memory.py discards the reasoning_content field when building / compressing conversation history.
- On the next API call, the
reasoning_content is missing from assistant message payload, proxy returns 400 bad request.
- Adding litellm parameter
preserve_reasoning_content into request dict does not fix this problem; it will cause another error unknown request field: preserve_reasoning_content, because this litellm‑specific parameter is mistakenly passed to upstream proxy.
Reproduction steps
- Configure
.env to use deepseek‑v4‑flash‑0731 via third‑party compatible proxy.
- Start PentestAgent.
- Submit a penetration‑testing task that requires multiple rounds of tool calling.
- First LLM response succeeds; second round request triggers this error and agent stops.
Expected behavior
- Either automatically preserve
reasoning_content field for DeepSeek thinking‑enabled models in message history.
- Or provide a config option to disable DeepSeek thinking mode via
extra_body={"thinking":{"type":"disabled"}}.
- Agent should continue multi‑turn workflow normally without 400 error.
Workaround available
Switch to non‑thinking model e.g. openai/gpt‑4o‑mini in .env, avoid DeepSeek‑V4‑Flash thinking mode.
Additional context
- This issue only occurs on DeepSeek‑V4 series with thinking mode enabled, third‑party compatible proxy environment.
- Clearing
__pycache__ cannot resolve this bug, it is pure application‑level logic defect.
Environment
deepseek‑v4‑flash‑0731(thinking mode enabled by default)Describe the bug
When running penetration‑testing tasks with multi‑turn tool calls, after the first round of model response, subsequent requests fail immediately:
Root cause analysis
reasoning_contentfield inside assistant message, which must be preserved and sent back within assistant message on next request.llm.py/memory.pydiscards thereasoning_contentfield when building / compressing conversation history.reasoning_contentis missing from assistant message payload, proxy returns 400 bad request.preserve_reasoning_contentinto request dict does not fix this problem; it will cause another errorunknown request field: preserve_reasoning_content, because this litellm‑specific parameter is mistakenly passed to upstream proxy.Reproduction steps
.envto usedeepseek‑v4‑flash‑0731via third‑party compatible proxy.Expected behavior
reasoning_contentfield for DeepSeek thinking‑enabled models in message history.extra_body={"thinking":{"type":"disabled"}}.Workaround available
Switch to non‑thinking model e.g.
openai/gpt‑4o‑miniin.env, avoid DeepSeek‑V4‑Flash thinking mode.Additional context
__pycache__cannot resolve this bug, it is pure application‑level logic defect.