Skip to content

Bug: DeepSeek‑V4‑Flash with thinking mode enabled throws reasoning_content must be passed back to the API in multi‑turn tool‑call workflow #96

Description

@MCzhao2006

Environment

  • OS: Windows 11
  • Python: 3.11
  • LiteLLM: latest
  • LLM Backend: third‑party OpenAI‑compatible proxy(tokenrhythm.studio/v1)
  • Model: deepseek‑v4‑flash‑0731 (thinking mode enabled by default)
  • PentestAgent config (.env):
OPENAI_API_KEY=sk_tr_xxx
OPENAI_API_BASE=https://tokenrhythm.studio/v1/
PENTESTAGENT_MODEL=openai/deepseek‑v4‑flash‑0731
PENTESTAGENT_EMBEDDINGS=openai

Describe the bug

When running penetration‑testing tasks with multi‑turn tool calls, after the first round of model response, subsequent requests fail immediately:

litellm.BadRequestError: OpenAIException - The `reasoning_content` in the thinking mode must be passed back to the API.

Root cause analysis

  1. DeepSeek‑V4‑Flash thinking mode returns an extra reasoning_content field inside assistant message, which must be preserved and sent back within assistant message on next request.
  2. Current PentestAgent implementation in llm.py / memory.py discards the reasoning_content field when building / compressing conversation history.
  3. On the next API call, the reasoning_content is missing from assistant message payload, proxy returns 400 bad request.
  4. Adding litellm parameter preserve_reasoning_content into request dict does not fix this problem; it will cause another error unknown request field: preserve_reasoning_content, because this litellm‑specific parameter is mistakenly passed to upstream proxy.

Reproduction steps

  1. Configure .env to use deepseek‑v4‑flash‑0731 via third‑party compatible proxy.
  2. Start PentestAgent.
  3. Submit a penetration‑testing task that requires multiple rounds of tool calling.
  4. First LLM response succeeds; second round request triggers this error and agent stops.

Expected behavior

  • Either automatically preserve reasoning_content field for DeepSeek thinking‑enabled models in message history.
  • Or provide a config option to disable DeepSeek thinking mode via extra_body={"thinking":{"type":"disabled"}}.
  • Agent should continue multi‑turn workflow normally without 400 error.

Workaround available

Switch to non‑thinking model e.g. openai/gpt‑4o‑mini in .env, avoid DeepSeek‑V4‑Flash thinking mode.

Additional context

  • This issue only occurs on DeepSeek‑V4 series with thinking mode enabled, third‑party compatible proxy environment.
  • Clearing __pycache__ cannot resolve this bug, it is pure application‑level logic defect.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions