Skip to content

reasoningEffort silently drops temperature for all 38 openai_compat providers, not just o1/o3/o4 #6002

Description

@GZY-SUPER-HACKER

Summary

Setting agents.defaults.reasoningEffort to anything other than null / "none" makes
nanobot stop sending temperature — for every provider, not only the reasoning models the
rule was written for.

registry.py defines 46 ProviderSpecs and 38 of them use backend="openai_compat", so
they all pass through the same predicate: DeepSeek, Zhipu, DashScope, Moonshot, Gemini,
OpenRouter, SiliconFlow, VolcEngine, Groq, Ollama, vLLM, LM Studio, and every custom
OpenAI-compatible endpoint.

The failure is silent: no warning, no log line, and the configured temperature stays
visible in config.json. The request simply goes out without it and the server applies its
own default. The only user-visible effect is that output randomness changes after a
reasoningEffort value is set.

1 · What actually goes out

Wrapping the provider's AsyncOpenAI client and logging the outbound kwargs, same config,
only reasoningEffort changed:

reasoningEffort = null
  responses.create kwargs: [..., 'store', 'stream', 'temperature', 'timeout', 'tool_choice', 'tools']
                                                        ^^^^^^^^^^^ present

reasoningEffort = "high"
  responses.create kwargs: [..., 'reasoning', 'store', 'stream', 'timeout', 'tool_choice', 'tools']
                                 ^^^^^^^^^ reasoning={"effort":"high"}     ^^^^^^^^^^^ gone

temperature is not rewritten or clamped — it is absent from the request body entirely.

2 · The predicate never looks at the model

nanobot/providers/openai_compat_provider.py:896-912

@staticmethod
def _supports_temperature(model_name, reasoning_effort=None) -> bool:
    """Return True when the model accepts a temperature parameter.

    Kimi K3 uses a fixed temperature that should be omitted. GPT-5 family
    and reasoning models (o1/o3/o4) reject temperature when
    reasoning_effort is set to anything other than ``"none"``.        # ← scope declared
    """
    if _model_slug(model_name) == _KIMI_K3_MODEL:
        return False                                                   # rule 1: by model
    if reasoning_effort and reasoning_effort.lower() != "none":
        return False                                                   # rule 2: no model check
    name = model_name.lower()
    return not any(token in name for token in ("gpt-5", "o1", "o3", "o4"))   # rule 3: by model

The docstring scopes the "effort ⇒ no temperature" rule to GPT-5 family and reasoning
models (o1/o3/o4)
. But the model-name check is rule 3, the last line of the function — it
is only reachable when reasoning_effort is None/"none", i.e. exactly when rule 2 did
not already return. For any model once reasoning_effort is set, rule 2 fires first and the
model name is never consulted:

model effort unset effort set
gpt-5 / o1 / o3 / o4 omitted ✅ intended omitted ✅ intended
gpt-4o sent ✅ omitted ❌
deepseek-v4-flash, glm-*, qwen-*, llama-* … sent ✅ omitted ❌

Both call paths are affected — the same predicate gates them:

  • openai_compat_provider.py:959 (Chat Completions) — if self._supports_temperature(...): kwargs["temperature"] = temperature
  • openai_compat_provider.py:1319 (Responses) — if self._supports_temperature(...): body["temperature"] = temperature

3 · The API being guarded against accepts both (measured)

The predicate presupposes that a target rejects temperature together with
reasoning.effort. Tested directly against api.deepseek.com (bypassing nanobot entirely,
one parameter pair per request):

model=deepseek-v4-flash  base_url=https://api.deepseek.com

[OK] responses · neither
[OK] responses · temperature only
[OK] responses · temperature + reasoning.effort=high     ← the pairing in question
[OK] responses · reasoning.effort=high only
[OK] responses · temperature + reasoning.effort=none
[OK] chat      · temperature only
[OK] chat      · temperature + reasoning_effort=high     ← the other call path

accepted 7/7

No 400, no error body, on either the Responses or the Chat Completions path. So for at least
this provider the suppression has no basis.

(Scope of this measurement: that the API accepts the pair — which is what the predicate
claims otherwise. Whether DeepSeek then actually honours reasoning.effort was not
measured, and is not what this issue is about.)

4 · Where the rule came from

_supports_temperature was added to the compat provider by PR #2788 ("feat(providers):
add GPT-5 model family support", merged 2026-04-04). Its description says:

Add _supports_temperature() helper to conditionally omit temperature for reasoning
models (o1/o3/o4) and when reasoning_effort is active, matching existing Azure provider
behaviour
…… Both changes are backward-compatible — older GPT-4 models work as before.

So the rule was inherited from the Azure provider rather than re-derived, the stated scope is
o1/o3/o4, and the author expected older models to be unaffected — which they are, until
reasoning_effort is set
, which is exactly the case the change was about.

The same intent is recorded a third time in a test name — see §5.

5 · Two test copies pin the current behaviour

(Possibly the part that needs a maintainer decision rather than a code change.)

AzureOpenAIProvider is a separate class (azure_openai_provider.py:94, does not inherit
from the compat provider) and carries its own copy of _supports_temperature
(azure_openai_provider.py:153). Both copies are pinned by tests that assert gpt-4o
is suppressed when an effort is set:

copy test assertion
azure tests/providers/test_azure_openai_provider.py:211 _supports_temperature("gpt-4o", reasoning_effort="medium") is False
compat tests/providers/test_litellm_kwargs.py:1012 _supports_temperature("gpt-4o", reasoning_effort="medium") is False

Note the compat test's own name —
test_openai_compat_supports_temperature_matches_reasoning_model_rules — describes
reasoning model rules while asserting a non-reasoning model (gpt-4o). Together with
the docstring (§2) and the PR description (§4), that is three independent statements of
intent ("o1/o3/o4 only") that the code does not implement.

For reference, the maintainers' direction on this area so far has consistently been to exempt
temperature per model: #2788 (GPT-5), #4685 (sonnet 5), #4967 (Moonshot K2.5/K2.6),
#4334 (opus-4-8 / fable) — all merged.

6 · Open questions

  1. What is the intended scope of the "effort ⇒ omit temperature" rule? The docstring, the
    PR that introduced it, and the test name all say o1/o3/o4 (plus the Azure deployment
    cases), while the implementation covers every model on every openai_compat provider.
    Which is correct?

  2. Should the check be by model name or by spec? A ProviderSpec flag (the predicate
    currently consults neither spec nor model_name for this branch) would let providers
    that accept both keep their configured temperature without relying on name matching.

  3. Should the include coupling be split? In the Responses path the same predicate
    result also decides an unrelated request field
    (openai_compat_provider.py:1322):

    if not self._supports_temperature(model_name, reasoning_effort) and not preserve_reasoning:
        body["include"] = ["reasoning.encrypted_content"]

    Setting an effort therefore also adds include: ["reasoning.encrypted_content"] for every
    non-OpenAI compat provider that doesn't already preserve reasoning (DeepSeek is exempt via
    preserve_reasoning = is_deepseek, :1287-1288). One condition controlling two unrelated
    request fields looks accidental.

I did not test the gpt-4o / Azure branch against the OpenAI or Azure API (no
credentials), so question 1 is the one I can't answer from here.

Steps to reproduce

  1. Configure any backend="openai_compat" provider (e.g. deepseek) with
    agents.defaults.temperature = 0.1 and agents.defaults.reasoningEffort = null.
  2. Run one turn and capture the outbound request: temperature is present.
  3. Set agents.defaults.reasoningEffort = "high", run the same turn: temperature is gone,
    no warning is logged, and config.json still shows 0.1.

Environment

  • nanobot 0.3.5 (local checkout), Python 3.12.0, Windows 11
  • Provider for the outbound capture and the direct API test: DeepSeek (deepseek-v4-flash)
  • Local test suite run in an isolated venv; working tree otherwise unmodified

Additional context

Reproduction scripts (self-contained, ~40 lines each, no nanobot internals needed for the
API probe): outbound kwargs capture, the direct DeepSeek pairing probe, and a listing of every
backend="openai_compat" spec. I can attach them if useful.

Happy to prepare a PR, but the scope question in §6.1 comes first — the answer determines
whether the compat predicate needs the model/spec check, and whether
test_litellm_kwargs.py:1012 is an assertion to keep or to update.

Activity

  1. GZY-SUPER-HACKER commented on Oct 2, 2026

    @GZY-SUPER-HACKER
    ContributorAuthor

    Updating the file:line anchors. Upstream moved ahead while this was open, so the
    line numbers in the report drifted by a few lines — the code itself is unchanged.
    Re-verified against d0d0a44e:

    in the report current
    openai_compat_provider.py:896-912 (_supports_temperature) :897-913
    :959 (Chat Completions call site) :960
    :1319 (Responses call site) :1323
    :1322 (include coupling) :1326

    The predicate is still verbatim what the report quotes, bare condition included,
    now at :910:

    if reasoning_effort and reasoning_effort.lower() != "none":
        return False

    Nothing else in the analysis changes.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions