Model config declares a context size the provider often does not actually serve, and in workflows that is per-node #42978
behrnt-slatgng
started this conversation in
Suggestion
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Disclosure: I build Grunz, a hosted chat and coding agent on open weights. Different product, but Dify's model-configuration layer is the natural place to fix something that currently costs users a lot of confused debugging.
The declared context size and the served context size are different numbers
When you add a model in Dify you tell it the context size, or it comes from the provider template. That number is a declaration. The window the endpoint actually serves is whatever the provider configured, and the two disagree far more often than people expect.
In practice most hosted endpoints serve around 32K tokens regardless of what the model card advertises. A few reach 256K. I have not found one genuinely serving the 1M figures that get quoted — including for models whose cards claim it.
It fails silently: no 413, no truncation notice. The front of the context is evicted and the model answers confidently from what remains.
Why Dify is unusually exposed
It is per-node, not per-app. Because each LLM node in a workflow can point at a different model and provider, the effective ceiling varies within a single run. A classifier node on a genuinely long endpoint feeding a generator node on a silently-32K one produces exactly the bug that is miserable to diagnose: each node looks correct in isolation, the workflow output drops context, and nothing in the run log says why.
Memory and conversation variables are sized against the declared number. If the declared window is 200K and the real one is 32K, Dify's own memory truncation is doing its arithmetic against a ceiling that does not exist, so it will hand the provider more than it can hold and the provider will quietly discard the oldest part — which is usually the system prompt and the instructions, not the recent chatter.
Knowledge retrieval makes it worse. Top-K chunks plus a long system prompt plus memory is easy to push past 32K without anyone noticing, and the symptom is "the model ignored my instructions" rather than anything that looks like a context problem.
Suggested
One more, for anyone here running open weights
Refusal-ablated variants lose instruction-following and output-format adherence before they lose knowledge. Prose stays good while structured output drifts off-spec — which in a Dify workflow means a classifier or parameter-extractor node starts failing to parse while the model looks perfectly healthy in a chat test. Perplexity will not warn you; test format compliance separately.
Happy to go deeper on any of it.
All reactions