What happens
When ranking fails or is skipped, LLMRanker returns every deduplicated input document and ignores top_k. The docstring says top_k is "The maximum number of documents to return", and the successful path applies it (ranked_documents[:top_k]), but the fallback paths do not.
Three paths are affected, in both run() and run_async():
- the chat generator raises and
raise_on_failure=False
- the reply cannot be parsed and
raise_on_failure=False
- the query is empty or not a string
Reproducer
from haystack import Document
from haystack.components.rankers.llm_ranker import LLMRanker
class FailingGenerator:
def run(self, messages):
raise RuntimeError("simulated API failure")
docs = [Document(content=f"doc {i}") for i in range(8)]
ranker = LLMRanker(chat_generator=FailingGenerator(), top_k=2)
result = ranker.run(query="q", documents=docs)
print(len(result["documents"])) # 8, expected at most 2
The same thing happens with a whitespace-only query and with a reply that is not valid JSON: all 8 documents come back with top_k=2. Tested on current main, twice with identical results.
Why it matters
top_k is usually set to keep the reranker output small enough for a downstream prompt. A transient API error or rate limit then silently passes the full candidate set, which can be hundreds of documents, to the next component. The pipeline produces an oversized prompt exactly when the system is already degraded.
Note that test_run_whitespace_query_returns_fallback currently pins this behavior (top_k=1, two documents returned), so a fix would need that test updated as well.
Happy to open a draft PR that slices the fallback results with top_k, keeping the original document order, if you want this changed.
What happens
When ranking fails or is skipped,
LLMRankerreturns every deduplicated input document and ignorestop_k. The docstring saystop_kis "The maximum number of documents to return", and the successful path applies it (ranked_documents[:top_k]), but the fallback paths do not.Three paths are affected, in both
run()andrun_async():raise_on_failure=Falseraise_on_failure=FalseReproducer
The same thing happens with a whitespace-only query and with a reply that is not valid JSON: all 8 documents come back with
top_k=2. Tested on currentmain, twice with identical results.Why it matters
top_kis usually set to keep the reranker output small enough for a downstream prompt. A transient API error or rate limit then silently passes the full candidate set, which can be hundreds of documents, to the next component. The pipeline produces an oversized prompt exactly when the system is already degraded.Note that
test_run_whitespace_query_returns_fallbackcurrently pins this behavior (top_k=1, two documents returned), so a fix would need that test updated as well.Happy to open a draft PR that slices the fallback results with
top_k, keeping the original document order, if you want this changed.