Skip to content

Knowledge Base: add a standalone recall test for retrieval tuning #2762

Description

@Bruce-Yii

Is your feature request related to a problem? Please describe.

When tuning a knowledge base, users currently need to go through an Agent or Workflow debugging flow to inspect what was actually retrieved. That works for debugging a complete downstream request, but it is relatively heavy when the goal is only to answer a narrower question:

For this knowledge base and this query, which chunks are retrieved under a given retrieval configuration, with what scores and sources?

There are several related community signals:

The problem is not that Coze has no retrieval observability at all. Agent/Workflow debugging can already show recalled chunks. The missing piece is a lightweight knowledge-base-level retrieval tuning loop that does not require first constructing or entering a downstream orchestration flow.

Describe the solution you'd like

Add a standalone Recall Test / Retrieval Playground on the Knowledge Base detail page.

A minimal version could:

  1. Let the user enter a single query.
  2. Let the user provide temporary, test-scoped retrieval settings such as search strategy, TopK, and MinScore.
  3. Reuse the same production knowledge retrieval path rather than introducing a separate retrieval implementation.
  4. Show ranked retrieved chunks with at least final score and source document information.
  5. Make zero-result cases explicit together with the active test settings.
  6. Keep all test settings ephemeral: running a test should not modify the Knowledge Base itself or any Agent/Workflow configuration.

For a first version, I would intentionally keep the scope small: no batch dataset evaluation, no Recall@K / MRR / NDCG metrics, no experiment history, and no retrieval algorithm changes.

Describe alternatives you've considered

Additional context

The current codebase already appears to have much of the underlying foundation:

  • the knowledge retrieval domain path returns retrieved slices with scores;
  • the Knowledge Base has its own IDE/detail surface where a test action could naturally live;
  • the public frontend still contains an autogenerated RetrieveTestReq/Resp contract and a generated client path for /api/devops/knowledge_platform/v1/retrieve_test originating from the earlier Bytedance code mirror.

I am not assuming that the old generated RetrieveTest contract is intended to be restored in the open-source product; it may be an internal-only historical artifact. I am mentioning it only because it suggests that this workflow has been modeled before.

Before implementing, I would like to align with the maintainers on three points:

  1. Is a standalone Knowledge-level retrieval test something you want in the open-source product, or should retrieval inspection intentionally remain inside Agent/Workflow debugging?
  2. Is the existing generated RetrieveTest contract still conceptually relevant, or should an OSS-native API be designed around the current knowledge retrieval domain path?
  3. If the direction is welcome, which Knowledge UI surface and API boundary would you prefer for a minimal first version?

If this direction fits the project, I would be happy to take ownership of the scoped implementation after we align on the boundary.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions