Problem
Everything the platform does with a knowledge corpus today is conversational: ask a question, get the top chunks across everything the agent can see, get a prose answer. There is one retrieval primitive and one output shape.
There is no way to ask the same set of questions of every document in a defined set and get back a structured, inspectable result. A user who needs "for each of these 200 documents, extract these 8 fields" has to drive it one conversation at a time and transcribe the answers by hand. The work is inherently tabular — documents down one axis, questions across the other — and we have no mode for it.
Example
The shape recurs across domains, which is the point — the engine is domain-neutral:
- A security team has 300 completed vendor questionnaires and wants, for each: which data residency region, whether SSO is supported, the stated breach-notification window, and whether a subprocessor list was provided
- A support team has a quarter of escalation write-ups and wants, for each: root-cause category, time to mitigation, whether a customer commitment was made, and the exact sentence that made it
- A finance team has several hundred supplier records and wants the renewal date, notice period, and price-escalation clause pulled out of each
Same operation every time: a fixed set of documents, a fixed set of questions, one row per document. The user then wants to sort, filter to the rows where something is missing, spot-check a handful of cells against the source, and hand the result to someone else.
Today that is hundreds of separate chats, no record of what was asked, no way to resume after an interruption, and no artifact at the end.
Where
The batch runner this needs already exists — it is a durable Postgres queue, not a new piece of infrastructure.
platform/backend/src/task-queue/task-queue.ts — TaskQueueService: per-lane concurrency caps, FOR UPDATE SKIP LOCKED dequeue, heartbeat renewal, stuck sweep, graceful-shutdown drain
platform/backend/src/types/task.ts — TaskTypeSchema is a closed enum (:19-32) and TASK_LANES maps types to lanes (:50-76); a review task type and its own lane go here. The lane doc comment is explicit that a saturated lane cannot head-of-line-block another
platform/backend/src/task-queue/handlers/batch-embedding-handler.ts — the fan-out template: children complete independently and the last one to finish finalizes the parent run
platform/backend/src/knowledge-base/connector-sync.ts:353-361 — the enqueue shape to copy ({ documentIds, connectorRunId })
platform/backend/src/database/schemas/connector-run.ts — parent-run progress counters (totalBatches, completedBatches, itemErrors) plus lease/epoch fencing, to model a review run on
platform/backend/src/database/schemas/task.ts — note the partial expression index on (payload ->> 'connectorRunId'), precedent for indexing a fan-out parent id out of JSONB
platform/backend/src/config.ts — taskWorkerMaxConcurrent defaults to 2, and that is the shared content-lane cap
Approach
Data model. A review owns a column config and a set of rows. A row is a document. A cell is (row, column) holding content, citations, and a status of pending | generating | done | error.
That cell status is the whole trick: it makes a run resumable (skip cells already done) and a single cell retryable, without a separate job-tracking structure. It is a job table wearing a spreadsheet.
A column is a prompt plus an output format — text, yes/no, date, number, list. Nothing domain-specific belongs in the engine. Domain-specific column presets should ship as a skill pack, not as product code.
Execution is whole-document, not retrieval. Stuff the document text and make one model call per document covering all of its columns. This needs no vector, chunk, or embedding work at all, which means the feature is independent of the RAG roadmap and does not queue behind it.
Runner. A new task type and a dedicated queue lane, structurally modelled on batch_embedding, with a parent run row modelled on connector_runs for progress counters. Give the lane its own concurrency budget — the content lane's default of 2 is sized for embedding, not for a user waiting on a review.
ACL. Rows resolve through the same document ACL filter retrieval already uses. A review narrows what is analysed; it must never widen what a user may read. Drop rows whose source documents are not all accessible.
The differentiator: let a column target a tool, not only a prompt. A column could extract a value from the document, look it up in another system over MCP, and flag the row — per document, across the whole set, under the proxy's cost limits and the audit log. A review engine that can only read the document is table stakes; one that can call tools per row is not.
Known unblockers
- No KB-scoped document enumeration endpoint.
GET /api/knowledge-bases/:id/documents does not exist — enumeration is connector-scoped, so callers must fan out over a KB's connectors themselves. Pagination is offset-only and capped at 100 (shared/pagination.ts:9)
- No
documentId filter on any retrieval path — queryService.query() takes connectorIds and nothing narrower, though kb_chunks.document_id is indexed and the SQL would accept it
- Content-lane worker concurrency of 2
Done when
- A user can define a set of documents and a set of columns, run it, and watch cells populate.
- An interrupted run resumes without recomputing completed cells, and a single cell can be retried on its own.
- Cells carry citations back to the source document.
- A review never surfaces content from a document the user could not otherwise read.
Notes
Sizing is genuinely unknown. Documents per review and columns per document are the two numbers that size the runner, and we have no evidence for either. Get real numbers before picking a fan-out width — the difference between 20×5 and 500×15 is the difference between a background job and a capacity problem.
There is no output surface today. No grid renderer and no CSV/XLSX writer exist anywhere in the repo; the only table-shaped output is a model-authored Markdown table rendered by Streamdown. Deciding where the grid lives is a real part of this work, not a detail to leave until the end.
This is the kind of capability that likely belongs in Enterprise-marked territory — see the license router in LICENSE.md before starting.
Problem
Everything the platform does with a knowledge corpus today is conversational: ask a question, get the top chunks across everything the agent can see, get a prose answer. There is one retrieval primitive and one output shape.
There is no way to ask the same set of questions of every document in a defined set and get back a structured, inspectable result. A user who needs "for each of these 200 documents, extract these 8 fields" has to drive it one conversation at a time and transcribe the answers by hand. The work is inherently tabular — documents down one axis, questions across the other — and we have no mode for it.
Example
The shape recurs across domains, which is the point — the engine is domain-neutral:
Same operation every time: a fixed set of documents, a fixed set of questions, one row per document. The user then wants to sort, filter to the rows where something is missing, spot-check a handful of cells against the source, and hand the result to someone else.
Today that is hundreds of separate chats, no record of what was asked, no way to resume after an interruption, and no artifact at the end.
Where
The batch runner this needs already exists — it is a durable Postgres queue, not a new piece of infrastructure.
platform/backend/src/task-queue/task-queue.ts—TaskQueueService: per-lane concurrency caps,FOR UPDATE SKIP LOCKEDdequeue, heartbeat renewal, stuck sweep, graceful-shutdown drainplatform/backend/src/types/task.ts—TaskTypeSchemais a closed enum (:19-32) andTASK_LANESmaps types to lanes (:50-76); a review task type and its own lane go here. The lane doc comment is explicit that a saturated lane cannot head-of-line-block anotherplatform/backend/src/task-queue/handlers/batch-embedding-handler.ts— the fan-out template: children complete independently and the last one to finish finalizes the parent runplatform/backend/src/knowledge-base/connector-sync.ts:353-361— the enqueue shape to copy ({ documentIds, connectorRunId })platform/backend/src/database/schemas/connector-run.ts— parent-run progress counters (totalBatches,completedBatches,itemErrors) plus lease/epoch fencing, to model a review run onplatform/backend/src/database/schemas/task.ts— note the partial expression index on(payload ->> 'connectorRunId'), precedent for indexing a fan-out parent id out of JSONBplatform/backend/src/config.ts—taskWorkerMaxConcurrentdefaults to 2, and that is the shared content-lane capApproach
Data model. A review owns a column config and a set of rows. A row is a document. A cell is
(row, column)holding content, citations, and a status ofpending | generating | done | error.That cell status is the whole trick: it makes a run resumable (skip cells already
done) and a single cell retryable, without a separate job-tracking structure. It is a job table wearing a spreadsheet.A column is a prompt plus an output format — text, yes/no, date, number, list. Nothing domain-specific belongs in the engine. Domain-specific column presets should ship as a skill pack, not as product code.
Execution is whole-document, not retrieval. Stuff the document text and make one model call per document covering all of its columns. This needs no vector, chunk, or embedding work at all, which means the feature is independent of the RAG roadmap and does not queue behind it.
Runner. A new task type and a dedicated queue lane, structurally modelled on
batch_embedding, with a parent run row modelled onconnector_runsfor progress counters. Give the lane its own concurrency budget — the content lane's default of 2 is sized for embedding, not for a user waiting on a review.ACL. Rows resolve through the same document ACL filter retrieval already uses. A review narrows what is analysed; it must never widen what a user may read. Drop rows whose source documents are not all accessible.
The differentiator: let a column target a tool, not only a prompt. A column could extract a value from the document, look it up in another system over MCP, and flag the row — per document, across the whole set, under the proxy's cost limits and the audit log. A review engine that can only read the document is table stakes; one that can call tools per row is not.
Known unblockers
GET /api/knowledge-bases/:id/documentsdoes not exist — enumeration is connector-scoped, so callers must fan out over a KB's connectors themselves. Pagination is offset-only and capped at 100 (shared/pagination.ts:9)documentIdfilter on any retrieval path —queryService.query()takesconnectorIdsand nothing narrower, thoughkb_chunks.document_idis indexed and the SQL would accept itDone when
Notes
Sizing is genuinely unknown. Documents per review and columns per document are the two numbers that size the runner, and we have no evidence for either. Get real numbers before picking a fan-out width — the difference between 20×5 and 500×15 is the difference between a background job and a capacity problem.
There is no output surface today. No grid renderer and no CSV/XLSX writer exist anywhere in the repo; the only table-shaped output is a model-authored Markdown table rendered by Streamdown. Deciding where the grid lives is a real part of this work, not a detail to leave until the end.
This is the kind of capability that likely belongs in Enterprise-marked territory — see the license router in
LICENSE.mdbefore starting.