Skip to content

Authenticated denial of service through unbounded XLSX embedded-image decompression

Moderate
liuruibin published GHSA-657x-9gfv-m4p8 Sep 3, 2026

Package

1Panel-dev/MaxKB

Affected versions

>= 2.7.0, <= 2.10.5-lts

Patched versions

v2.10.6-lts

Description

Summary

An ordinary authenticated MaxKB user can upload a small XLSX file containing an embedded
image with very large pixel dimensions.

MaxKB globally disables Pillow's decompression-bomb pixel limit. It then decodes the image,
converts it to RGB, and encodes it as JPEG inside a shared HTTP worker. The upload check uses
only the small outer XLSX file size.

In the official v2.10.5-lts all-in-one image, a 7,307-byte XLSX caused an OOM kill. Three
concurrent requests caused more OOM kills and timeouts for a separate ordinary user. The
service recovered by replacing workers, so this report claims transient availability loss,
not permanent container termination.

Tested version and deployment

  • MaxKB: v2.10.5-lts
  • Commit: 01b21db88145278d98bf5e9bd55e6abd6b3aad43
  • Official image: 1panel/maxkb:v2.10.5-lts
  • Image digest:
    sha256:23920f160a799844ecf78947a31efcc20f9e796a544ea16d0efb3ee684cb6b76
  • Architecture: ARM64
  • Resources: 4 GiB memory, 1 CPU, 512 PIDs
  • Server: the image's default all-in-one PostgreSQL, Redis, Gunicorn, local-model, and Celery
    processes
  • Gunicorn: three workers with the image's default configuration

The relevant source files in the image have the same SHA-256 hashes as the tested release.
The same vulnerable common_handle.py hash is present in every checked release from
v2.7.0 through v2.10.5-lts.

Why this is a security issue

  • The attacker is a normal workspace USER, not an administrator.
  • The vulnerable /document/split route explicitly permits the ordinary-user role.
  • The attack uses a normal XLSX upload to a knowledge base the user can access.
  • The request is only 7,307 bytes and does not need high network traffic.
  • The embedded image has 900 million pixels. RGB conversion needs at least 2.7 GB before
    JPEG-encoding and other parser memory are counted.
  • MaxKB explicitly sets Pillow's existing safety limit to None.
  • A separate authenticated user experienced request timeouts during the attack.
  • No LLM request, special Agent feature, API key, or victim action is required.

Root cause

The document-split route accepts an ordinary workspace user:

class Split(APIView):
    authentication_classes = [TokenAuth]
    parser_classes = [MultiPartParser]

    # @extend_schema(...) omitted
    @has_permissions(
        PermissionConstants.KNOWLEDGE_DOCUMENT_READ.get_workspace_knowledge_permission(),
        PermissionConstants.KNOWLEDGE_DOCUMENT_READ.get_workspace_permission_workspace_manage_role(),
        RoleConstants.WORKSPACE_MANAGE.get_workspace_role(),
        ViewPermission(
            [RoleConstants.USER.get_workspace_role()],
            [PermissionConstants.KNOWLEDGE.get_workspace_knowledge_permission()],
            CompareConstants.AND,
        ),
    )

It validates only the uploaded XLSX byte count. The default knowledge-base limit is 100 MiB:

for f in files:
    if f.size > 1024 * 1024 * knowledge.file_size_limit:
        raise AppApiException(...)

The XLSX handler loads the workbook and processes embedded cell images synchronously:

workbook = openpyxl.load_workbook(io.BytesIO(buffer))
image_dict: dict = xlsx_embed_cells_images(io.BytesIO(buffer))

At module import, MaxKB disables Pillow's decompression-bomb protection:

ImageFile.LOAD_TRUNCATED_IMAGES = True
PILImage.MAX_IMAGE_PIXELS = None

The embedded image is then fully decoded, RGB-converted, and JPEG-encoded without a pixel,
decoded-byte, expansion-ratio, time, or memory limit:

f = archive.open(img.target)
img_byte = io.BytesIO()
im = PILImage.open(f).convert('RGB')
im.save(img_byte, format='JPEG')

Source links:

  • class Split(APIView):
    authentication_classes = [TokenAuth]
    parser_classes = [MultiPartParser]
    @extend_schema(
    methods=["POST"],
    description=_("Segmented document"),
    summary=_("Segmented document"),
    operation_id=_("Segmented document"), # type: ignore
    parameters=DocumentSplitAPI.get_parameters(),
    request=DocumentSplitAPI.get_request(),
    responses=DocumentSplitAPI.get_response(),
    tags=[_("Knowledge Base/Documentation")], # type: ignore
    )
    @has_permissions(
    PermissionConstants.KNOWLEDGE_DOCUMENT_READ.get_workspace_knowledge_permission(),
    PermissionConstants.KNOWLEDGE_DOCUMENT_READ.get_workspace_permission_workspace_manage_role(),
    RoleConstants.WORKSPACE_MANAGE.get_workspace_role(),
    ViewPermission(
    [RoleConstants.USER.get_workspace_role()],
    [PermissionConstants.KNOWLEDGE.get_workspace_knowledge_permission()],
    CompareConstants.AND,
    ),
    )
    def post(self, request: Request, workspace_id: str, knowledge_id: str):
    split_data = {"file": request.FILES.getlist("file")}
    request_data = request.data
    if (
    "patterns" in request.data
    and request.data.get("patterns") is not None
    and len(request.data.get("patterns")) > 0
    ):
    split_data.__setitem__("patterns", request_data.getlist("patterns"))
    if "limit" in request.data:
    split_data.__setitem__("limit", request_data.get("limit"))
    if "with_filter" in request.data:
    split_data.__setitem__("with_filter", request_data.get("with_filter"))
    return result.success(
    DocumentSerializers.Split(
    data={
    "workspace_id": workspace_id,
    "knowledge_id": knowledge_id,
    }
  • class Split(serializers.Serializer):
    workspace_id = serializers.CharField(required=False, label=_("workspace id"), allow_null=True)
    knowledge_id = serializers.UUIDField(required=True, label=_("knowledge id"))
    def is_valid(self, *, instance=None, raise_exception=True):
    super().is_valid(raise_exception=True)
    workspace_id = self.data.get("workspace_id")
    query_set = QuerySet(Knowledge).filter(id=self.data.get("knowledge_id"))
    if workspace_id:
    query_set = query_set.filter(workspace_id=workspace_id)
    if not query_set.exists():
    raise AppApiException(500, _("Knowledge id does not exist"))
    files = instance.get("file")
    knowledge = Knowledge.objects.filter(id=self.data.get("knowledge_id")).first()
    for f in files:
    if f.size > 1024 * 1024 * knowledge.file_size_limit:
    raise AppApiException(
    500,
    _("The maximum size of the uploaded file cannot exceed {}MB").format(knowledge.file_size_limit),
    )
    def parse(self, instance):
    self.is_valid(instance=instance, raise_exception=True)
    DocumentSplitRequest(data=instance).is_valid(raise_exception=True)
    file_list = instance.get("file")
    return reduce(
    lambda x, y: [*x, *y],
    [
    self.file_to_paragraph(
    f,
    instance.get("patterns", None),
    instance.get("with_filter", None),
    instance.get("limit", 4096),
    )
    for f in file_list
    ],
    [],
    )
    def save_image(self, image_list):
    if image_list is not None and len(image_list) > 0:
    exist_image_list = [
    str(i.get("id")) for i in QuerySet(File).filter(id__in=[i.id for i in image_list]).values("id")
    ]
    save_image_list = [image for image in image_list if not exist_image_list.__contains__(str(image.id))]
    save_image_list = list({img.id: img for img in save_image_list}.values())
    # save image
    for file in save_image_list:
    file_bytes = file.meta.pop("content")
    file.meta["knowledge_id"] = self.data.get("knowledge_id")
    file.source_type = FileSourceType.KNOWLEDGE
    file.source_id = self.data.get("knowledge_id")
    file.save(file_bytes)
    def file_to_paragraph(self, file, pattern_list: List, with_filter: bool, limit: int):
    # 保存源文件
    file_id = uuid.uuid7()
    raw_file = File(
    id=file_id,
    file_name=file.name,
    file_size=file.size,
    source_type=FileSourceType.KNOWLEDGE,
    source_id=self.data.get("knowledge_id"),
    )
    raw_file.save(file.read())
    file.seek(0)
    get_buffer = FileBufferHandle().get_buffer
    for split_handle in split_handles:
    if split_handle.support(file, get_buffer):
    result = split_handle.handle(file, pattern_list, with_filter, limit, get_buffer, self.save_image)
    if isinstance(result, list):
    for item in result:
    item["source_file_id"] = file_id
    return result
    result["source_file_id"] = file_id
    return [result]
    result = default_split_handle.handle(file, pattern_list, with_filter, limit, get_buffer, self.save_image)
    if isinstance(result, list):
    for item in result:
    item["source_file_id"] = file_id
    return result
    result["source_file_id"] = file_id
    return [result]
  • def handle(self, file, pattern_list: List, with_filter: bool, limit: int, get_buffer, save_image):
    buffer = get_buffer(file)
    try:
    if type(limit) is str:
    limit = int(limit)
    workbook = openpyxl.load_workbook(io.BytesIO(buffer))
    try:
    image_dict: dict = xlsx_embed_cells_images(io.BytesIO(buffer))
    save_image([item for item in image_dict.values()])
    except Exception as e:
    image_dict = {}
    worksheets = workbook.worksheets
    worksheets_size = len(worksheets)
    return [row for row in
    [handle_sheet(file.name,
    sheet,
    image_dict,
    limit) if worksheets_size == 1 and sheet.title == 'Sheet1' else handle_sheet(
    sheet.title, sheet, image_dict, limit) for sheet
    in worksheets] if row is not None]
    except Exception as e:
    maxkb_logger.error(f"Error processing XLSX file {file.name}: {e}, {traceback.format_exc()}")
    return [{'name': file.name, 'content': []}]
  • from PIL import ImageFile
    ImageFile.LOAD_TRUNCATED_IMAGES = True
    PILImage.MAX_IMAGE_PIXELS = None
  • def handle_images(deps, archive: ZipFile) -> []:
    images = []
    if not PILImage: # Pillow not installed, drop images
    return images
    for dep in deps:
    try:
    image_io = archive.read(dep.target)
    image = openpyxl_Image(BytesIO(image_io))
    except Exception as e:
    maxkb_logger.error(f"Error reading image {dep.target}: {e}, {traceback.format_exc()}")
    continue
    image.embed = dep.id # 文件rId
    image.target = dep.target # 文件地址
    images.append(image)
    return images
    def xlsx_embed_cells_images(buffer) -> {}:
    archive = ZipFile(buffer)
    # 解析cellImage.xml文件
    deps = get_dependents(archive, get_rels_path("xl/cellimages.xml"))
    image_rel = handle_images(deps=deps, archive=archive)
    # 工作表及其中图片ID
    sheet_list = {}
    for item in archive.namelist():
    if not item.startswith('xl/worksheets/sheet'):
    continue
    key = item.split('/')[-1].split('.')[0].split('sheet')[-1]
    sheet_list[key] = parse_element_sheet_xml(fromstring(archive.read(item)))
    cell_images_xml = parse_element(fromstring(archive.read("xl/cellimages.xml")))
    cell_images_rel = {}
    for image in image_rel:
    cell_images_rel[image.embed] = image
    for cnv, embed in cell_images_xml.items():
    cell_images_xml[cnv] = cell_images_rel.get(embed)
    result = {}
    for key, img in cell_images_xml.items():
    all_cells = [
    cell
    for _sheet_id, sheet in sheet_list.items()
    if sheet is not None
    for cell in sheet or []
    ]
    image_excel_id_list = [
    cell for cell in all_cells
    if isinstance(cell, str) and key in cell
    ]
    # print(key, img)
    if img is None:
    continue
    if len(image_excel_id_list) > 0:
    image_excel_id = image_excel_id_list[-1]
    f = archive.open(img.target)
    img_byte = io.BytesIO()
    im = PILImage.open(f).convert('RGB')
    im.save(img_byte, format='JPEG')
    image = File(id=uuid.uuid7(), file_name=img.path, meta={'debug': False, 'content': img_byte.getvalue()})
    result['=' + image_excel_id] = image
    archive.close()

Proof of concept

maxkb-xlsx-embedded-image-dos-poc-v1.zip

Please attach and extract:

maxkb-xlsx-embedded-image-dos-poc-v1.zip

The included README provides a direct official-image test and a full HTTP test. All traffic
uses an isolated Docker network with no published host port.

The full test performs this request as a default ordinary user:

POST /admin/api/workspace/default/knowledge/{knowledge_id}/document/split
Authorization: Bearer <ordinary-user token>
Content-Type: multipart/form-data

file=@probe-30000.xlsx
limit=4096

Fixture details:

  • Outer XLSX size: 7,307 bytes
  • Embedded PNG size: 874,852 bytes
  • Image dimensions: 30,000 x 30,000
  • Pixel count: 900,000,000
  • Minimum RGB data: 2,700,000,000 bytes
  • XLSX SHA-256:
    40aef1fd024a0c4299e06b6a9875208087655884037707aadbef5e02705e5923

The lab used a placeholder embedding-model UUID only when creating its isolated knowledge
base because no external model was configured. The API accepted it. The vulnerable split
and image parsing path does not call the embedding model. A normal deployment can use an
existing accessible knowledge base or a configured embedding model.

Results

Full HTTP application

A 100x100 control file returned HTTP 200 in 0.057 seconds. The container remained healthy
and Docker reported no OOM.

One 7,307-byte attack request produced an empty reply after 1.126 seconds and changed
Docker's OOM state to true. The Gunicorn master and container remained running and replaced
the affected worker.

I then sent three attack requests concurrently while a second default USER repeatedly
called GET /admin/api/user/profile:

Result Value
Aggregate attacker input 21,921 bytes
Attacker requests ending with an empty reply 3 / 3
cgroup oom_kill counter 3 to 5
Second-user profile requests 60
Second-user HTTP 200 responses 46
Second-user 500 ms timeouts 14
Post-attack recovery request HTTP 200 in 0.113 s

The oom_kill counter was already 3 from earlier validation. This burst added two OOM
kills; it did not cause five new kills.

This demonstrates transient cross-user availability impact.

Direct official-image test

The exact image source function was called in a networkless 512 MiB container:

Fixture Result
8,000x8,000 control success, about 357 MiB (365,328 KiB) peak RSS
20,000x20,000 attack, Run 1 exit 137, OOMKilled=true
20,000x20,000 attack, Run 2 exit 137, OOMKilled=true
20,000x20,000 attack, Run 3 exit 137, OOMKilled=true

Impact

Confirmed: a normal authenticated user can cause repeatable OOM kills in shared MaxKB HTTP
workers with very small requests. A short three-request burst caused timeouts for another
authenticated user.

Inferred: a sustained request stream can keep recycling workers and degrade the service.
More memory or more replicas can reduce the effect, but each concurrent decode still has an
unbounded multi-gigabyte memory cost.

A single request kills one worker and ends that request. Sustained cross-user degradation
requires continuing or concurrent requests. I did not measure a long-duration flood or a
deployment other than the 4 GiB all-in-one image.

Suggested fix

  • Remove the global PILImage.MAX_IMAGE_PIXELS = None assignment.
  • Check image dimensions before decoding and enforce per-image and aggregate pixel limits.
  • Limit XLSX member count, decoded bytes, and expansion ratio.
  • Parse untrusted documents in an isolated worker with a hard memory limit.
  • Limit concurrent document parsing per user.
  • Return a bounded 4xx error and delete partially saved files when a parser limit is reached.

Severity

Moderate

CVSS overall score

This score calculates overall vulnerability severity from 0 to 10 and is based on the Common Vulnerability Scoring System (CVSS).
/ 10

CVSS v3 base metrics

Attack vector
Network
Attack complexity
Low
Privileges required
Low
User interaction
None
Scope
Unchanged
Confidentiality
None
Integrity
None
Availability
Low

CVSS v3 base metrics

Attack vector: More severe the more the remote (logically and physically) an attacker can be in order to exploit the vulnerability.
Attack complexity: More severe for the least complex attacks.
Privileges required: More severe if no privileges are required.
User interaction: More severe when no user interaction is required.
Scope: More severe when a scope change occurs, e.g. one vulnerable component impacts resources in components beyond its security scope.
Confidentiality: More severe when loss of data confidentiality is highest, measuring the level of data access available to an unauthorized user.
Integrity: More severe when loss of data integrity is the highest, measuring the consequence of data modification possible by an unauthorized user.
Availability: More severe when the loss of impacted component availability is highest.
CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:L

CVE ID

No known CVE

Weaknesses

Uncontrolled Resource Consumption

The product does not properly control the allocation and maintenance of a limited resource. Learn more on MITRE.

Improper Handling of Highly Compressed Data (Data Amplification)

The product does not handle or incorrectly handles a compressed input with a very high compression ratio that produces a large output. Learn more on MITRE.

Credits