Summary
An ordinary authenticated MaxKB user can upload a small XLSX file containing an embedded
image with very large pixel dimensions.
MaxKB globally disables Pillow's decompression-bomb pixel limit. It then decodes the image,
converts it to RGB, and encodes it as JPEG inside a shared HTTP worker. The upload check uses
only the small outer XLSX file size.
In the official v2.10.5-lts all-in-one image, a 7,307-byte XLSX caused an OOM kill. Three
concurrent requests caused more OOM kills and timeouts for a separate ordinary user. The
service recovered by replacing workers, so this report claims transient availability loss,
not permanent container termination.
Tested version and deployment
- MaxKB:
v2.10.5-lts
- Commit:
01b21db88145278d98bf5e9bd55e6abd6b3aad43
- Official image:
1panel/maxkb:v2.10.5-lts
- Image digest:
sha256:23920f160a799844ecf78947a31efcc20f9e796a544ea16d0efb3ee684cb6b76
- Architecture: ARM64
- Resources: 4 GiB memory, 1 CPU, 512 PIDs
- Server: the image's default all-in-one PostgreSQL, Redis, Gunicorn, local-model, and Celery
processes
- Gunicorn: three workers with the image's default configuration
The relevant source files in the image have the same SHA-256 hashes as the tested release.
The same vulnerable common_handle.py hash is present in every checked release from
v2.7.0 through v2.10.5-lts.
Why this is a security issue
- The attacker is a normal workspace
USER, not an administrator.
- The vulnerable
/document/split route explicitly permits the ordinary-user role.
- The attack uses a normal XLSX upload to a knowledge base the user can access.
- The request is only 7,307 bytes and does not need high network traffic.
- The embedded image has 900 million pixels. RGB conversion needs at least 2.7 GB before
JPEG-encoding and other parser memory are counted.
- MaxKB explicitly sets Pillow's existing safety limit to
None.
- A separate authenticated user experienced request timeouts during the attack.
- No LLM request, special Agent feature, API key, or victim action is required.
Root cause
The document-split route accepts an ordinary workspace user:
class Split(APIView):
authentication_classes = [TokenAuth]
parser_classes = [MultiPartParser]
# @extend_schema(...) omitted
@has_permissions(
PermissionConstants.KNOWLEDGE_DOCUMENT_READ.get_workspace_knowledge_permission(),
PermissionConstants.KNOWLEDGE_DOCUMENT_READ.get_workspace_permission_workspace_manage_role(),
RoleConstants.WORKSPACE_MANAGE.get_workspace_role(),
ViewPermission(
[RoleConstants.USER.get_workspace_role()],
[PermissionConstants.KNOWLEDGE.get_workspace_knowledge_permission()],
CompareConstants.AND,
),
)
It validates only the uploaded XLSX byte count. The default knowledge-base limit is 100 MiB:
for f in files:
if f.size > 1024 * 1024 * knowledge.file_size_limit:
raise AppApiException(...)
The XLSX handler loads the workbook and processes embedded cell images synchronously:
workbook = openpyxl.load_workbook(io.BytesIO(buffer))
image_dict: dict = xlsx_embed_cells_images(io.BytesIO(buffer))
At module import, MaxKB disables Pillow's decompression-bomb protection:
ImageFile.LOAD_TRUNCATED_IMAGES = True
PILImage.MAX_IMAGE_PIXELS = None
The embedded image is then fully decoded, RGB-converted, and JPEG-encoded without a pixel,
decoded-byte, expansion-ratio, time, or memory limit:
f = archive.open(img.target)
img_byte = io.BytesIO()
im = PILImage.open(f).convert('RGB')
im.save(img_byte, format='JPEG')
Source links:
|
class Split(APIView): |
|
authentication_classes = [TokenAuth] |
|
parser_classes = [MultiPartParser] |
|
|
|
@extend_schema( |
|
methods=["POST"], |
|
description=_("Segmented document"), |
|
summary=_("Segmented document"), |
|
operation_id=_("Segmented document"), # type: ignore |
|
parameters=DocumentSplitAPI.get_parameters(), |
|
request=DocumentSplitAPI.get_request(), |
|
responses=DocumentSplitAPI.get_response(), |
|
tags=[_("Knowledge Base/Documentation")], # type: ignore |
|
) |
|
@has_permissions( |
|
PermissionConstants.KNOWLEDGE_DOCUMENT_READ.get_workspace_knowledge_permission(), |
|
PermissionConstants.KNOWLEDGE_DOCUMENT_READ.get_workspace_permission_workspace_manage_role(), |
|
RoleConstants.WORKSPACE_MANAGE.get_workspace_role(), |
|
ViewPermission( |
|
[RoleConstants.USER.get_workspace_role()], |
|
[PermissionConstants.KNOWLEDGE.get_workspace_knowledge_permission()], |
|
CompareConstants.AND, |
|
), |
|
) |
|
def post(self, request: Request, workspace_id: str, knowledge_id: str): |
|
split_data = {"file": request.FILES.getlist("file")} |
|
request_data = request.data |
|
if ( |
|
"patterns" in request.data |
|
and request.data.get("patterns") is not None |
|
and len(request.data.get("patterns")) > 0 |
|
): |
|
split_data.__setitem__("patterns", request_data.getlist("patterns")) |
|
if "limit" in request.data: |
|
split_data.__setitem__("limit", request_data.get("limit")) |
|
if "with_filter" in request.data: |
|
split_data.__setitem__("with_filter", request_data.get("with_filter")) |
|
return result.success( |
|
DocumentSerializers.Split( |
|
data={ |
|
"workspace_id": workspace_id, |
|
"knowledge_id": knowledge_id, |
|
} |
|
class Split(serializers.Serializer): |
|
workspace_id = serializers.CharField(required=False, label=_("workspace id"), allow_null=True) |
|
knowledge_id = serializers.UUIDField(required=True, label=_("knowledge id")) |
|
|
|
def is_valid(self, *, instance=None, raise_exception=True): |
|
super().is_valid(raise_exception=True) |
|
workspace_id = self.data.get("workspace_id") |
|
query_set = QuerySet(Knowledge).filter(id=self.data.get("knowledge_id")) |
|
if workspace_id: |
|
query_set = query_set.filter(workspace_id=workspace_id) |
|
if not query_set.exists(): |
|
raise AppApiException(500, _("Knowledge id does not exist")) |
|
files = instance.get("file") |
|
knowledge = Knowledge.objects.filter(id=self.data.get("knowledge_id")).first() |
|
for f in files: |
|
if f.size > 1024 * 1024 * knowledge.file_size_limit: |
|
raise AppApiException( |
|
500, |
|
_("The maximum size of the uploaded file cannot exceed {}MB").format(knowledge.file_size_limit), |
|
) |
|
|
|
def parse(self, instance): |
|
self.is_valid(instance=instance, raise_exception=True) |
|
DocumentSplitRequest(data=instance).is_valid(raise_exception=True) |
|
|
|
file_list = instance.get("file") |
|
return reduce( |
|
lambda x, y: [*x, *y], |
|
[ |
|
self.file_to_paragraph( |
|
f, |
|
instance.get("patterns", None), |
|
instance.get("with_filter", None), |
|
instance.get("limit", 4096), |
|
) |
|
for f in file_list |
|
], |
|
[], |
|
) |
|
|
|
def save_image(self, image_list): |
|
if image_list is not None and len(image_list) > 0: |
|
exist_image_list = [ |
|
str(i.get("id")) for i in QuerySet(File).filter(id__in=[i.id for i in image_list]).values("id") |
|
] |
|
save_image_list = [image for image in image_list if not exist_image_list.__contains__(str(image.id))] |
|
save_image_list = list({img.id: img for img in save_image_list}.values()) |
|
# save image |
|
for file in save_image_list: |
|
file_bytes = file.meta.pop("content") |
|
file.meta["knowledge_id"] = self.data.get("knowledge_id") |
|
file.source_type = FileSourceType.KNOWLEDGE |
|
file.source_id = self.data.get("knowledge_id") |
|
file.save(file_bytes) |
|
|
|
def file_to_paragraph(self, file, pattern_list: List, with_filter: bool, limit: int): |
|
# 保存源文件 |
|
file_id = uuid.uuid7() |
|
raw_file = File( |
|
id=file_id, |
|
file_name=file.name, |
|
file_size=file.size, |
|
source_type=FileSourceType.KNOWLEDGE, |
|
source_id=self.data.get("knowledge_id"), |
|
) |
|
raw_file.save(file.read()) |
|
file.seek(0) |
|
|
|
get_buffer = FileBufferHandle().get_buffer |
|
for split_handle in split_handles: |
|
if split_handle.support(file, get_buffer): |
|
result = split_handle.handle(file, pattern_list, with_filter, limit, get_buffer, self.save_image) |
|
if isinstance(result, list): |
|
for item in result: |
|
item["source_file_id"] = file_id |
|
return result |
|
result["source_file_id"] = file_id |
|
return [result] |
|
result = default_split_handle.handle(file, pattern_list, with_filter, limit, get_buffer, self.save_image) |
|
if isinstance(result, list): |
|
for item in result: |
|
item["source_file_id"] = file_id |
|
return result |
|
result["source_file_id"] = file_id |
|
return [result] |
|
def handle(self, file, pattern_list: List, with_filter: bool, limit: int, get_buffer, save_image): |
|
buffer = get_buffer(file) |
|
try: |
|
if type(limit) is str: |
|
limit = int(limit) |
|
workbook = openpyxl.load_workbook(io.BytesIO(buffer)) |
|
try: |
|
image_dict: dict = xlsx_embed_cells_images(io.BytesIO(buffer)) |
|
save_image([item for item in image_dict.values()]) |
|
except Exception as e: |
|
image_dict = {} |
|
worksheets = workbook.worksheets |
|
worksheets_size = len(worksheets) |
|
return [row for row in |
|
[handle_sheet(file.name, |
|
sheet, |
|
image_dict, |
|
limit) if worksheets_size == 1 and sheet.title == 'Sheet1' else handle_sheet( |
|
sheet.title, sheet, image_dict, limit) for sheet |
|
in worksheets] if row is not None] |
|
except Exception as e: |
|
maxkb_logger.error(f"Error processing XLSX file {file.name}: {e}, {traceback.format_exc()}") |
|
return [{'name': file.name, 'content': []}] |
|
from PIL import ImageFile |
|
ImageFile.LOAD_TRUNCATED_IMAGES = True |
|
PILImage.MAX_IMAGE_PIXELS = None |
|
def handle_images(deps, archive: ZipFile) -> []: |
|
images = [] |
|
if not PILImage: # Pillow not installed, drop images |
|
return images |
|
for dep in deps: |
|
try: |
|
image_io = archive.read(dep.target) |
|
image = openpyxl_Image(BytesIO(image_io)) |
|
except Exception as e: |
|
maxkb_logger.error(f"Error reading image {dep.target}: {e}, {traceback.format_exc()}") |
|
continue |
|
image.embed = dep.id # 文件rId |
|
image.target = dep.target # 文件地址 |
|
images.append(image) |
|
return images |
|
|
|
|
|
def xlsx_embed_cells_images(buffer) -> {}: |
|
archive = ZipFile(buffer) |
|
# 解析cellImage.xml文件 |
|
deps = get_dependents(archive, get_rels_path("xl/cellimages.xml")) |
|
image_rel = handle_images(deps=deps, archive=archive) |
|
# 工作表及其中图片ID |
|
sheet_list = {} |
|
for item in archive.namelist(): |
|
if not item.startswith('xl/worksheets/sheet'): |
|
continue |
|
key = item.split('/')[-1].split('.')[0].split('sheet')[-1] |
|
sheet_list[key] = parse_element_sheet_xml(fromstring(archive.read(item))) |
|
cell_images_xml = parse_element(fromstring(archive.read("xl/cellimages.xml"))) |
|
cell_images_rel = {} |
|
for image in image_rel: |
|
cell_images_rel[image.embed] = image |
|
for cnv, embed in cell_images_xml.items(): |
|
cell_images_xml[cnv] = cell_images_rel.get(embed) |
|
result = {} |
|
for key, img in cell_images_xml.items(): |
|
all_cells = [ |
|
cell |
|
for _sheet_id, sheet in sheet_list.items() |
|
if sheet is not None |
|
for cell in sheet or [] |
|
] |
|
|
|
image_excel_id_list = [ |
|
cell for cell in all_cells |
|
if isinstance(cell, str) and key in cell |
|
] |
|
# print(key, img) |
|
if img is None: |
|
continue |
|
if len(image_excel_id_list) > 0: |
|
image_excel_id = image_excel_id_list[-1] |
|
f = archive.open(img.target) |
|
img_byte = io.BytesIO() |
|
im = PILImage.open(f).convert('RGB') |
|
im.save(img_byte, format='JPEG') |
|
image = File(id=uuid.uuid7(), file_name=img.path, meta={'debug': False, 'content': img_byte.getvalue()}) |
|
result['=' + image_excel_id] = image |
|
archive.close() |
Proof of concept
maxkb-xlsx-embedded-image-dos-poc-v1.zip
Please attach and extract:
maxkb-xlsx-embedded-image-dos-poc-v1.zip
The included README provides a direct official-image test and a full HTTP test. All traffic
uses an isolated Docker network with no published host port.
The full test performs this request as a default ordinary user:
POST /admin/api/workspace/default/knowledge/{knowledge_id}/document/split
Authorization: Bearer <ordinary-user token>
Content-Type: multipart/form-data
file=@probe-30000.xlsx
limit=4096
Fixture details:
- Outer XLSX size: 7,307 bytes
- Embedded PNG size: 874,852 bytes
- Image dimensions: 30,000 x 30,000
- Pixel count: 900,000,000
- Minimum RGB data: 2,700,000,000 bytes
- XLSX SHA-256:
40aef1fd024a0c4299e06b6a9875208087655884037707aadbef5e02705e5923
The lab used a placeholder embedding-model UUID only when creating its isolated knowledge
base because no external model was configured. The API accepted it. The vulnerable split
and image parsing path does not call the embedding model. A normal deployment can use an
existing accessible knowledge base or a configured embedding model.
Results
Full HTTP application
A 100x100 control file returned HTTP 200 in 0.057 seconds. The container remained healthy
and Docker reported no OOM.
One 7,307-byte attack request produced an empty reply after 1.126 seconds and changed
Docker's OOM state to true. The Gunicorn master and container remained running and replaced
the affected worker.
I then sent three attack requests concurrently while a second default USER repeatedly
called GET /admin/api/user/profile:
| Result |
Value |
| Aggregate attacker input |
21,921 bytes |
| Attacker requests ending with an empty reply |
3 / 3 |
cgroup oom_kill counter |
3 to 5 |
| Second-user profile requests |
60 |
| Second-user HTTP 200 responses |
46 |
| Second-user 500 ms timeouts |
14 |
| Post-attack recovery request |
HTTP 200 in 0.113 s |
The oom_kill counter was already 3 from earlier validation. This burst added two OOM
kills; it did not cause five new kills.
This demonstrates transient cross-user availability impact.
Direct official-image test
The exact image source function was called in a networkless 512 MiB container:
| Fixture |
Result |
| 8,000x8,000 control |
success, about 357 MiB (365,328 KiB) peak RSS |
| 20,000x20,000 attack, Run 1 |
exit 137, OOMKilled=true |
| 20,000x20,000 attack, Run 2 |
exit 137, OOMKilled=true |
| 20,000x20,000 attack, Run 3 |
exit 137, OOMKilled=true |
Impact
Confirmed: a normal authenticated user can cause repeatable OOM kills in shared MaxKB HTTP
workers with very small requests. A short three-request burst caused timeouts for another
authenticated user.
Inferred: a sustained request stream can keep recycling workers and degrade the service.
More memory or more replicas can reduce the effect, but each concurrent decode still has an
unbounded multi-gigabyte memory cost.
A single request kills one worker and ends that request. Sustained cross-user degradation
requires continuing or concurrent requests. I did not measure a long-duration flood or a
deployment other than the 4 GiB all-in-one image.
Suggested fix
- Remove the global
PILImage.MAX_IMAGE_PIXELS = None assignment.
- Check image dimensions before decoding and enforce per-image and aggregate pixel limits.
- Limit XLSX member count, decoded bytes, and expansion ratio.
- Parse untrusted documents in an isolated worker with a hard memory limit.
- Limit concurrent document parsing per user.
- Return a bounded 4xx error and delete partially saved files when a parser limit is reached.
Summary
An ordinary authenticated MaxKB user can upload a small XLSX file containing an embedded
image with very large pixel dimensions.
MaxKB globally disables Pillow's decompression-bomb pixel limit. It then decodes the image,
converts it to RGB, and encodes it as JPEG inside a shared HTTP worker. The upload check uses
only the small outer XLSX file size.
In the official
v2.10.5-ltsall-in-one image, a 7,307-byte XLSX caused an OOM kill. Threeconcurrent requests caused more OOM kills and timeouts for a separate ordinary user. The
service recovered by replacing workers, so this report claims transient availability loss,
not permanent container termination.
Tested version and deployment
v2.10.5-lts01b21db88145278d98bf5e9bd55e6abd6b3aad431panel/maxkb:v2.10.5-ltssha256:23920f160a799844ecf78947a31efcc20f9e796a544ea16d0efb3ee684cb6b76processes
The relevant source files in the image have the same SHA-256 hashes as the tested release.
The same vulnerable
common_handle.pyhash is present in every checked release fromv2.7.0throughv2.10.5-lts.Why this is a security issue
USER, not an administrator./document/splitroute explicitly permits the ordinary-user role.JPEG-encoding and other parser memory are counted.
None.Root cause
The document-split route accepts an ordinary workspace user:
It validates only the uploaded XLSX byte count. The default knowledge-base limit is 100 MiB:
The XLSX handler loads the workbook and processes embedded cell images synchronously:
At module import, MaxKB disables Pillow's decompression-bomb protection:
The embedded image is then fully decoded, RGB-converted, and JPEG-encoded without a pixel,
decoded-byte, expansion-ratio, time, or memory limit:
Source links:
MaxKB/apps/knowledge/views/document.py
Lines 218 to 260 in 01b21db
MaxKB/apps/knowledge/serializers/document.py
Lines 1200 to 1284 in 01b21db
MaxKB/apps/common/handle/impl/text/xlsx_split_handle.py
Lines 103 to 125 in 01b21db
MaxKB/apps/common/handle/impl/common_handle.py
Lines 25 to 27 in 01b21db
MaxKB/apps/common/handle/impl/common_handle.py
Lines 71 to 130 in 01b21db
Proof of concept
maxkb-xlsx-embedded-image-dos-poc-v1.zip
Please attach and extract:
maxkb-xlsx-embedded-image-dos-poc-v1.zipThe included README provides a direct official-image test and a full HTTP test. All traffic
uses an isolated Docker network with no published host port.
The full test performs this request as a default ordinary user:
Fixture details:
40aef1fd024a0c4299e06b6a9875208087655884037707aadbef5e02705e5923The lab used a placeholder embedding-model UUID only when creating its isolated knowledge
base because no external model was configured. The API accepted it. The vulnerable split
and image parsing path does not call the embedding model. A normal deployment can use an
existing accessible knowledge base or a configured embedding model.
Results
Full HTTP application
A 100x100 control file returned HTTP 200 in 0.057 seconds. The container remained healthy
and Docker reported no OOM.
One 7,307-byte attack request produced an empty reply after 1.126 seconds and changed
Docker's OOM state to true. The Gunicorn master and container remained running and replaced
the affected worker.
I then sent three attack requests concurrently while a second default
USERrepeatedlycalled
GET /admin/api/user/profile:oom_killcounterThe
oom_killcounter was already 3 from earlier validation. This burst added two OOMkills; it did not cause five new kills.
This demonstrates transient cross-user availability impact.
Direct official-image test
The exact image source function was called in a networkless 512 MiB container:
OOMKilled=trueOOMKilled=trueOOMKilled=trueImpact
Confirmed: a normal authenticated user can cause repeatable OOM kills in shared MaxKB HTTP
workers with very small requests. A short three-request burst caused timeouts for another
authenticated user.
Inferred: a sustained request stream can keep recycling workers and degrade the service.
More memory or more replicas can reduce the effect, but each concurrent decode still has an
unbounded multi-gigabyte memory cost.
A single request kills one worker and ends that request. Sustained cross-user degradation
requires continuing or concurrent requests. I did not measure a long-duration flood or a
deployment other than the 4 GiB all-in-one image.
Suggested fix
PILImage.MAX_IMAGE_PIXELS = Noneassignment.