Skip to content

fix(helpers): correct data: URL base64 size estimate for maxContentLength guard - #11061

Merged
jasonsaayman merged 4 commits into
axios:v1.xfrom
spokodev:fix/datauri-base64-size-estimate
Jul 14, 2026
Merged

jasonsaayman merged 4 commits into
axios:v1.xfrom
spokodev:fix/datauri-base64-size-estimate

Conversation

@spokodev

@spokodev spokodev commented Jul 1, 2026 •

Copy link
Copy Markdown
Contributor

The base64 branch of estimateDataURLDecodedBytes() under-reports the actual decoded byte length, which weakens the maxContentLength guard the Node http and fetch adapters apply to data: URLs before decoding.

Steps to reproduce

fromDataURI() decodes the raw body with Buffer.from(body, 'base64'), so the estimate must never be lower than the length that call produces. It currently is:

import estimateDataURLDecodedBytes from './lib/helpers/estimateDataURLDecodedBytes.js';

// hex digits inside %XX are valid base64 chars, so Node keeps them as data
const body = 'QQ' + '%41'.repeat(4000);
const url = 'data:application/octet-stream;base64,' + body;

const actual = Buffer.from(body, 'base64').length; // 6001
const estimate = estimateDataURLDecodedBytes(url); // 3000 (before this change)

console.log({ estimate, actual }); // estimate < actual  -> guard bypassed

// smaller case
const url2 = 'data:text/plain;base64,TQ%3D%3D';
Buffer.from('TQ%3D%3D', 'base64').length;   // 4  (EXPECTED)
estimateDataURLDecodedBytes(url2);           // 1  (ACTUAL, before this change)

Because the estimate can be roughly half of the real decoded size, a body up to about twice maxContentLength passes the check and is then fully decoded into a Buffer.

Root cause

The base64 branch assumed the body was percent-decoded before base64 decoding: it subtracted 2 for every %XX escape and treated a trailing %3D as == padding. Node's base64 decoder does neither. It never percent-decodes, it ignores any character outside the base64 alphabet (a bare % is dropped), and it keeps the surrounding hex digits as ordinary base64 data. %3D is not padding: the % is dropped and 3D decode as data.

Fix

Rewrite the base64 branch to mirror Node's decoder: count only the significant base64/base64url alphabet characters, ignore everything else, stop at the first literal =, then derive bytes from the group count plus the trailing remainder (2 chars -> 1 byte, 3 chars -> 2 bytes). The estimate now satisfies the invariant estimate >= Buffer.from(body, 'base64').length, so it may over-estimate but never under-report.

Tests

The existing unit test for TQ%3D%3D asserted an incorrect oracle value of 1; the real decoded length is 4, and the assertion is corrected to compare against Buffer.from(body, 'base64').length. A regression test is added for percent-embedded base64 bodies that asserts the >= invariant.

  • npm run test:vitest:unit -- tests/unit/estimateDataURLDecodedBytes.test.js: 8/8 pass with the fix; the two updated assertions fail on the current branch without it.
  • Full unit suite npm run test:vitest:unit: 948/948 pass.
  • npx eslint lib/helpers/estimateDataURLDecodedBytes.js: clean.

🏄


Summary by cubic

Fixes base64 data: URL size under-estimation so maxContentLength is enforced correctly in both Node HTTP and fetch. For fetch, the size check now ignores URL fragments.

Description

  • Summary of changes

    • Split estimator into estimateDataURLDecodedBytes() (fetch) and estimateDataURLBufferAllocation() (Node HTTP).
    • Rewrote base64 logic to match each adapter:
      • Fetch: percent-decodes first, validates forgiving-base64/base64url, then computes decoded bytes; ignores URL fragments.
      • Node: computes Buffer allocation bound from the raw body length and padding, including ignored chars and content after =.
    • Updated lib/adapters/http.js to use estimateDataURLBufferAllocation().
  • Reasoning

    • Node’s Buffer.from(body, 'base64') never percent-decodes; old logic undercounted when %XX appeared, letting oversized bodies pass.
    • Fetch rejects malformed base64 and processes data: URLs without fragments; estimators now mirror these behaviors.
  • Additional context

    • Supports base64url (-, _).
    • Estimators may overcount but never undercount for their adapter.

Docs

Suggest updating /docs/ to explain:

  • How data: URL size checks differ: fetch (percent-decoded decoded size, ignores fragments) vs Node HTTP (raw base64 Buffer allocation, counts ignored chars and post-= input).
  • Estimators may over-estimate but never under-estimate.

Testing

  • Updated unit tests for both estimators and adapters:
    • Fetch: rejects percent-embedded base64 over maxContentLength, allows percent-encoded padding at the limit, ignores fragments.
    • Node HTTP: allows at the allocation limit; rejects when ignored chars or trailing content increase the raw allocation; counts input after =.
  • No new fixtures required.

Semantic version impact

Patch: bug fix tightening maxContentLength enforcement for data: URLs without public API changes.

Written for commit 1940751. Summary will update on new commits.

Review in cubic

…ngth guard

estimateDataURLDecodedBytes is the pre-decode size guard the Node http and
fetch adapters use to reject data: URLs whose decoded body exceeds
config.maxContentLength before allocating a Buffer. Its base64 branch assumed
the body was percent-decoded before base64 decoding: it subtracted 2 for each
%XX escape and treated a trailing %3D as == padding.

fromDataURI decodes the raw body with Buffer.from(body, 'base64'), which never
percent-decodes. Node's decoder ignores characters outside the base64 alphabet
(including a bare %) and keeps the surrounding hex digits as real base64 data,
and %3D is not padding. A base64 body sprinkled with %XX (hex digits that are
valid base64 chars) was therefore under-estimated by up to ~2x, letting a body
larger than maxContentLength through.

Rewrite the base64 branch to mirror Node's decoder: count only significant
base64/base64url alphabet characters, ignore everything else, stop at the first
literal =, then compute bytes from the group count and remainder. The estimate
now satisfies the invariant estimate >= Buffer.from(body, 'base64').length.

Also correct the existing unit oracle for TQ%3D%3D (was 1, actual decoded
length is 4) and add a regression test for percent-embedded base64 bodies.
@spokodev
spokodev requested a review from jasonsaayman as a code owner July 1, 2026 21:37

@cubic-dev-ai cubic-dev-ai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No issues found across 2 files

Confidence score: 5/5

  • Automated review surfaced no issues in the provided summaries.
  • No files require special attention.

Re-trigger cubic

@Ziiyodullayevv Ziiyodullayevv left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Important fix for base64 size estimation. Correct data: URL size calculation ensures maxContentLength guard works properly for embedded content.

@jasonsaayman jasonsaayman added the commit::fix The PR is related to a bugfix label Jul 13, 2026

@cubic-dev-ai cubic-dev-ai Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All reported issues were addressed across 6 files (changes from recent commits).

Reply with feedback, questions, or to request a fix.

Re-trigger cubic

Comment thread lib/helpers/estimateDataURLDecodedBytes.js Outdated
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

commit::fix The PR is related to a bugfix

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants