Skip to content

[vulnhunter] VulnHunter findings in github/gh-aw #66044

Description

@github-actions

Overview

A single-agent VulnHunter pass (Injection class, verified with adversarial falsification) over the pre-ranked scan scope surfaced one confirmed, PoC-verified finding: a complete sandbox-escape in the custom JavaScript grader runtime, bypassing both its static source blocklist and its Node vm isolation to achieve arbitrary command execution.

No other candidates in scope survived verification as exploitable — the Go CLI code (git/docker/grype/poutine/grant/remote-download command construction) is consistently hardened with argument-array exec.Command calls, explicit ref/path validators (gitutil.ValidateGitRef/ValidateGitPath), and #nosec justifications that checked out on review.

Key finding

ID Component Class Severity Status
VULN-001 pkg/cli/graders_run.cjs (runInline) + pkg/workflow/graders_config.go blocklist Code Injection / Sandbox Escape (CWE-94) High CONFIRMED (runnable PoC)
VULN-001 — Grader script sandbox escape → arbitrary command execution (CWE-94)

Affected files / functions

  • pkg/cli/graders_run.cjs — runInline() (the "sandboxed" JS grader evaluator, embedded via (go/redacted):embed from pkg/cli/graders_run.go:42)
  • pkg/workflow/graders_config.go:454 — the forbiddenPatterns static blocklist, the only other gate on this input
  • Entry point: gh aw graders run <workflow> <grader-id> [run-id] (pkg/cli/graders_command.go → pkg/cli/graders_run.go:66 runGrader → runJavaScriptGrader)

Attacker path / preconditions
A workflow's frontmatter can declare a custom (non-builtin, non-operational-value) grader with an inline script field (graders.<id>.script). gh aw graders run stages the embedded graders_run.cjs and executes it with Node against that script. This command is the execution primitive behind gh-aw's trial/evaluation tooling (pkg/cli/trial_helpers.go), i.e. it is designed to run workflow-declared graders automatically against agent-run artifacts — plausibly in CI contexts holding GH_TOKEN and other secrets (see actions/setup/js/validate_secrets.cjs for the kind of secrets typically present alongside this tooling).

Why it's exploitable

  1. The only static gate on graders.<id>.script is a literal substring blocklist: require(, import(, import , fetch(, eval(, process.exit, child_process, execSync, spawnSync, Function(.
  2. graders_run.cjs::runInline() evaluates the script inside a Node vm context whose sandbox nulls out process, require, Function, eval, global, globalThis, and sets codeGeneration: { strings: false, wasm: false } to block in-context eval/Function string compilation.
  3. However, the trace object passed into that sandbox is built via structuredClone(trace) in the host (unrestricted) Node realm, then only Object.freeze-deep-frozen. Freezing does not sever the prototype chain — any array/object inside trace still resolves .constructor to the host realm's Array/Function, which is not subject to the vm context's codeGeneration.strings:false restriction (that restriction binds to the context where code is compiled, not to functions reached via a leaked host-realm reference).

Proof of concept (falsified against the exact blocklist from graders_config.go — contains none of its literal substrings):

const cp = trace.toolCalls.constructor.constructor("return process")()["mainModule"]["require"]("child_" + "process");
const method = "exec" + "Sync";
return cp[method]("id").toString();

Running this through a faithful reproduction of runInline() prints:

Passes graders_config.go blocklist? true
Command output from inside sandbox: uid=1001(runner) gid=1001(runner) groups=1001(runner),4(adm),100(users),117(docker),999(systemd-journal)

i.e. attacker-authored workflow configuration achieves full native command execution in the process running gh aw graders run, defeating both the blocklist and the vm sandbox at once.

Why this is a real security boundary, not a design assumption
The combined use of a frozen sandbox, nulled globals, codeGeneration.strings:false, and a separate static blocklist whose own error message states "inline grader scripts must be pure functions without side effects" is explicit evidence the project intends graders.<id>.script to be treated as less-trusted, sandboxed input — not developer-equivalent-trust code.

Remediation

  • Don't rely on bare Node vm + a frozen global sandbox for untrusted code execution — Node's own docs state the vm module "is not a security mechanism." Use real isolation (separate process/container with no inherited secrets, or a proper engine such as isolated-vm configured to deep-copy — not just freeze — data entering the context).
  • At minimum, any data handed into a vm context must be constructed inside that context (e.g. via vm.runInContext(JSON.stringify(data), context)), so its prototype chain binds to the restricted realm rather than the host realm.
  • Replace the substring blocklist with an AST-based allowlist analyzer, and stop presenting it as a security control if the vm sandbox is the intended actual boundary.

Next actions

  • Treat this as the primary remediation item; it allows a workflow-authored grader script to achieve arbitrary command execution wherever gh aw graders run executes it with live credentials.
  • Re-run the blocklist/vm-sandbox PoC after any fix to confirm the escape is closed (not just the specific PoC payload).

Generated by 🛡️ Daily VulnHunter Scan · claude · sonnet50 · 251.3 AIC · ⌖ 28.7 AIC · ⊞ 6.5K · ◷

Activity

  1. github-actions commented on Oct 7, 2026

    @github-actions
    ContributorAuthor

    🍪 Issue Monster selected this for Copilot

    I've identified this issue as a good candidate for automated resolution and requested assignment to the Copilot coding agent.

    If assignment succeeds, the Copilot coding agent will analyze the issue and create a pull request with the fix.

    Om nom nom! 🍪

    🍪 Om nom nom by Issue Monster · pi · gpt54 · 28.3 AIC · ⌖ 10.3 AIC · ⊞ 13.2K · ◷

  2. github-actions commented on Oct 7, 2026

    @github-actions
    ContributorAuthor

    🍪 Issue Monster selected this for Copilot

    I've identified this issue as a good candidate for automated resolution and requested assignment to the Copilot coding agent.

    If assignment succeeds, the Copilot coding agent will analyze the issue and create a pull request with the fix.

    Om nom nom! 🍪

    🍪 Om nom nom by Issue Monster · pi · gpt54 · 15.1 AIC · ⌖ 8.78 AIC · ⊞ 13K · ◷

  3. github-actions commented on Oct 7, 2026

    @github-actions
    ContributorAuthor

    Caution

    agentic threat detected
    Threat detection flagged this output in warn mode. Manual review is REQUIRED before any follow-up automation.

    Details

    Potential security threats were detected in the agent output.

    Review the workflow run logs for details.

    🍪 Issue Monster selected this for Copilot

    I've identified this issue as a good candidate for automated resolution and requested assignment to the Copilot coding agent.

    If assignment succeeds, the Copilot coding agent will analyze the issue and create a pull request with the fix.

    Om nom nom! 🍪

    🍪 Om nom nom by Issue Monster · pi · gpt54 · 19.7 AIC · ⌖ 14.8 AIC · ⊞ 12.9K · ◷

  4. github-actions commented on Oct 8, 2026

    @github-actions
    ContributorAuthor

    🍪 Issue Monster selected this for Copilot

    I've identified this issue as a good candidate for automated resolution and requested assignment to the Copilot coding agent.

    If assignment succeeds, the Copilot coding agent will analyze the issue and create a pull request with the fix.

    Om nom nom! 🍪

    🍪 Om nom nom by Issue Monster · pi · gpt54 · 21.1 AIC · ⌖ 8.92 AIC · ⊞ 14.3K · ◷

  5. github-actions commented on Oct 8, 2026

    @github-actions
    ContributorAuthor

    🍪 Issue Monster selected this for Copilot

    I've identified this issue as a good candidate for automated resolution and requested assignment to the Copilot coding agent.

    If assignment succeeds, the Copilot coding agent will analyze the issue and create a pull request with the fix.

    Om nom nom! 🍪

    🍪 Om nom nom by Issue Monster · pi · gpt54 · 12 AIC · ⌖ 9.76 AIC · ⊞ 14.5K · ◷

  6. github-actions commented on Oct 9, 2026

    @github-actions
    ContributorAuthor

    🍪 Issue Monster selected this for Copilot

    I've identified this issue as a good candidate for automated resolution and requested assignment to the Copilot coding agent.

    If assignment succeeds, the Copilot coding agent will analyze the issue and create a pull request with the fix.

    Om nom nom! 🍪

    🍪 Om nom nom by Issue Monster · pi · gpt54 · 15.8 AIC · ⌖ 8.87 AIC · ⊞ 14.1K · ◷

  7. github-actions commented on Oct 10, 2026

    @github-actions
    ContributorAuthor

    🍪 Issue Monster selected this for Copilot

    I've identified this issue as a good candidate for automated resolution and requested assignment to the Copilot coding agent.

    If assignment succeeds, the Copilot coding agent will analyze the issue and create a pull request with the fix.

    Om nom nom! 🍪

    🍪 Om nom nom by Issue Monster · pi · gpt54 · 11.3 AIC · ⌖ 8.53 AIC · ⊞ 14K · ◷

  8. pelikhan commented on Oct 11, 2026

    @pelikhan
    Collaborator

    Reviewed against completed essentials issue #67454 (#67454) and merged PR #67475 (#67475).

    The merged fix isolates inline graders in separate restricted Node processes, strips inherited environment capabilities, and adds regressions for process/require access and the reported command-execution escape.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions