Skip to content

Threat detection inherits the main engine's provider env when it runs on a different engine #66203

Description

@theletterf

When the main agent and threat detection use different engines, the detection step still inherits the main engine's engine.env. Provider settings such as ANTHROPIC_BASE_URL or OPENAI_BASE_URL then reach a detector that can't use them, and detection fails before it analyzes anything.

Our docs review runs the agent on the Claude engine through OpenRouter, with threat detection on Copilot:

engine:
  id: claude
  env:
    ANTHROPIC_BASE_URL: https://openrouter.ai/api
    ANTHROPIC_API_KEY: ${{ secrets.OPENROUTER_API_KEY }}
safe-outputs:
  threat-detection:
    engine:
      id: copilot
      model: sonnet

The compiled detection step gets ANTHROPIC_BASE_URL=https://openrouter.ai/api, and its AWF config sets an anthropic target on openrouter.ai. The firewall health check then stops the container before threat-detect starts:

[health-check][ERROR] Cannot connect to Anthropic API proxy at https://openrouter.ai/api
Threat Detection Engine Failure — The analysis engine could not complete.

Because detection continues on error and safe_outputs only checks the detection job result, every run posted its outputs without a scan, and nothing showed it. We found this in all our docs review runs since at least early October, and in a second workflow that runs Codex through OpenRouter (Cannot connect to OpenAI API proxy).

The cause is mergeThreatDetectionEngineEnv in pkg/workflow/threat_detection_helpers.go, which copies the full main engine.env into detection. It's the same in v0.89.17, v0.91.1, and main. #54047 fixed the model half of this split-engine setup through #54366. This is the env half. It also means the main engine's API key reaches the detection step.

The workaround is an empty override per key on the detection engine (env: { OPENAI_BASE_URL: "" }), which works because hasCustomTarget checks for a non-empty value. We verified it for both the Codex and the Claude case: with the overrides, detection starts and writes detection_result.json with a verdict. But you only find it after detection has failed silently.

Ask: when threat-detection.engine.id differs from the main engine, don't inherit the main engine's provider settings (*_BASE_URL, API keys, custom headers, model aliases), or at least warn at compile time. Separately, it would help if a detection run with no verdict were visible, for example as a warning on the conclusion job.

Seen on gh-aw v0.89.17.


Generated by Claude Code

Activity

  1. locked and limited conversation to collaborators on Oct 6, 2026
  2. unlocked this conversation on Oct 6, 2026
  3. theletterf commented on Oct 7, 2026

    @theletterf
    Author

    @pelikhan Not sure if more recent releases fixed this.

  4. locked and limited conversation to collaborators on Oct 7, 2026
  5. unlocked this conversation on Oct 7, 2026
  6. pelikhan commented on Oct 7, 2026

    @pelikhan
    Collaborator

    cooking

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions