Skip to content

security guide is outdated #1071

Description

@pelikhan

It was written pre safe-outputs and needs a refresh.

https://githubnext.github.io/gh-aw/guides/security/

Activity

  1. pelikhan commented on Oct 6, 2025

    @pelikhan
    CollaboratorAuthor

    /poem

  2. pelikhan commented on Oct 6, 2025

    @pelikhan
    CollaboratorAuthor

    /scout


    🤖 Workflow run triggered by this comment

  3. github-actions commented on Oct 6, 2025

    @github-actions
    Contributor

    🔍 Scout Research Report

    Triggered by @pelikhan

    Executive Summary

    The security guide at https://githubnext.github.io/gh-aw/guides/security/ was written before the introduction of the safe-outputs feature and needs significant updates to reflect modern security best practices. Based on comprehensive research into agentic AI security, GitHub Actions hardening, and zero-trust architecture patterns, the guide should be restructured to emphasize safe-outputs as the primary security pattern rather than a supplementary feature.

    The current guide mentions safe-outputs briefly but doesn't adequately convey that this architecture represents a fundamental security paradigm shift: separating AI processing (read-only permissions) from GitHub API write operations (isolated, validated jobs). This separation of concerns is critical for defending against prompt injection, unauthorized actions, and data exfiltration in agentic workflows.

    Research Findings

    1. Current State Analysis

    Existing Security Guide Structure:
    The current security guide (23,944 bytes, last updated Oct 6, 2024) covers:

    • Threat model and security principles
    • Workflow permissions and triggers
    • Human-in-the-loop patterns
    • Strict mode validation
    • MCP tool hardening and network isolation
    • Sanitized context text usage
    • Engine-specific security considerations

    Safe-Outputs Coverage (Lines 380-386):
    Currently only 6 lines briefly mention safe-outputs:

    ### Safe Outputs Security Model
    
    Safe outputs provide a security-first approach to GitHub API interactions by separating AI processing from write operations...
    
    See the [Safe Outputs Reference](/gh-aw/reference/safe-outputs/) for complete configuration details.

    This drastically understates the importance of safe-outputs as the recommended security architecture for agentic workflows.

    2. Industry Security Best Practices for Agentic AI

    Prompt Injection Defense (2024-2025 Threat Landscape):

    • Escalating Threat: Research shows ~3.3 AI-agent security incidents per day across 3,000 U.S. companies in 2025, with ~1.3/day tied to prompt injection or agent abuse
    • OWASP Top 10 for LLMs: LLM01:2025 identifies Prompt Injection as the rejig docs #1 vulnerability
    • Zero-Click Exploits: New research (EchoLeak) demonstrates sophisticated prompt injection can occur without user interaction
    • Defense-in-Depth Required: No single measure suffices; layered defenses combining input sanitization, output validation, and privilege separation are essential

    Key Defense Patterns Identified:

    1. Separation of Planning and Execution: The "Plan-then-Execute" pattern separates AI reasoning (untrusted) from action execution (validated)
    2. Policy Proxy Pattern: Dedicated systems validate AI outputs before execution
    3. Output Validation Gates: Treat all AI-generated content as untrusted until verified
    4. Least Privilege Access: Grant only minimum permissions required for specific functions

    3. GitHub Actions Security Hardening

    Permission Management Crisis:

    • 86% of GitHub Actions workflows don't limit token permissions (Legit Security 2024 Report)
    • Over 2,500 workflows execute untrusted code without proper sandboxing
    • 7,000+ workflows interpolate untrusted input directly into execution contexts

    Best Practices from Industry Research:

    1. Least Privilege by Default: Set minimal permissions at job level, not workflow level
    2. Job Isolation: Separate jobs for build, test, and deploy with explicit dependencies
    3. Secret Management: Never hard-code secrets; use GitHub Secrets with environment scoping
    4. Supply Chain Security: Pin all dependencies to immutable SHAs
    5. Branch Protection: Require PR reviews for all workflow changes

    GitHub Actions Workflow Hijacking Patterns (2025):
    Recent attacks demonstrate how malicious workflows can exfiltrate secrets through:

    • Encoding secrets with ${{ toJSON(secrets) }}
    • Using curl/base64 to transmit to external servers
    • Exploiting overly broad permissions

    4. Zero Trust Architecture for AI Workflows

    Core Principles (NIST & Industry Standards):

    1. Never Trust, Always Verify: Assume breach; validate every access request
    2. Least Privilege Access: Grant minimum rights necessary for specific functions
    3. Microsegmentation: Isolate resources based on context and risk levels
    4. Continuous Monitoring: Real-time verification and anomaly detection
    5. Assume Breach Posture: Design systems to minimize impact when breaches occur

    Application to Agentic Workflows:

    • Identity Verification: Validate AI agent actions through policy gates
    • Context-Based Access: Permissions should be request-specific, not blanket
    • Audit Trail: Every action must be traceable and reviewable
    • Granular Permissions: Separate read from write, planning from execution

    5. Safe-Outputs as Zero Trust Implementation

    How Safe-Outputs Embodies Zero Trust:

    Separation of Concerns:

    • AI Processing Job: Runs with minimal read-only permissions (contents: read, actions: read)
    • Write Operations Job: Separate job with specific write permissions, activated only after AI completes successfully
    • Validation Layer: Automatic output parsing and threat detection between jobs

    Least Privilege in Practice:

    # AI job: minimal permissions
    permissions:
      contents: read
      actions: read
    
    # Separate job handles writes
    safe-outputs:
      create-issue:      # Issue creation in isolated job
      add-comment:       # Comment creation in isolated job
      create-pull-request: # PR creation in isolated job

    Defense Against Common Attacks:

    1. Prompt Injection → Limited Blast Radius: Even if AI is manipulated, it can't directly write to GitHub
    2. Secret Exfiltration → No Write Access: AI job can't create issues/PRs to leak secrets
    3. Unauthorized Actions → Validation Gates: Output processing validates before execution
    4. Code Injection → Sandboxed Output: Generated code reviewed before PR creation

    Threat Detection Integration:
    Safe-outputs includes automatic threat detection analyzing:

    • Prompt injection attempts in AI output
    • Secret leaks in generated content
    • Malicious code patterns in patches
    • Using AI-powered analysis with workflow context to reduce false positives

    6. Comparative Analysis: Direct Permissions vs Safe-Outputs

    Traditional Pattern (Risky):

    permissions:
      contents: write
      issues: write
      pull-requests: write
      
    # AI has direct write access throughout execution

    Vulnerabilities:

    • Single point of failure
    • AI can be manipulated to perform unauthorized writes
    • No validation layer between decision and action
    • Difficult to audit what AI intended vs. what it executed

    Safe-Outputs Pattern (Recommended):

    permissions:
      contents: read
      actions: read
    
    safe-outputs:
      create-issue:
        labels: [ai-generated]
        max: 5

    Security Benefits:

    • Permission Separation: AI never has write access
    • Validation Layer: Output parsed and validated before execution
    • Audit Trail: Clear separation between AI reasoning and GitHub actions
    • Threat Detection: Automatic scanning for malicious content
    • Rate Limiting: Max counts prevent runaway automation
    • Explicit Intent: Output format requires structured, reviewable data

    7. AI Output Validation and Sanitization

    Input/Output Security for AI Agents:

    Input Sanitization (Already Implemented):
    The guide correctly emphasizes needs.activation.outputs.text which provides:

    • @mention neutralization
    • Bot trigger escaping
    • XML tag conversion
    • URI filtering (only HTTPS from trusted domains)
    • Content size limits (0.5MB, 65k lines)
    • Control character removal

    Output Validation (Safe-Outputs Provides):
    Research shows output validation is equally critical:

    • Schema Validation: Ensure AI output matches expected format
    • Content Filtering: Block sensitive data, secrets, malicious URLs
    • Anomaly Detection: Flag unusual patterns or behaviors
    • Policy Enforcement: Verify actions align with workflow intent

    Industry Patterns:

    1. Dual-Model Validation: Secondary AI scans primary AI output for threats
    2. Constrained Decoding: Enforce JSON schemas, apply regex/stop-sequences
    3. Provenance Tracking: Tag and track source of all content
    4. Runtime Monitoring: Continuous behavioral analysis during execution

    8. Documentation Structure Recommendations

    Current Guide Issues:

    1. Buried Lead: Safe-outputs mentioned at line 380 of 508 (75% through document)
    2. Minimal Coverage: Only 6 lines for the most important security feature
    3. No Examples: Lacks concrete before/after comparisons
    4. No Threat Modeling: Doesn't explain what attacks safe-outputs prevents
    5. Supplementary Framing: Presented as optional feature, not primary pattern

    Recommended Structure:

    Early Prominence (Top 25% of Guide):

    ## Security Architecture: Safe-Outputs First
    
    ### The Safe-Outputs Security Model
    
    GitHub Agentic Workflows implements a zero-trust security architecture 
    through the `safe-outputs` system. This is the RECOMMENDED approach for 
    all agentic workflows that need to create issues, comments, or pull requests.
    
    **Core Principle:** Separate AI processing (read-only) from GitHub API 
    writes (validated, isolated jobs).
    
    [Detailed explanation with threat model, diagrams, examples]
    
    ### When to Use Safe-Outputs vs Direct Permissions
    
    Use safe-outputs for: [95% of workflows]
    Consider direct permissions only for: [edge cases with justification]

    Concrete Examples:

    • Before/after workflow comparisons
    • Attack scenarios that safe-outputs prevents
    • Real-world security incidents and how safe-outputs would have prevented them
    • Performance and usability benefits

    Visual Aids:

    • Architecture diagrams showing job separation
    • Flow charts for output validation pipeline
    • Threat model mappings

    Recommendations

    Immediate Actions (High Priority)

    1. Restructure Security Guide:

      • Move safe-outputs section to top 25% of document (after "Before You Begin")
      • Expand from 6 lines to comprehensive section (200-300 lines)
      • Add subsections:
        • Security Architecture Overview
        • How Safe-Outputs Implements Zero Trust
        • Threat Model and Attack Prevention
        • Available Output Types and Security Features
        • Migration Guide (Direct Permissions → Safe-Outputs)
        • When NOT to Use Safe-Outputs (rare edge cases)
    2. Add Concrete Examples:

      • Side-by-side workflow comparisons (unsafe vs. safe)
      • Attack scenarios with mitigation demonstrations
      • Real-world use cases for each safe-output type
      • Security audit checklist
    3. Enhance Threat Detection Documentation:

      • Explain automatic threat scanning in safe-outputs
      • Document what threats are detected
      • Show example threat detection logs
      • Guidance on handling detected threats
    4. Create Visual Documentation:

      • Architecture diagram: AI job → Output → Validation → Write job
      • Permission flow diagram
      • Threat model mapping diagram

    Content Additions (Medium Priority)

    1. Comparative Security Analysis:

      • Table comparing security postures
      • Risk assessment matrix
      • Decision tree for choosing security patterns
    2. Migration Guide:

      • Step-by-step conversion from direct permissions
      • Common migration pitfalls
      • Testing strategies for migrated workflows
    3. Security Validation Checklist:

      • Pre-deployment security review items
      • Runtime monitoring recommendations
      • Incident response procedures

    Supporting Documentation (Lower Priority)

    1. Cross-Reference Updates:

      • Ensure all workflow examples use safe-outputs
      • Update getting-started guides
      • Add security notes to frontmatter reference
    2. Tutorials and Videos:

      • "Securing Your First Agentic Workflow" tutorial
      • Video walkthrough of safe-outputs architecture
      • Security workshop materials
    3. Community Resources:

      • Security best practices checklist
      • Example secure workflows repository
      • Security audit template

    Key Sources

    Academic and Research Papers

    • NIST AI 100-2e2023: "Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations"
    • OWASP LLM01:2025: "Prompt Injection" - Top security risk for LLM applications
    • ArXiv 2509.08646: "Architecting Resilient LLM Agents: A Guide to Secure Plan-then-Execute Patterns"
    • ArXiv 2509.10540: "EchoLeak: The First Real-World Zero-Click Prompt Injection Exploit"
    • Medium (2024): "When Hacks Go Awry: The Rising Tide of AI Prompt Injection Attacks" - Documents ~3.3 incidents/day in 2025

    Industry Security Standards

    • Zero Trust Architecture Technology Book (GSA, March 2025): Comprehensive ZTA implementation guide
    • Legit Security 2024 Report: "86% of GitHub Actions workflows don't limit token permissions"
    • Microsoft Security (2024): Zero Trust principles and implementation
    • SANS Institute: AI Security Guidelines - LLM guardrails and validation

    GitHub Actions Security

    • Graphite Dev: "GitHub Actions permissions" - Principle of least privilege implementation
    • StatusNeo: "Best Practices for Securing GitHub Actions Workflows"
    • Wiz Security: "Advanced GitHub Security Best Practices Cheat Sheet"
    • Aembit.io: "Secretless Access for GitHub Actions and Workflows"

    AI Agent Security

    • DataCamp: "AI Security: A Comprehensive Guide" - Input validation and sanitization
    • Dev.to: "Hardening Your AI Agent Against Prompt Injection via MCP"
    • Galileo AI: "How to Test AI Agents Effectively" - Safety and security evaluation
    • Varonis: "Detecting Agentic AI Threats with Agentic AI"

    Relevant Documentation

    • Current Security Guide: /docs/src/content/docs/guides/security.md (lines 380-386)
    • Safe-Outputs Reference: /docs/src/content/docs/reference/safe-outputs.md (33,599 bytes)
    • Example Workflow: .github/workflows/scout.md (demonstrates safe-outputs usage)

    Suggested Next Steps

    1. Draft Updated Security Guide (Week 1):

      • Create outline with new structure
      • Write safe-outputs comprehensive section
      • Develop attack scenario examples
      • Create architecture diagrams
    2. Internal Review (Week 2):

      • Security team review
      • Technical accuracy validation
      • Example workflow testing
      • Diagram review
    3. Community Feedback (Week 3):

      • Beta documentation release
      • Solicit feedback from early adopters
      • Security researcher review
      • Accessibility and clarity testing
    4. Final Publication (Week 4):

      • Incorporate feedback
      • Publish updated guide
      • Update all cross-references
      • Announce changes in release notes
    5. Ongoing Maintenance:

      • Monitor for new security research
      • Update as safe-outputs features expand
      • Collect real-world security incidents and lessons learned
      • Quarterly security guide reviews

    Critical Message for Documentation Update

    The fundamental insight that must be conveyed:

    Safe-outputs is not an optional feature or nice-to-have security enhancement. It is the recommended security architecture for GitHub Agentic Workflows. By separating AI processing (which should be treated as untrusted) from GitHub API write operations (which are validated and isolated), safe-outputs implements zero-trust principles that defend against prompt injection, unauthorized actions, and data exfiltration.

    When building agentic workflows that create issues, comments, or pull requests, start with safe-outputs. Only consider direct write permissions for rare edge cases where the architecture genuinely doesn't fit—and document the security trade-offs explicitly.

    This shift from "here's a feature" to "this is the security model" is critical for helping users build secure agentic workflows by default.

    AI generated by Scout

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions