Skip to content

Agent Persona Exploration - 2026-10-09 #67228

Description

@github-actions

Persona Overview

  • Agent: agentic-workflows custom agent (via persona-evaluator sub-agent)
  • Personas This Run: Program Manager, Designer, Legal / Compliance
  • Scenarios Tested: 4 of 6 generated (PM-A, DS-A, LC-A, LC-B)
  • Average Quality Score: 4.6/5.0

Key Findings

  • All 4 invocations succeeded with no tool-call errors.
  • The agent consistently produced correctly-scoped pull_request/schedule/issues triggers with paths:/types: filters matched to the task, and defaulted to minimal read-only permissions with writes routed exclusively through safe-outputs.
  • Two of four sub-agent invocations (LC-A, LC-B) declined to self-score, citing an instruction embedded in their own agent definition that reserves scoring for "the parent evaluator" — this conflicts with this workflow's direct per-scenario invocation pattern. Same issue was seen on a prior run's IW-1 scenario; it recurred here on 2 of 4 calls, suggesting it's systemic rather than a one-off.
  • The agent showed good judgment on not over-provisioning: it explicitly reasoned about and rejected playwright for a static-diff accessibility check (DS-A), and reasoned network access down to "omit entirely" when the compliance comparison target was in-repo (LC-B).
  • Non-obvious safe-output detail handled correctly: PM-A needed add-comment with target: "*" (not the default triggering) to post on the parent milestone issue rather than the issue that fired the trigger — the agent caught this without being prompted.

Top Patterns

  1. Triggers: pull_request with extension/dir paths: scoping (DS-A, LC-A) for diff-reactive checks; schedule + workflow_dispatch (LC-B) for periodic policy audits; issues: [closed, reopened] (PM-A) for state-change alerts.
  2. Tools: github MCP in gh-proxy mode with the default toolset was the near-universal choice; sandboxed bash/edit left at their defaults rather than redeclared. No scenario reached for playwright or cli-proxy unnecessarily.
  3. Security practices: every recommendation used contents: read (+ narrow issues/pull-requests: read as needed), zero write scopes on the agent job, and safe-outputs as the sole write path. Network access was denied by default in 3/4 cases and scoped to "GitHub API only" in the fourth.
View High Quality Responses (Top 2)
  • DS-A (Designer, accessibility check) — 4.8 avg. Diff-scoped review (not whole-file), explicit rejection of a browser-rendering tool as unnecessary, and clear fix-suggestion format per finding. Only gap: no sticky-comment/dedup strategy across repeated pushes.
  • LC-B (Legal/Compliance, scheduled disclosure-contact audit) — 4.8 avg. Proposed a skip-if-match previous-result strategy keyed on an open compliance-labeled issue to prevent duplicate findings piling up — a thoughtful addition not explicitly requested.
View Areas for Improvement (Top 2)
  • Self-scoring conflict recurs: LC-A and LC-B both declined to self-score due to embedded agent-definition guidance that assumes a different invocation context (dynamic workflow materialization with a parent evaluator). This is now observed across two separate runs (this run and a prior run's IW-1) — worth resolving so scoring behavior is consistent regardless of how the sub-agent is invoked.
  • Comment-update strategy underspecified: both PR-comment scenarios (DS-A, LC-A) would benefit from explicit guidance on updating a single comment across pushes vs. posting fresh ones, to avoid notification noise on iterative PRs.

Recommendations

  1. Clarify in .github/aw/create-agentic-workflow.md (or wherever the sub-agent's self-scoring instructions live) that direct ad hoc invocation should always self-score — the "parent evaluator will score" caveat should only apply when the agent is dynamically materialized inside its own orchestrated workflow, not when invoked standalone.
  2. Add a short authoring note to .github/aw/create-agentic-workflow.md on PR-triggered comment workflows: recommend a sticky/updatable comment pattern (vs. one-comment-per-push) as the default for diff-reactive checks like accessibility or license scans, to reduce noise across synchronize events.
  3. Continue to reinforce the "omit network access when comparison target is in-repo" and "reject tools not needed for static analysis" judgment calls observed this run — these were correct but worth calling out explicitly in authoring guidance as a named pattern so they're reliably repeated.

Generated by 🎭 Agent Persona Explorer · claude · agent · 196.9 AIC · ⌖ 0.717 AIC · ⊞ 6.6K · ◷

Activity

  1. github-actions commented on Oct 10, 2026

    @github-actions
    ContributorAuthor

    This issue is being closed as outdated. A newer issue has been created: #67465

    View newer issue


    This action was performed automatically by the Agent Persona Explorer workflow.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions