Persona Overview
- Agent: agentic-workflows custom agent (via
persona-evaluator sub-agent)
- Personas This Run: Program Manager, Designer, Legal / Compliance
- Scenarios Tested: 4 of 6 generated (PM-A, DS-A, LC-A, LC-B)
- Average Quality Score: 4.6/5.0
Key Findings
- All 4 invocations succeeded with no tool-call errors.
- The agent consistently produced correctly-scoped
pull_request/schedule/issues triggers with paths:/types: filters matched to the task, and defaulted to minimal read-only permissions with writes routed exclusively through safe-outputs.
- Two of four sub-agent invocations (LC-A, LC-B) declined to self-score, citing an instruction embedded in their own agent definition that reserves scoring for "the parent evaluator" — this conflicts with this workflow's direct per-scenario invocation pattern. Same issue was seen on a prior run's IW-1 scenario; it recurred here on 2 of 4 calls, suggesting it's systemic rather than a one-off.
- The agent showed good judgment on not over-provisioning: it explicitly reasoned about and rejected
playwright for a static-diff accessibility check (DS-A), and reasoned network access down to "omit entirely" when the compliance comparison target was in-repo (LC-B).
- Non-obvious safe-output detail handled correctly: PM-A needed
add-comment with target: "*" (not the default triggering) to post on the parent milestone issue rather than the issue that fired the trigger — the agent caught this without being prompted.
Top Patterns
- Triggers:
pull_request with extension/dir paths: scoping (DS-A, LC-A) for diff-reactive checks; schedule + workflow_dispatch (LC-B) for periodic policy audits; issues: [closed, reopened] (PM-A) for state-change alerts.
- Tools:
github MCP in gh-proxy mode with the default toolset was the near-universal choice; sandboxed bash/edit left at their defaults rather than redeclared. No scenario reached for playwright or cli-proxy unnecessarily.
- Security practices: every recommendation used
contents: read (+ narrow issues/pull-requests: read as needed), zero write scopes on the agent job, and safe-outputs as the sole write path. Network access was denied by default in 3/4 cases and scoped to "GitHub API only" in the fourth.
View High Quality Responses (Top 2)
- DS-A (Designer, accessibility check) — 4.8 avg. Diff-scoped review (not whole-file), explicit rejection of a browser-rendering tool as unnecessary, and clear fix-suggestion format per finding. Only gap: no sticky-comment/dedup strategy across repeated pushes.
- LC-B (Legal/Compliance, scheduled disclosure-contact audit) — 4.8 avg. Proposed a
skip-if-match previous-result strategy keyed on an open compliance-labeled issue to prevent duplicate findings piling up — a thoughtful addition not explicitly requested.
View Areas for Improvement (Top 2)
- Self-scoring conflict recurs: LC-A and LC-B both declined to self-score due to embedded agent-definition guidance that assumes a different invocation context (dynamic workflow materialization with a parent evaluator). This is now observed across two separate runs (this run and a prior run's IW-1) — worth resolving so scoring behavior is consistent regardless of how the sub-agent is invoked.
- Comment-update strategy underspecified: both PR-comment scenarios (DS-A, LC-A) would benefit from explicit guidance on updating a single comment across pushes vs. posting fresh ones, to avoid notification noise on iterative PRs.
Recommendations
- Clarify in
.github/aw/create-agentic-workflow.md (or wherever the sub-agent's self-scoring instructions live) that direct ad hoc invocation should always self-score — the "parent evaluator will score" caveat should only apply when the agent is dynamically materialized inside its own orchestrated workflow, not when invoked standalone.
- Add a short authoring note to
.github/aw/create-agentic-workflow.md on PR-triggered comment workflows: recommend a sticky/updatable comment pattern (vs. one-comment-per-push) as the default for diff-reactive checks like accessibility or license scans, to reduce noise across synchronize events.
- Continue to reinforce the "omit network access when comparison target is in-repo" and "reject tools not needed for static analysis" judgment calls observed this run — these were correct but worth calling out explicitly in authoring guidance as a named pattern so they're reliably repeated.
Generated by 🎭 Agent Persona Explorer · claude · agent · 196.9 AIC · ⌖ 0.717 AIC · ⊞ 6.6K · ◷
Persona Overview
persona-evaluatorsub-agent)Key Findings
pull_request/schedule/issuestriggers withpaths:/types:filters matched to the task, and defaulted to minimal read-only permissions with writes routed exclusively through safe-outputs.playwrightfor a static-diff accessibility check (DS-A), and reasoned network access down to "omit entirely" when the compliance comparison target was in-repo (LC-B).add-commentwithtarget: "*"(not the defaulttriggering) to post on the parent milestone issue rather than the issue that fired the trigger — the agent caught this without being prompted.Top Patterns
pull_requestwith extension/dirpaths:scoping (DS-A, LC-A) for diff-reactive checks;schedule+workflow_dispatch(LC-B) for periodic policy audits;issues: [closed, reopened](PM-A) for state-change alerts.githubMCP ingh-proxymode with thedefaulttoolset was the near-universal choice; sandboxedbash/editleft at their defaults rather than redeclared. No scenario reached forplaywrightorcli-proxyunnecessarily.contents: read(+ narrowissues/pull-requests: readas needed), zero write scopes on the agent job, andsafe-outputsas the sole write path. Network access was denied by default in 3/4 cases and scoped to "GitHub API only" in the fourth.View High Quality Responses (Top 2)
skip-if-matchprevious-result strategy keyed on an open compliance-labeled issue to prevent duplicate findings piling up — a thoughtful addition not explicitly requested.View Areas for Improvement (Top 2)
Recommendations
.github/aw/create-agentic-workflow.md(or wherever the sub-agent's self-scoring instructions live) that direct ad hoc invocation should always self-score — the "parent evaluator will score" caveat should only apply when the agent is dynamically materialized inside its own orchestrated workflow, not when invoked standalone..github/aw/create-agentic-workflow.mdon PR-triggered comment workflows: recommend a sticky/updatable comment pattern (vs. one-comment-per-push) as the default for diff-reactive checks like accessibility or license scans, to reduce noise acrosssynchronizeevents.