You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
[agentic-token-optimizer] Optimize Deep Report workflow: trim prompt, fix memory patch-size failure, consider task-mining sub-agent #54273
Deep Report (.github/workflows/deep-report.md → deep-report.lock.yml) — a scheduled (every 6h) intelligence-gathering agent that reviews discussions/issues/workflow logs and creates up to 7 quick-win issues.
Why selected: Highest AIC among candidates not recently optimized and not self-targeting ("Token" workflows excluded). No entry in optimization-log.json for this workflow in the last 14 days. Runs at ~153 AIC / 38K raw tokens on a single-run sample, with the workflow's own historical run cadence showing consistently long agent-turn durations (14–18 minutes) across the last 15 scheduled runs.
Analysis Period + Runs Audited
Window: last 7 days of scheduled runs (6-hour cadence).
Runs inspected: 15 recent runs via gh api actions/workflows/deep-report.lock.yml/runs, plus job-level detail on 6 of them.
Conclusion mix: 14 success, 1 failure (a downstream push_repo_memory job, not the main agent job).
Cost Profile
Metric
Value
Total AIC (sample run)
153.48
Avg AIC/run
153.48 (1-run window sample; historical cadence is every 6h)
Raw tokens
38,039
Avg turns/run
not reported (0)
Action minutes
15–18 min per run (agent job alone: ~10–14 min)
Cache efficiency
not exposed in current audit data
Ranked Recommendations
Trim and de-duplicate the Intelligence Collection prompt (Steps 0–4) — Est. savings: ~15–20 AIC/run
The prompt runs to 390 lines with 5 sequential "Step" sections (Step 0 memory check, Step 1 discussions, Step 2 workflow logs, Step 2.5 issues sub-agent, Step 2.7 code-quality mining, Step 3 cross-reference, Step 4 memory store) plus a 7-issue creation section with its own dedup-gate instructions repeated per-issue.
Several instructions overlap: dedup guidance appears in both Step 2.7 and the "Actionable Task Creation" dedup gate; discussion-mining logic in Step 2.7 largely restates Step 1's discussion-loading. Consolidating shared instructions (single "load discussions once, then branch into intelligence vs. task-mining passes") would cut redundant prompt tokens without changing behavior.
Step 1 instructs the agent to "ingest the filtered discussion data into AgentDB memory" and run semantic/hybrid searches over the full corpus every run, even though Step 0 already checks whether the last analysis was <20h ago. Confirm the AgentDB ingestion step actually skips/limits scope on the "recent-only" fast path — if it re-ingests the full corpus regardless, that's a repeated fixed cost that could be gated behind the incremental-analysis branch.
Root-cause and fix intermittent push_repo_memory patch-size failures — reliability, not directly AIC, but wastes a full run's setup/checkout cost
§32252000817 failed at push_repo_memory with: Patch diff size (13 KB, 12767 bytes) exceeds maximum allowed size (12 KB, 12288 bytes, configured limit: 10 KB with 20% overhead allowance).
This is unexpected: the workflow's frontmatter and lock file both configure max-patch-size: 51200 (50 KB) for this exact memory (tools.repo-memory.max-patch-size and the push_repo_memory safe-output config), yet the runtime enforcement in that run used a 10 KB limit. Worth confirming whether the configured 50 KB limit is actually being threaded through to the runtime patch-size check, since a stale/default 10 KB ceiling would explain sporadic failures on larger memory-update days (Step 4 writes 4 markdown files every run).
All other 14 sampled runs succeeded, so this looks like an edge case tied to unusually large single-run diffs rather than a systemic misconfiguration — but each failure means the memory update (last_analysis_timestamp, known_patterns, trend_data, flagged_items) silently doesn't persist, forcing the next run to redo more analysis from scratch.
Audit the 9 imported shared components for overlap — Est. savings: ~5–10 AIC/run
imports: pulls in shared/meta-analysis-base.md (with toolsets: [all]), the jqschema skill, shared/discussions-data-fetch.md, shared/mcp/agentdb.md, shared/weekly-issues-data-fetch.md, shared/reporting.md, shared/github-mcp-pagination-wrappers.md, shared/otlp.md, and shared/default-ai-credits-pricing.md. Several of these (reporting + github-mcp-pagination-wrappers, and otlp + default-ai-credits-pricing) are plausible candidates for consolidation into a single shared import if their guidance overlaps, reducing the fixed per-run prompt overhead that's paid on every 6-hour cycle regardless of how much new data is available.
The workflow already has one inline sub-agent (#### agent: issues-analyst, Step 2.5), so per the guardrail we only recommend a second sub-agent when it is a clearly separate, still-extractive task.
Candidate: task-miner sub-agent for Step 2.7 (Mine Discussions for Code Quality Tasks)
Task: Given the pre-fetched discussions JSON and the previously-processed-discussions list from repo-memory, extract candidate code-quality tasks (refactoring, testing gaps, docs, performance, security, tech debt, tooling) meeting the stated specific/actionable/valuable/scoped/independent criteria, and emit a structured list of {title, rationale, area, files_if_known}.
Why a smaller model fits: This is primarily extractive/classificatory work over structured JSON — filtering by 7-day window, checking against a fixed set of task categories, and formatting a candidate list. It does not require the strategic cross-referencing or final issue authoring that the main agent does in Steps 3–4.
Score: Independence 2/3 (needs the processed-discussions memory file as input, but doesn't depend on Steps 1–2 outputs) + Small-model adequacy 2/3 (extraction/classification) + Parallelism 1/2 (could run alongside Step 1's discussion intelligence pass) + Size 2/2 (processes potentially dozens of discussions) = 7/10 — strong candidate.
Exact invocation change: Replace the inline "### Step 2.7: Mine Discussions for Code Quality Tasks" instructions with:
Use the task-miner sub-agent to analyze /tmp/gh-aw/agent/discussions-data/discussions.json (excluding entries already in /tmp/gh-aw/repo-memory/default/deep-report/processed-discussions.json) and produce /tmp/gh-aw/repo-memory/default/deep-report/extracted-tasks.json with candidate code-quality tasks meeting the specific/actionable/valuable/scoped/independent criteria.
agent: task-miner
Extract code-quality quick-win tasks from the discussions JSON. For each unprocessed discussion (last 7 days), identify tasks that are specific, actionable, valuable, scoped to 1–3 days, and independent. Focus on refactoring, testing gaps, documentation, performance, security, technical debt, tooling. Exclude vague suggestions, feature requests, bug reports, and architectural decisions. Output a JSON array of {title, rationale, area} objects to /tmp/gh-aw/repo-memory/default/deep-report/extracted-tasks.json.
This keeps the main agent focused on cross-referencing (Step 3), dedup-gating, and final issue authoring, while offloading the mechanical extraction pass.
Caveats
This audit is based on a 1-run AIC/token sample for Deep Report in the 7-day pre-aggregated window (top-workflows.json); the schedule cadence (every 6h) means there were likely more historical runs, but detailed run-level financials for those weren't in the pre-aggregated dataset. Job-timing evidence across 15 runs (via gh api) was used to corroborate consistent duration, but per-run AIC could not be independently verified beyond the single sample.
The push_repo_memory limit discrepancy (10 KB enforced vs. 50 KB configured) is reported as observed evidence from one failed run; a full root-cause requires checking the safe-outputs runtime code path, which was out of scope for this audit.
No tool-removal recommendation is made: this audit did not have enough successful-run tool-call telemetry to confidently identify an unused tool among bash, edit, cli-proxy, repo-memory, or the many mcp__github__* read tools.
References:
§32368201153 — most recent successful run (AIC 153.48, sampled)
Target Workflow
Deep Report (
.github/workflows/deep-report.md→deep-report.lock.yml) — a scheduled (every 6h) intelligence-gathering agent that reviews discussions/issues/workflow logs and creates up to 7 quick-win issues.Why selected: Highest AIC among candidates not recently optimized and not self-targeting ("Token" workflows excluded). No entry in
optimization-log.jsonfor this workflow in the last 14 days. Runs at ~153 AIC / 38K raw tokens on a single-run sample, with the workflow's own historical run cadence showing consistently long agent-turn durations (14–18 minutes) across the last 15 scheduled runs.Analysis Period + Runs Audited
gh api actions/workflows/deep-report.lock.yml/runs, plus job-level detail on 6 of them.push_repo_memoryjob, not the mainagentjob).Cost Profile
Ranked Recommendations
Trim and de-duplicate the Intelligence Collection prompt (Steps 0–4) — Est. savings: ~15–20 AIC/run
importsare merged into the same context on top of this (see Add workflow: githubnext/agentics/weekly-research #4 below).Investigate AgentDB full-corpus semantic search cost — Est. savings: ~10–15 AIC/run
Root-cause and fix intermittent
push_repo_memorypatch-size failures — reliability, not directly AIC, but wastes a full run's setup/checkout costpush_repo_memorywith:Patch diff size (13 KB, 12767 bytes) exceeds maximum allowed size (12 KB, 12288 bytes, configured limit: 10 KB with 20% overhead allowance).max-patch-size: 51200(50 KB) for this exact memory (tools.repo-memory.max-patch-sizeand thepush_repo_memorysafe-output config), yet the runtime enforcement in that run used a 10 KB limit. Worth confirming whether the configured 50 KB limit is actually being threaded through to the runtime patch-size check, since a stale/default 10 KB ceiling would explain sporadic failures on larger memory-update days (Step 4 writes 4 markdown files every run).Audit the 9 imported shared components for overlap — Est. savings: ~5–10 AIC/run
imports:pulls inshared/meta-analysis-base.md(withtoolsets: [all]), thejqschemaskill,shared/discussions-data-fetch.md,shared/mcp/agentdb.md,shared/weekly-issues-data-fetch.md,shared/reporting.md,shared/github-mcp-pagination-wrappers.md,shared/otlp.md, andshared/default-ai-credits-pricing.md. Several of these (reporting + github-mcp-pagination-wrappers, and otlp + default-ai-credits-pricing) are plausible candidates for consolidation into a single shared import if their guidance overlaps, reducing the fixed per-run prompt overhead that's paid on every 6-hour cycle regardless of how much new data is available.Optional Structural Optimization: Inline Sub-Agent Candidate
The workflow already has one inline sub-agent (
#### agent: issues-analyst, Step 2.5), so per the guardrail we only recommend a second sub-agent when it is a clearly separate, still-extractive task.Candidate:
task-minersub-agent for Step 2.7 (Mine Discussions for Code Quality Tasks)Task: Given the pre-fetched discussions JSON and the previously-processed-discussions list from repo-memory, extract candidate code-quality tasks (refactoring, testing gaps, docs, performance, security, tech debt, tooling) meeting the stated specific/actionable/valuable/scoped/independent criteria, and emit a structured list of
{title, rationale, area, files_if_known}.Why a smaller model fits: This is primarily extractive/classificatory work over structured JSON — filtering by 7-day window, checking against a fixed set of task categories, and formatting a candidate list. It does not require the strategic cross-referencing or final issue authoring that the main agent does in Steps 3–4.
Score: Independence 2/3 (needs the processed-discussions memory file as input, but doesn't depend on Steps 1–2 outputs) + Small-model adequacy 2/3 (extraction/classification) + Parallelism 1/2 (could run alongside Step 1's discussion intelligence pass) + Size 2/2 (processes potentially dozens of discussions) = 7/10 — strong candidate.
Exact invocation change: Replace the inline "### Step 2.7: Mine Discussions for Code Quality Tasks" instructions with:
This keeps the main agent focused on cross-referencing (Step 3), dedup-gating, and final issue authoring, while offloading the mechanical extraction pass.
Caveats
top-workflows.json); the schedule cadence (every 6h) means there were likely more historical runs, but detailed run-level financials for those weren't in the pre-aggregated dataset. Job-timing evidence across 15 runs (viagh api) was used to corroborate consistent duration, but per-run AIC could not be independently verified beyond the single sample.push_repo_memorylimit discrepancy (10 KB enforced vs. 50 KB configured) is reported as observed evidence from one failed run; a full root-cause requires checking the safe-outputs runtime code path, which was out of scope for this audit.bash,edit,cli-proxy,repo-memory, or the manymcp__github__*read tools.References:
push_repo_memorypatch-size failure