Blog

Agent of the Day – September 28, 2026

Every repository accumulates issue clutter — related bugs filed separately, follow-ups scattered across weeks, nobody quite sure which ticket is the “real” one. Today’s Agent of the Day is the workflow built to fight that entropy: Issue Arborist, a daily codex-powered agent that reads the last 100 open issues in github/gh-aw and cultivates order out of them.

Issue Arborist’s job description is disarmingly simple: fetch open issues without a parent, spot the ones that belong together, and link them as sub-issues under a fresh tracking parent. No opinions about priority, no code changes — just topology. It runs once a day on a schedule, reads with github: mode: local in read-only issue scope, and writes exclusively through safe outputs (create-issue, link-sub-issue, create-discussion), so every action it takes is logged, reviewable, and reversible.

Its most recent run on September 28 is a good showcase of both the value and the honesty built into these agents. Working from 63.2k tokens and 11.6 AIC of compute, Issue Arborist scanned the open backlog and found two clusters worth grouping:

It also opened discussion #63933 in the audits category to summarize the pass, following the same reporting convention every gh-aw daily agent uses so a human skimming discussions gets the same story without opening 45 issues.

Here’s the part that makes this a genuinely interesting pick rather than a tidy success story: the run’s audit trail shows it attempted 34 link_sub_issue calls and most of them failed with Target is "triggering" but not running in issue context, skipping link_sub_issue — a context mismatch between the scheduled trigger and the sub-issue linking tool. The workflow’s own audit report flags this plainly as a workflow_failed critical finding, not something smoothed over. And yet the two parent issues and the discussion still landed cleanly, because gh-aw’s safe-outputs model treats each output independently — a broken tool call doesn’t roll back the ones that already succeeded. Compare that to the previous day’s run, which completed with no findings and no issues to group, a perfectly quiet, uneventful pass.

That contrast is the whole point of running the same agent daily: some days there’s nothing to say, and some days there’s a real cluster to surface — and when a downstream tool call breaks, gh-aw’s audit tooling makes sure the failure is visible instead of buried in a green checkmark. Issue Arborist doesn’t pretend the run was flawless. It reports the win and the wart in the same breath, which is exactly what you want from automation touching your issue tracker unsupervised.

Want to see the compiled workflow, the safe-outputs config, or borrow the pattern for your own repo? Start here: github/gh-aw.

Weekly Update – September 28, 2026

It was a big week for github/gh-aw: three releases shipped, a legacy sandbox runtime was retired, and the firewall got noticeably stricter about what agents can reach. Here’s the rundown.

v0.89.22 landed on September 27th as the week’s headline release, building on v0.89.21 and v0.89.20 from earlier in the week.

  • Docker sbx and gVisor sandbox runtimes removed (#63034) — isolated agent execution now runs exclusively on Cloud Hypervisor. If your workflow frontmatter still sets sandbox.agent.runtime: docker-sbx or gvisor, switch it to cloud-hypervisor before you next recompile.
  • Hosted-web domain policies in frontmatter (#63212) — the new network.hosted-web key lets you allow or block domains that provider-hosted Claude and Codex web tools reach outside AWF’s network boundary.
  • Firewall enforcement got sharper — Copilot and web tools now get enforced firewall compatibility checks (#63474), scoped precisely to the web tools a workflow actually enables (#63632), while false blocked-domain warnings for otherwise-successful requests were squashed (#63246).
  • Grouped audit findings mode (#63032) — gh aw audit --group now aggregates findings by run and code with occurrence counts, in pretty, markdown, or JSON output.
  • More audit visibility — an opt-out for automatic audit baseline downloads (#63012) and MCP payload size reporting in audit (#63687) give workflow authors finer control over what audit runs measure and fetch.
  • Large Copilot prompts streamed through stdin (#62766, thanks @davidslater!) — prompts over 100 KiB are now delivered via stdin instead of being truncated, preserving full context for big agentic workflows.
  • Repo-memory backend for daily AIC guardrail (#62958) — memory persistence now supports repository-backed storage for cost accounting.
  • Fixed dangling safe-outputs-app-token references when safe outputs are staged (#63490).
  • Fixed repo memory retry head refresh authentication (#63497) and prevented cache-memory validation marker EACCES failures (#63498).
  • Pinned pull_request activation checkout to the base SHA for improved supply-chain safety (#63499).
  • Fixed slash-command PR prefetch caching and error handling (#63678).

Meet ci-coach, the CI Optimization Coach that runs daily to hunt down slow, imbalanced, or wasteful spots in gh-aw’s CI pipeline and proposes fixes as pull requests.

This week ci-coach had its hands full with cgo.yml’s notoriously lopsided unit-test shard. On September 24th it noticed the alphabetic A-C shard was taking ~106 seconds — more than double the other four shards’ 40-48 seconds — and traced nearly 20% of that gap to a single test full of pointless time.Sleep-backed cooldowns. It opened a fix (#63187) that reconfigured the mocked rate limit so the test skips its legacy 500ms sleep loop entirely. The very next day, still not satisfied, it went after the same shard again — this time proposing a full matrix rebalance that carves out a dedicated “Linters” shard, projected to cut the A-C shard’s time from ~93s down to ~29s.

Two days, two PRs, one stubbornly slow test shard — ci-coach is basically the CI equivalent of someone who keeps rearranging the furniture until the room finally feels right.

Usage tip: Run this class of workflow on a schedule against your CI config files — it’s great at spotting shard imbalances and dead-weight sleeps that are easy to miss by eye but add up fast across hundreds of runs.

→ View the workflow on GitHub

Update to v0.89.22 today, double-check your sandbox runtime settings if you were on docker-sbx or gvisor, and give the new gh aw audit --group mode a spin. As always, questions and contributions are welcome in github/gh-aw.

Agent of the Day – September 25, 2026

Most security scanners are optimized for one thing: finding something to report. That incentive quietly rewards noise — every string concatenation near a shell call becomes a “finding,” every developer inherits a triage queue full of things that were never actually exploitable. Today’s Agent of the Day takes the opposite approach. Meet the Daily VulnHunter Scan, a Claude Code workflow that hunts for injection-class bugs in gh-aw’s own source and treats “no findings” as a perfectly good outcome.

VulnHunter runs Capital One’s open-sourced vulnhunt methodology inside a sandboxed bundle of the repository — a snapshot of pkg/cli, pkg/parser, pkg/workflow, and actions/setup/js, plus a reference guide (phase2_class_inj.md) covering the usual dangerous-sink suspects: SQL injection, command injection, path traversal, SSRF, XXE, unrestricted file upload, XSS, open redirect, LDAP injection, and server-side template injection. Each day it works through a ranked candidate list, checking whether attacker- or LLM-controlled data can reach one of those sinks unsanitized.

On September 24 (run 35961274043), the agent reviewed 25 of 41 candidates and flagged exactly one soft lead: an unsanitized branch argument passed into git checkout -b inside actions/setup/js/apply_samples.cjs, which on paper looks like a textbook CWE-88 flag-injection risk. But instead of filing an issue on pattern-match alone, VulnHunter ran it through a reachability gate — tracing where that branch value actually originates. It found the input comes from GH_AW_SAMPLES, a compile-time deterministic-replay fixture produced by gh aw compile --use-samples and authored by the workflow developer, never by live agent output or an external actor at runtime. No trust boundary is crossed, so the candidate was falsified and no issue was created — a correct “no finding” rather than a false alarm dressed up as a vulnerability report.

The very next run, on September 25 (run 36099855973), the agent went further: all 40 ranked candidates reviewed, zero surviving even initial construction, so nothing needed the deeper Phase 2b falsification pass at all. Its reasoning notes are worth reading as a mini security audit in themselves — every os/exec call across the codebase uses array-based arguments rather than shell-string concatenation, path-construction sinks are guarded by explicit validators like isSafeGitRevisionArg and fileutil.ValidatePathWithinBase, Docker-based scanners validate image references before building argv, and values interpolated into generated YAML are consistently escaped or sourced from trusted workflow-author configuration.

That run also came with a candid self-assessment from gh aw’s own audit tooling: 50 turns and 1.94M tokens for a single-agent scan, flagged as a “resource heavy” profile for its task domain, with roughly half the turns doing data-gathering that could in principle move to deterministic pre-agent steps. Comparing it against a matched September 22 baseline showed turn count climbing from 44 to 50 while blocked network requests dropped from 3 to 0 — a small but visible behavioral drift that the audit surfaced automatically, without anyone having to eyeball two log files side by side.

gh-aw workflow activity chart

What makes VulnHunter interesting isn’t that it never finds anything — it’s that “nothing to report” is a real, load-bearing output, backed by explicit falsification logic instead of a shrug. In a codebase that runs LLM agents against live repositories every day, a scanner that can tell the difference between “this pattern looks scary” and “this pattern is actually reachable by an attacker” is worth more than one that just counts regex matches.

Curious how workflows like this get built, scoped, and kept honest run after run? Check out github/gh-aw.

Agent of the Day – September 24, 2026

Every maintainer knows the feeling: you open your pull request list on a Monday morning and it’s wall-to-wall Dependabot noise — a dozen individual bumps for docker/login-action, @vitest/ui, prettier, github.com/cli/go-gh, each waiting for a separate review pass. Today’s Agent of the Day exists purely to make that Monday less painful. Say hello to the Dependabot Burner.

Dependabot Burner runs on a weekly schedule (with manual dispatch and a /dependabot-burner slash command as backups) and does one job: collect the grouped Dependabot PRs targeting generated workflow manifests, trace them back to their source workflow markdown, apply the equivalent update there, recompile, and open a single replacement pull request. Instead of a reviewer wading through ten mechanical diffs, they get one coherent change with a clear story.

The workflow was born from PR #40396, which added centralized grouping and retry-aware remediation so a single run can absorb an entire batch of dependency noise rather than nibbling at it one PR at a time. It’s a gpt-5.4-mini Copilot CLI run with read-only GitHub access by default — it only writes through the create_pull_request safe output, keeping its blast radius small even though its job is to touch a lot of files.

Recent run history is a good demonstration of why “Agent of the Day” doesn’t require a flawless track record. Digging into run #33845039019 from September 4, the agent completed successfully in 27 turns and ~10 minutes, walking through candidate PRs, grouping the applicable ones, and firing off its replacement pull request. A week later on September 11, and again on September 18, the same workflow hit a driver_exit failure and stopped after essentially zero turns — and that’s exactly the point. gh aw’s audit tooling flagged it immediately: workflow conclusion failure, a threat_detection_job_failed finding noting the security job never started, and a clear recommendation to check the raw error logs before assuming anything shipped. No silent PR, no partial state — just a loud, attributable stop.

That loud failure mode matters more than a shiny green checkmark. A dependency-remediation agent that silently half-applies changes is far more dangerous than one that occasionally refuses to proceed. Comparing the September 4 success against the September 18 failure through gh aw’s baseline-cohort matching also surfaced a useful signal for the maintainers: the classification engine flagged the failed run as “risky” purely from turn-count divergence (27 turns → 0 turns) before even reading the error text, which is the kind of anomaly detection that turns a scheduled job from a black box into something debuggable.

There’s a second lesson baked into the successful run too. The audit noted the agent used a “resource heavy” execution profile for a general-automation task and that roughly half its turns were data-gathering that could, in principle, move to deterministic pre-agent steps — a nudge toward the DeterministicOps pattern that shows up across gh-aw’s more mature workflows. Grouping ten PRs into one is already a win for reviewer time; trimming the agent’s own overhead is the next iteration.

gh-aw workflow activity chart

Dependabot Burner is a small reminder that “boring” agents — the ones that just tidy up dependency churn — are some of the most operationally valuable in a large repository. It doesn’t need creativity. It needs discipline, a clean failure mode, and a habit of turning ten PRs into one.

Want to see how workflows like this are built, audited, and kept honest? Check out github/gh-aw.

Agent of the Day – September 23, 2026

Some agents chase bugs. Some agents chase flaky tests. Today’s Agent of the Day chases something much quieter: the extra words your documentation didn’t need. Meet Daily Caveman Optimizer, the workflow that reads gh-aw’s own instruction files one at a time and asks a single, blunt question — “why use many token when few do trick?”

Agent of the Day: Daily Caveman Optimizer

Section titled “Agent of the Day: Daily Caveman Optimizer ”

The name is a deliberate joke, borrowed from the “caveman optimization” principle: strip prose down to its essentials without losing a single fact. The workflow runs daily on a schedule, round-robins through the .github/aw and .github/agents instruction directories, and only opens a pull request when it finds real redundancy worth removing.

Its most recent successful run landed on file 31 of a 71-file queue: .github/aw/loop.md, the shared spec for gh-aw’s Autoloop, Goal, and Crane automation patterns. The agent noticed that a ## Shared architecture section near the top of the file was restating — almost line for line — six concepts that a ## Pattern inventory section immediately below already covered in more depth: the single-item scheduler, canonical branch and single-PR conventions, ratcheting acceptance, durable repo-memory state, the human control-plane issue, and pause semantics.

Rather than just deleting the redundant section, the agent did the more careful thing: it checked whether any unique details lived only in the shorter, duplicate version, and migrated those into the pattern entries that were keeping them. That meant preserving branch-naming templates like autoloop/<program> and goal/<issue>-<slug>, the repo-memory branch names (memory/autoloop, memory/goal, memory/crane), the status-comment sentinel format, and the “discard the change but still record the run” rule — nothing structural was lost, only the repetition.

The result, opened as PR #62470, was refreshingly small: 8 lines added, 32 removed, net 17% shorter, one file touched. The PR body even included a table of files it reviewed but chose not to touch — intent.md, jobs.md, linter-workflows.md, llms.md — each with a one-line reason (“already imperative bullets with no filler,” “mostly code blocks and a port/credential table”), which is a good sign the agent isn’t optimizing for PR volume, just genuine wins. Maintainer @pelikhan merged it about 27 minutes after it was opened.

This run was also part of a live A/B experiment comparing claude-sonnet-5 against claude-haiku-4.5 on this exact task — testing whether the cheaper model produces equivalent documentation trims at lower token cost. This run drew the Haiku variant and still shipped a clean, mergeable PR, one data point toward the workflow’s hypothesis that a lighter model can hold its own on this kind of surgical editing.

Not every run finds something to fix — the agent’s if-no-changes: "ignore" setting means it stays quiet when a file is already tight, and one recent run did close without a PR, exactly as designed. Zero-signal days are a feature here, not a failure: pruning the last-lines-standing kind of redundancy is 17% of the time, and that’s fine, because loud PRs full of arguable rewrites would be worse than a quiet queue.

It’s a small, unglamorous job — nobody puts “reduced loop.md by 24 lines” on a highlight reel. But instruction files are the load-bearing walls of an agentic-workflow repo like gh-aw: every workflow, every linter, every sub-agent reads them before doing anything else. An agent that keeps them lean, accurate, and duplicate-free is doing maintenance work that pays compounding interest every time a human — or another agent — has to parse those docs under a deadline.

Want to see how a scheduled cleanup agent turns “read this file” into a mergeable pull request? Explore the source at github/gh-aw and see what your own instruction files might be hiding.