fix(coderd/x/chatd): improve agent task and completion guidance - #29401
Conversation
|
@codex review |
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 7acc00a8f3
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
|
@codex review |
|
Codex Review: Didn't find any major issues. Keep them coming! Reviewed commit: ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
If Codex has suggestions, it will comment; otherwise it will react with 👍. Codex can also answer questions or update the PR. Try commenting "@codex address that feedback". |
Adopt peer-harness mechanisms the prompt comparison surfaced but did not carry over: an action-safety section naming the operations that need authorization and banning destructive Git shortcuts, a preference for read_file, edit_files, and write_file over shell equivalents, no unrequested files, no introduced vulnerabilities, a findings-first review stance, and delegation restraint in the root-only orchestration block. Prompt contract tests cover each addition.
Coder Agents explores code less deliberately than peer harnesses. Add an investigation section that asks for search, reading surrounding code, a reference implementation, an end-to-end trace, and depth scaled to the change. Give the Explore sub-agent overlay a search-first process with location citations and negative findings, and tell root chats to brief agents with known facts and to keep the understanding they need themselves. Prompt contract tests cover each addition.
|
@codex review |
|
Codex Review: Didn't find any major issues. Breezy! Reviewed commit: ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
If Codex has suggestions, it will comment; otherwise it will react with 👍. Codex can also answer questions or update the PR. Try commenting "@codex address that feedback". |
Why
The default agent prompt asked for as many tool calls as possible and gave conflicting instructions about when to ask for clarification. Replace those directives and the misleading completion examples with focused investigation, routine-choice autonomy, preservation of existing edits, dependency-aware tool use, and evidence-backed completion.
A follow-up comparison added mechanisms all three peers share and the first pass did not carry over: an action-safety section naming the operations that need authorization and banning destructive Git shortcuts, a preference for
read_file,edit_files, andwrite_fileover shell equivalents, no unrequested files, no introduced vulnerabilities, a findings-first review stance, and delegation restraint. User feedback that Coder Agents explores code less deliberately than peer harnesses added an investigation section, a search-first Explore sub-agent overlay, and briefing guidance for root chats.Stack context
This is a single-PR change based on
main. It updates the built-in prompt and its regression contracts without changing protected-branch confirmations, workspace-template selection, the Plan overlays, or the root-only subagent instruction boundary. The Explore overlay gains search discipline; its opening line thatchatd_test.goasserts on is unchanged.Verification and review record
Validated commit:
50147b36b72288f3734e307494b1b400b3190179.chatpromptpackage passed.go test -overlayand green with the change:TestDefaultSystemPromptTaskDiscipline,TestDefaultSystemPromptContainsSubagentOrchestration, andTestExploreSubagentOverlayPromptSearchDiscipline.98ebb8dc0a, read-only prompt/tool-boundary review of the diff committed as7acc00a8f39. No remaining findings in that review.7acc00a8f39found a missing explicit verification opt-out. Confirmed with a failing assertion, fixed inb582cc4098f, and verified green. The thread has an evidence reply and is resolved.b582cc4098. Commits430f35b43e7and50147b36b72are newer than that verdict; a fresh Codex review is requested below.Latest focused test output: