Skip to content

fix(coderd/x/chatd): improve agent task and completion guidance - #29401

Merged
ibetitsmike merged 4 commits into
mainfrom
mike/agents-system-prompt
Sep 16, 2026
Merged

ibetitsmike merged 4 commits into
mainfrom
mike/agents-system-prompt

Conversation

@ibetitsmike

@ibetitsmike ibetitsmike commented Sep 16, 2026 •

Copy link
Copy Markdown
Collaborator

Why

The default agent prompt asked for as many tool calls as possible and gave conflicting instructions about when to ask for clarification. Replace those directives and the misleading completion examples with focused investigation, routine-choice autonomy, preservation of existing edits, dependency-aware tool use, and evidence-backed completion.

A follow-up comparison added mechanisms all three peers share and the first pass did not carry over: an action-safety section naming the operations that need authorization and banning destructive Git shortcuts, a preference for read_file, edit_files, and write_file over shell equivalents, no unrequested files, no introduced vulnerabilities, a findings-first review stance, and delegation restraint. User feedback that Coder Agents explores code less deliberately than peer harnesses added an investigation section, a search-first Explore sub-agent overlay, and briefing guidance for root chats.

Stack context

This is a single-PR change based on main. It updates the built-in prompt and its regression contracts without changing protected-branch confirmations, workspace-template selection, the Plan overlays, or the root-only subagent instruction boundary. The Explore overlay gains search discipline; its opening line that chatd_test.go asserts on is unchanged.

Verification and review record

Validated commit: 50147b36b72288f3734e307494b1b400b3190179.

  • Focused prompt, assembly, workspace-awareness, mode, and skill tests passed on every commit. The complete chatprompt package passed.
  • The repository pre-commit hook passed generation, formatting, lint, and the slim build on all four commits (94s, 91s, and 71s for the latest three). Local pre-push checks were not run, as requested.
  • Every new prompt contract assertion was run red against the previous commit's prompt via go test -overlay and green with the change: TestDefaultSystemPromptTaskDiscipline, TestDefaultSystemPromptContainsSubagentOrchestration, and TestExploreSubagentOverlayPromptSearchDiscipline.
  • Independent initial review: Xum task 98ebb8dc0a, read-only prompt/tool-boundary review of the diff committed as 7acc00a8f39. No remaining findings in that review.
  • Remote UAT and live model behavior were not evaluated. String and assembly tests do not demonstrate better model behavior or prompt-injection resistance. Human verification remains outstanding.
  • Codex review on 7acc00a8f39 found a missing explicit verification opt-out. Confirmed with a failing assertion, fixed in b582cc4098f, and verified green. The thread has an evidence reply and is resolved.
  • Codex re-review returned no major issues for b582cc4098. Commits 430f35b43e7 and 50147b36b72 are newer than that verdict; a fresh Codex review is requested below.

Latest focused test output:

ok  github.com/coder/coder/v2/coderd/x/chatd             6.132s
ok  github.com/coder/coder/v2/coderd/x/chatd/chatprompt  0.230s

Xum acted on behalf of @ibetitsmike.

@ibetitsmike ibetitsmike changed the title mike/agents system prompt fix(coderd/x/chatd): improve agent task and completion guidance Sep 16, 2026
@ibetitsmike

Copy link
Copy Markdown
Collaborator Author

@codex review

@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 16, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review 🔄 Running since 2026-09-16T20:00:12.089636Z 50147b3 Draft marked ready
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 7acc00a8f3

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread coderd/x/chatd/prompt.go Outdated
@ibetitsmike

Copy link
Copy Markdown
Collaborator Author

@codex review

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex Review: Didn't find any major issues. Keep them coming!

Reviewed commit: b582cc4098

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Adopt peer-harness mechanisms the prompt comparison surfaced but did not
carry over: an action-safety section naming the operations that need
authorization and banning destructive Git shortcuts, a preference for
read_file, edit_files, and write_file over shell equivalents, no
unrequested files, no introduced vulnerabilities, a findings-first
review stance, and delegation restraint in the root-only orchestration
block. Prompt contract tests cover each addition.
Coder Agents explores code less deliberately than peer harnesses. Add
an investigation section that asks for search, reading surrounding
code, a reference implementation, an end-to-end trace, and depth
scaled to the change. Give the Explore sub-agent overlay a search-first
process with location citations and negative findings, and tell root
chats to brief agents with known facts and to keep the understanding
they need themselves. Prompt contract tests cover each addition.
@ibetitsmike

Copy link
Copy Markdown
Collaborator Author

@codex review

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex Review: Didn't find any major issues. Breezy!

Reviewed commit: 50147b36b7

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

@ibetitsmike
ibetitsmike marked this pull request as ready for review September 16, 2026 20:00
@ibetitsmike
ibetitsmike merged commit eb800dc into main Sep 16, 2026
41 checks passed
@ibetitsmike
ibetitsmike deleted the mike/agents-system-prompt branch September 16, 2026 20:01
@github-actions github-actions Bot locked and limited conversation to collaborators Sep 16, 2026
Sign up for free to subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants