Repository navigation
Normalize Claude session diagnostics and snapshot evidence - #67351
Conversation
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
|
✅ Design Decision Gate 🏗️ completed the design decision gate check. See the comment below for the result and any generated ADR draft. No ADR enforcement needed for PR #67351: the PR does not carry the 'implementation' label and has 0 new lines in default business logic directories (threshold: 100, no custom .design-gate.yml present).
|
|
🧠 Matt Pocock Skills Reviewer is reviewing this pull request using Matt Pocock's engineering skills... |
|
✅ Ponytail Reviewer completed successfully! Lean already. Ship.
|
|
✅ Test Quality Sentinel completed test quality analysis. Test Quality Sentinel skipped because pre-fetch PR data was unavailable: unable to fetch test file diff
|
|
✅ PR Code Quality Reviewer completed the code quality review.
|
There was a problem hiding this comment.
🟡 Changes recommended
Stream normalization still duplicates recognized blocks and leaves several mapped transport envelopes as separate native events.
1 open finding
What changed in this PR
Normalizes Claude diagnostics and streamed evidence into the unified session model while preserving native metadata.
Changes:
- Canonicalizes errors, permission denials, refusals, and structured content.
- Reconciles streamed events with finalized snapshots.
- Adds regression coverage and Claude documentation.
| File | Description |
|---|---|
.changeset/patch-claude-session-audit.md |
Records the patch release changes. |
actions/setup/js/agent_session_render.test.cjs |
Updates stream projection expectations. |
actions/setup/js/claude_session.cjs |
Implements Claude session normalization. |
actions/setup/js/claude_session.test.cjs |
Updates streaming snapshot coverage. |
actions/setup/js/claude_session_audit.test.cjs |
Adds diagnostic and evidence regressions. |
actions/setup/js/claude_session_normalization.test.cjs |
Updates denial normalization coverage. |
actions/setup/js/provider_refusal.test.cjs |
Tests attached refusal snapshots. |
actions/setup/js/unified_session.test.cjs |
Tests canonical provider errors. |
docs/src/content/docs/engines/claude.md |
Documents Claude session mapping. |
🧠 Review effort: Balanced
💡 Add a code-review agent skill for context-aware, tailored reviews. Learn more in the docs.
Comment MemoryPeek at saved memory (pr-code-quality-reviewer)Note This comment is managed by comment memory.Expand the saved memory block to view or edit the persistent context for this thread.
|
There was a problem hiding this comment.
Request changes
These changes still regress Claude trace fidelity in three places: successful camelCase tool results can become "unknown", and the only finalized SDK snapshot can be lost for both streamed blocks and late refusals.
Blocking themes
isError: falseis not normalized symmetrically with the new failure aliases, so some successful tool results no longer project as explicit successes.- First finalized SDK snapshots are not always retained on the canonical event once duplicate
claude.assistant_snapshotrecords are removed. - Late refusal reclassification drops the originating finalized snapshot, which defeats the evidence-preservation goal of this change.
🔎 Code quality review by PR Code Quality Reviewer · copilot · gpt54 · 49.3 AIC · ⌖ 7.87 AIC · ⊞ 19.9K
Comment /review to run again
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Why
Live Claude transcripts contain duplicate provider-error envelopes and permission-denial notices presented as assistant answers. This aligns those observations with the unified agent session contract while preserving native evidence and partial traces.
Changes
claude.*records. Preserve refusal policy metadata, observed partials, reasoning signatures, and recovered snapshot suffixes.Includes sanitized regression coverage and Claude-specific documentation. Shared implementation, types, projection, and specification files are unchanged; shared test edits only update Claude-specific expectations.
Evidence and limitations
Inspected existing success, API failure, documentation/subagent, and dynamic-workflow artifacts. The documentation run persisted 15 denial notices as assistant answers in both canonical and unified traces; the corrected replay retains 15 canonical denials once. The failure replay preserves all 22 error observations without duplicate native envelopes.
The newer two runs contain
agent-session.jsonlandaw_session.jsonl; older runs only support native-log replay. Streaming and structured-content regressions use synthetic SDK/API shapes, not live coverage claims. No workflow runs were triggered.Validation
make fmt-cjs && make lint-cjsPATH=/opt/homebrew/bin:$PATH make agent-report-progress: build, lint, schema freshness, and 237 impacted JavaScript tests passed; no Go changes.npm run typecheckandnpm run schema:session:checkfromactions/setup/js.