AI error explanations: prompt builder and evaluation hook - #122
Draft
CatarinaGamboa wants to merge 1 commit into
Draft
CatarinaGamboa wants to merge 1 commit into
CatarinaGamboa wants to merge 1 commit into
Conversation
Adds a LiquidJava primer and a pure prompt builder that turns a LiquidJava diagnostic into a system/user prompt, with configurable context level (L0-L4), primer variant and explanation style (nudge or fix). activate() now returns the builder and the current diagnostics so the evaluation harness in liquidjava-dev can build prompts from real verifier output. Refs #113 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Part of #113. This PR adds the prompt side of the "Explain error" feature so it can be evaluated before we build the UI.
What's here
client/src/ai/primer.ts: LiquidJava background for the model, in three variants (none,syntax,full). Adapted from theliquidjava-mcpskill (Syntax section) and the docs page Understanding Refinement Errors (how to read "found/expected", counterexamples and superscript names likex⁷⁵).client/src/ai/prompt.ts:buildExplanationPrompt(diagnostic, readFile, { level, primer, style }). A pure function with novscodeimport, so the extension and the eval harness use the same code.nudge(next step to try) orfixactivate()returns{ getDiagnostics, buildExplanationPrompt }so the headless harness inliquidjava-dev(scripts/explain-capture.sh) can build prompts from real diagnostics.Evaluation
The six study exercises × 14 prompt variants × 3 models (GPT-6-Luna medium and high, GPT-6-Sol medium, called through Codex) were scored against the exercise ground truth by a separate judge model (GPT-6-Astra). Results and the report are in CatarinaGamboa/building-examples-liquidjava (
exercises/reports/ai-explanations/).Results (6 study exercises, 251 explanations)
"good" = correct diagnosis, useful step or fix, no hallucination, correct format (graded by a GPT-6-Astra judge against the sidecar ground truth).
cause_line(48) disagrees with its intended fix, which changes the reported line (60).Suggested default: L3 or L4, full primer, nudge style, with Luna medium for speed or Sol medium for the fewest hallucinations and the shortest answers. Which models Copilot actually offers through
vscode.lmstill needs checking.Full report: CatarinaGamboa/building-examples-liquidjava#2,
exercises/reports/ai-explanations/report.html. Open it locally; it has per-explanation comment boxes.Not in this PR (next steps)
vscode.lmwith the chosen model and prompt settings (needsengines.vscode≥ 1.90)🤖 Generated with Claude Code