Skip to content

AI error explanations: prompt builder and evaluation hook - #122

Draft
CatarinaGamboa wants to merge 1 commit into
mainfrom
feat/113-ai-explanations
Draft

CatarinaGamboa wants to merge 1 commit into
mainfrom
feat/113-ai-explanations

Conversation

@CatarinaGamboa

@CatarinaGamboa CatarinaGamboa commented Sep 25, 2026 •

Copy link
Copy Markdown
Collaborator

Part of #113. This PR adds the prompt side of the "Explain error" feature so it can be evaluated before we build the UI.

What's here

  • client/src/ai/primer.ts: LiquidJava background for the model, in three variants (none, syntax, full). Adapted from the liquidjava-mcp skill (Syntax section) and the docs page Understanding Refinement Errors (how to read "found/expected", counterexamples and superscript names like x⁷⁵).
  • client/src/ai/prompt.ts: buildExplanationPrompt(diagnostic, readFile, { level, primer, style }). A pure function with no vscode import, so the extension and the eval harness use the same code.
    • Context levels: L0 message and hint · L1 ±15 lines · L2 whole file · L3 + the violated spec and a compact state machine · L4 + verifier internals (expected/found, counterexample, internal-name table)
    • Styles: nudge (next step to try) or fix
    • The answer format is fixed: What went wrong / Where it comes from (line N) / Try or Fix, at most 90 words
    • Converts the language server's 0-based positions to editor line numbers
  • activate() returns { getDiagnostics, buildExplanationPrompt } so the headless harness in liquidjava-dev (scripts/explain-capture.sh) can build prompts from real diagnostics.

Evaluation

The six study exercises × 14 prompt variants × 3 models (GPT-6-Luna medium and high, GPT-6-Sol medium, called through Codex) were scored against the exercise ground truth by a separate judge model (GPT-6-Astra). Results and the report are in CatarinaGamboa/building-examples-liquidjava (exercises/reports/ai-explanations/).

Results (6 study exercises, 251 explanations)

"good" = correct diagnosis, useful step or fix, no hallucination, correct format (graded by a GPT-6-Astra judge against the sidecar ground truth).

context level (full primer) good names cause line hallucinates median time
L0 message only 0% 0% 34% 13.6s
L1 ±15 lines 94% 94% 3% 6.7s
L2 whole file 83% 89% 8% 7.6s
L3 + spec and state machine 83% 86% 3% 7.1s
L4 + verifier internals 89% 83% 3% 6.9s
model (all variants) good hallucinates words median time
GPT-6-Luna medium 71% 12% 70 5.8s
GPT-6-Luna high 76% 8% 61 9.4s
GPT-6-Sol medium 76% 4% 54 7.7s
  • The error message alone is not enough. Every variant must include code.
  • L3/L4 are the safest settings. With the spec in the prompt, Luna medium made no hallucinations at L3 or L4, compared with 25% at L2. L1 looks best here only because every exercise's cause is 4–12 lines above the reported line.
  • The primer hardly matters when the spec is included (at L3: none 86%, full 83%, syntax 83% good).
  • Nudges never gave the fix away (0/126) and scored at least as well as full fixes.
  • E_A2_08 is an outlier (34% good) because its cause_line (48) disagrees with its intended fix, which changes the reported line (60).

Suggested default: L3 or L4, full primer, nudge style, with Luna medium for speed or Sol medium for the fewest hallucinations and the shortest answers. Which models Copilot actually offers through vscode.lm still needs checking.

Full report: CatarinaGamboa/building-examples-liquidjava#2, exercises/reports/ai-explanations/report.html. Open it locally; it has per-explanation comment boxes.

Not in this PR (next steps)

  • The "Explain error" action in the webview, calling Copilot through vscode.lm with the chosen model and prompt settings (needs engines.vscode ≥ 1.90)
  • An "AI-generated" label, and study logging of explain events (Opt-in study logging of extension interactions #121)

🤖 Generated with Claude Code

Adds a LiquidJava primer and a pure prompt builder that turns a LiquidJava
diagnostic into a system/user prompt, with configurable context level
(L0-L4), primer variant and explanation style (nudge or fix). activate()
now returns the builder and the current diagnostics so the evaluation
harness in liquidjava-dev can build prompts from real verifier output.

Refs #113

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant