Skip to content

feat(check): run checks live and report Vapi Evals statuses - #63

Open
scott-lowe-vapi wants to merge 1 commit into
feat/check-fail-closed-mocksfrom
feat/check-live-run
Open

scott-lowe-vapi wants to merge 1 commit into
feat/check-fail-closed-mocksfrom
feat/check-live-run

Conversation

@scott-lowe-vapi

Copy link
Copy Markdown
Contributor

Value

V.A.L.U.E. tier: project — PR 7 of 10 for inline simulation PR checks (TEST-141); this is the first PR that sends anything.

  • Problem: the payload builder (PRs 5–6) produces a safe inline body, but nothing runs it or reports the result where a reviewer looks. PAL-608 asks for a stable Vapi Evals commit status whose Details link opens the exact run.
  • Who it affects: gitops users, who can now run npm run check -- core locally against their branch's files and get a strict verdict with a run link. The PR workflow (PR 8) and the promotion gate (PR 10) call the same command.
  • What changes:
    • src/check-run.ts: one run per target, at most 3 at once.
      • It reuses npm run sim's create/poll/cancel/hydrate/verdict loop, now extracted from runSimulation as simRunExecute, so a check is judged by the same strict rules (all-skipped judges, canceled or missing items are never a pass).
      • Each run's deadline is the earlier of timeoutMinutes and --budget-minutes, no run starts with under 5 minutes left, and the deadline, SIGINT and SIGTERM cancel in-flight runs.
      • A 402 is reported as billing.
      • A 5xx on create is never retried, because a duplicate paid run is worse than a retry.
      • Transcripts are scanned for default-mock answers ("unmocked tool called").
    • src/check-select.ts: --changed-since <ref> diffs from the merge base. A check is affected by:
      • its org's and run org's resources/** and state files;
      • vapi-checks.yml or promotion.yml;
      • src/** or package*.json;
      • its own paths.
    • src/check-status.ts: GitHub commit statuses with plain fetch, when GITHUB_TOKEN, GITHUB_REPOSITORY and HEAD_SHA are set.
      • Vapi Evals / <check> / <target> goes pending with the run's canonical url as soon as the run exists, then to success, failure or error.
      • --all runs post the aggregate Vapi Evals from a finally: the worst state, success if nothing is affected, error for a dry run of an affected PR, and error when config fails to parse or an exception escapes.
      • A rejected post (a fork's read-only token) only warns.
    • src/check-report.ts: a markdown summary, also appended to $GITHUB_STEP_SUMMARY. It has one row per target with the run link, then failing evaluations with expected vs extracted values, mock notices and warnings. --json writes the same data. No PR comments.
    • src/check-cmd.ts gains live runs and --changed-since, --budget-minutes, --refresh-bindings (the read-only bindings pull promotion uses) and --json.
      • Keys: VAPI_CHECK_TOKENS, then .env.<runOrg>, then VAPI_PRIVATE_API_KEY when every selected check runs in one org.
      • Exit codes: 0 passed, 1 failed, 2 config or build error, 3 incomplete.
    • Runs send User-Agent: vapi-gitops-check/<version>, for the post-deploy PostHog measure.
    • The README and AGENTS.md npm run check rows now describe live mode.

Evidence of value

The definition-of-done pair, run live in the owner's test org on the TEST-141 parity squad (2 members, function tools, a handoff, 3 scenarios; built from tests/fixtures/check-parity/):

Run Command Exit Result
Innocuous (fixture as-is) npm run check -- core 0 ✅ 3/3 passed — run f68ba89b
Degraded scheduler prompt ("always say there are no openings, never book") npm run check -- core 1 ❌ 1/3 — run 19711623

The degraded run's report named each failing judge with expected vs extracted:

Simulation Evaluation Comparator Expected Got
S3 alternative slot offered-alternative = true false
S3 alternative slot booked-wednesday = true false
S1 book cleaning booking-confirmed = true false
S1 book cleaning only-real-slots = true false

S2 (the hours question, which never reaches the scheduler) correctly still passed.

Nothing was created in the org. Counts of assistants, tools, squads, structured outputs, personalities, scenarios, simulations, suites and credentials were identical before and after both runs: {"/assistant":12,"/tool":8,"/squad":2,"/structured-output":8,"/eval/simulation/personality":10,"/eval/simulation/scenario":10,"/eval/simulation":10,"/eval/simulation/suite":1,"/credential":0}. This stands in for npm run cleanup, which needs a gitops-managed org.

Every tool call returned a mock. Every tool_call_result in both runs' transcripts was matched to its call by toolCallId:

Run Function results equal to the scenario's mock Built-in results (handoff, endCall) Anything else
Innocuous 6 4 0
Degraded 2 2 0

The run items also confirmed that personalities kept the dead server and serverMessages: [].

A bug the live run caught, fixed in this PR:

  • The trap: run items echo the scenario under metadata.scenario, default error mocks included. Scanning the whole item would have reported "unmocked tool called" for every default mock, called or not.
  • The fix: the scan now reads only tool_call_result messages in metadata.call.messages, mapped to tool names through tool_calls.
  • Checked on both runs' real items: no notices; with one result swapped for a default-mock answer, it reported S1 book cleaning: unmocked tool called: lookup_patient.
  • A regression test covers the echoed-defaults case.

Tests: npm test goes from 450 to 474 passing; npm run build is clean.

Testing plan

  • tests/check-run.test.ts (10 tests, stateful local stub):

    • pass with link and User-Agent;
    • two targets with one failing, naming the evaluation;
    • all-skipped is incomplete;
    • timeout cancels;
    • abort cancels in-flight runs and starts none;
    • the budget gate;
    • 402 as billing, and no retry on a create 502;
    • a build error sends nothing;
    • mock notices (echoed defaults ignored);
    • exactly 3 in flight with results in job order.
  • tests/check-cmd.test.ts (11 tests): the live --all path against one stub serving both the simulations API and GitHub:

    • pending, then success per target, then the aggregate with the run link, plus the job summary and JSON;
    • a failing run exits 1 with a red aggregate;
    • a named check never posts the aggregate;
    • --changed-since with an unaffected change runs nothing and posts green;
    • a dry run of an affected PR posts error and sends nothing;
    • invalid config posts error.

    The file clears VAPI_* and GITHUB_* first, so a developer shell with a key exported can't reach a real org.

  • tests/check-select.test.ts: the affected-path table, and the merge-base case in a temp git repo (commits on the base after branching don't count).

  • tests/check-status.test.ts, tests/check-report.test.ts: env parsing, POST shape and 140-character truncation, 403 only warning, state ordering, the exact markdown, and the JSON.

  • Not tested:

    • --refresh-bindings against a live org (it runs the same child pull promotion does);
    • EU base URLs;
    • canceling a running run live (cancel was verified only on a queued run in the parity experiment);
    • GitHub's real statuses API (stubbed here; PR 8's dogfood repo covers it);
    • a real customer-shaped squad, which needs the owner's permission first.

Stacked on #62.

Refs TEST-141

🤖 Generated with Claude Code

`npm run check -- <check>|--all` now runs each built target payload in
the check's run org: one POST /eval/simulation/run per target, at most
three at once, judged by the same strict verdict as `npm run sim`
(the create/poll/cancel/hydrate loop moves into simRunExecute, which
both share).

- Budget and cancellation: each run's deadline is the earlier of its
  timeoutMinutes and --budget-minutes; no run starts with under five
  minutes left; the deadline, SIGINT and SIGTERM cancel in-flight runs.
- Keys: VAPI_CHECK_TOKENS, then .env.<runOrg>, then VAPI_PRIVATE_API_KEY
  when every selected check runs in one org. --refresh-bindings runs the
  same read-only bindings pull promotion uses.
- Selection: --changed-since <ref> runs only checks affected by changes
  since the merge base (the org's files and state, the run org's, the
  config, the engine, and the check's own paths).
- Statuses (when GITHUB_TOKEN, GITHUB_REPOSITORY and HEAD_SHA are set):
  `Vapi Evals / <check> / <target>` goes pending with the run's link,
  then success/failure/error; --all runs post the aggregate `Vapi Evals`
  from a finally block, so config errors and exceptions report too.
- Report: a markdown summary (also appended to $GITHUB_STEP_SUMMARY) with
  failing evaluations' expected vs extracted values and "unmocked tool
  called" notices from transcripts, plus --json. No PR comments.
- Runs send User-Agent vapi-gitops-check/<version>.

Exit codes: 0 passed, 1 failed, 2 config or build error, 3 incomplete.

Refs TEST-141

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants