feat(check): run checks live and report Vapi Evals statuses - #63
Open
scott-lowe-vapi wants to merge 1 commit into
Open
scott-lowe-vapi wants to merge 1 commit into
scott-lowe-vapi wants to merge 1 commit into
Conversation
`npm run check -- <check>|--all` now runs each built target payload in the check's run org: one POST /eval/simulation/run per target, at most three at once, judged by the same strict verdict as `npm run sim` (the create/poll/cancel/hydrate loop moves into simRunExecute, which both share). - Budget and cancellation: each run's deadline is the earlier of its timeoutMinutes and --budget-minutes; no run starts with under five minutes left; the deadline, SIGINT and SIGTERM cancel in-flight runs. - Keys: VAPI_CHECK_TOKENS, then .env.<runOrg>, then VAPI_PRIVATE_API_KEY when every selected check runs in one org. --refresh-bindings runs the same read-only bindings pull promotion uses. - Selection: --changed-since <ref> runs only checks affected by changes since the merge base (the org's files and state, the run org's, the config, the engine, and the check's own paths). - Statuses (when GITHUB_TOKEN, GITHUB_REPOSITORY and HEAD_SHA are set): `Vapi Evals / <check> / <target>` goes pending with the run's link, then success/failure/error; --all runs post the aggregate `Vapi Evals` from a finally block, so config errors and exceptions report too. - Report: a markdown summary (also appended to $GITHUB_STEP_SUMMARY) with failing evaluations' expected vs extracted values and "unmocked tool called" notices from transcripts, plus --json. No PR comments. - Runs send User-Agent vapi-gitops-check/<version>. Exit codes: 0 passed, 1 failed, 2 config or build error, 3 incomplete. Refs TEST-141 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This was referenced Oct 1, 2026
Contributor
Author
|
Warning This pull request is not mergeable via GitHub because a downstack PR is open. Once all requirements are satisfied, merge this PR as a stack on Graphite.
This stack of pull requests is managed by Graphite. Learn more about stacking. |
This was referenced Oct 1, 2026
scott-lowe-vapi
marked this pull request as ready for review
October 1, 2026 23:52
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.

Value
V.A.L.U.E. tier: project — PR 7 of 10 for inline simulation PR checks (TEST-141); this is the first PR that sends anything.
Vapi Evalscommit status whose Details link opens the exact run.npm run check -- corelocally against their branch's files and get a strict verdict with a run link. The PR workflow (PR 8) and the promotion gate (PR 10) call the same command.src/check-run.ts: one run per target, at most 3 at once.npm run sim's create/poll/cancel/hydrate/verdict loop, now extracted fromrunSimulationassimRunExecute, so a check is judged by the same strict rules (all-skipped judges, canceled or missing items are never a pass).timeoutMinutesand--budget-minutes, no run starts with under 5 minutes left, and the deadline, SIGINT and SIGTERM cancel in-flight runs.src/check-select.ts:--changed-since <ref>diffs from the merge base. A check is affected by:resources/**and state files;vapi-checks.ymlorpromotion.yml;src/**orpackage*.json;paths.src/check-status.ts: GitHub commit statuses with plainfetch, whenGITHUB_TOKEN,GITHUB_REPOSITORYandHEAD_SHAare set.Vapi Evals / <check> / <target>goes pending with the run's canonicalurlas soon as the run exists, then to success, failure or error.--allruns post the aggregateVapi Evalsfrom afinally: the worst state,successif nothing is affected,errorfor a dry run of an affected PR, anderrorwhen config fails to parse or an exception escapes.src/check-report.ts: a markdown summary, also appended to$GITHUB_STEP_SUMMARY. It has one row per target with the run link, then failing evaluations with expected vs extracted values, mock notices and warnings.--jsonwrites the same data. No PR comments.src/check-cmd.tsgains live runs and--changed-since,--budget-minutes,--refresh-bindings(the read-only bindings pull promotion uses) and--json.VAPI_CHECK_TOKENS, then.env.<runOrg>, thenVAPI_PRIVATE_API_KEYwhen every selected check runs in one org.User-Agent: vapi-gitops-check/<version>, for the post-deploy PostHog measure.npm run checkrows now describe live mode.Evidence of value
The definition-of-done pair, run live in the owner's test org on the TEST-141 parity squad (2 members, function tools, a handoff, 3 scenarios; built from
tests/fixtures/check-parity/):npm run check -- corenpm run check -- coreThe degraded run's report named each failing judge with expected vs extracted:
S2 (the hours question, which never reaches the scheduler) correctly still passed.
Nothing was created in the org. Counts of assistants, tools, squads, structured outputs, personalities, scenarios, simulations, suites and credentials were identical before and after both runs:
{"/assistant":12,"/tool":8,"/squad":2,"/structured-output":8,"/eval/simulation/personality":10,"/eval/simulation/scenario":10,"/eval/simulation":10,"/eval/simulation/suite":1,"/credential":0}. This stands in fornpm run cleanup, which needs a gitops-managed org.Every tool call returned a mock. Every
tool_call_resultin both runs' transcripts was matched to its call bytoolCallId:The run items also confirmed that personalities kept the dead server and
serverMessages: [].A bug the live run caught, fixed in this PR:
metadata.scenario, default error mocks included. Scanning the whole item would have reported "unmocked tool called" for every default mock, called or not.tool_call_resultmessages inmetadata.call.messages, mapped to tool names throughtool_calls.S1 book cleaning: unmocked tool called: lookup_patient.Tests:
npm testgoes from 450 to 474 passing;npm run buildis clean.Testing plan
tests/check-run.test.ts(10 tests, stateful local stub):tests/check-cmd.test.ts(11 tests): the live--allpath against one stub serving both the simulations API and GitHub:--changed-sincewith an unaffected change runs nothing and posts green;errorand sends nothing;error.The file clears
VAPI_*andGITHUB_*first, so a developer shell with a key exported can't reach a real org.tests/check-select.test.ts: the affected-path table, and the merge-base case in a temp git repo (commits on the base after branching don't count).tests/check-status.test.ts,tests/check-report.test.ts: env parsing, POST shape and 140-character truncation, 403 only warning, state ordering, the exact markdown, and the JSON.Not tested:
--refresh-bindingsagainst a live org (it runs the same child pull promotion does);runningrun live (cancel was verified only on a queued run in the parity experiment);Stacked on #62.
Refs TEST-141
🤖 Generated with Claude Code