Minimal localhost dashboard for coder_eval runs. Next.js App Router, reads runs from a local runs directory in server components — no database, no persistent backend.
cd coder_eval/evalboard
pnpm install
EVALBOARD_LOCAL_RUNS_DIR=/path/to/runs pnpm dev
# open http://localhost:3030Point EVALBOARD_LOCAL_RUNS_DIR at any coder_eval runs directory; the listing
comes from the filesystem. pnpm dev:local is a shortcut for coder_eval's own
runs/ (relative to evalboard/). Only directories containing a run.json
show up in the index — empty shells and the latest symlink are filtered out.
/— the 20 most recent runs, one row each, clickable. Includes a daily success-rate chart and tag rails for filtering./trends— per-task pass rate and avg duration/cost/turns across the last 10 runs, with a tag filter and expandable per-task history./path-to-ga— GA-readiness report for the tasks taggedpath-to-ga. Its task table deliberately answers under stricter rules than every other surface, and both differences are load-bearing:- De-tagged tasks are dropped. A
run.jsontag is a historical stamp, so elsewhere (including/trends) a task de-tagged upstream lingers until the last run that predates the removal ages out. Here a task is dropped once a newer run in the window carries it without the tag — proof of removal. A task that merely stopped appearing is unknowable, so it is kept and dated. - Mature carry-forwards are not passes. Elsewhere a skipped-but-carried-
forward row counts as a pass; here it is excluded from both the numerator
and the denominator, so the rate reports only measured runs (
—when nothing executed). The headline tile and chart above the table keep the ordinary mature-blind, union-over-window semantics — they feed the front page — which is why they read higher than the table, and why the page says so in prose.
- De-tagged tasks are dropped. A
/watchlist— what needs attention, ranked over the recent-runs window: tasks and skills scored on failures, regressions and turn-budget pressure (lib/watchlist.ts)./scribe— the Autopilot (aria/Composer) suite, run bycoder_eval_uipath'sUiPath.Autopilot.Eval.Manualpipeline. This is the one surface that reads a different blob container (aria-runs, notruns) — see Sources below./runs/latest— shortcut that redirects to the newest run id./runs/<run-id>— run summary (pass rate, cost, duration) + one row per task. A "Download run (.zip)" button bundles the whole run folder./runs/<run-id>/<task-id>— per-task detail: success-criteria cards, artifact downloads, flow debug table, tool timeline, message timeline (per-message generation / exec time and output / cache-write / cache-read tokens, with each row expandable into thinking / tool / text sub-rows), tail oftask.log. A "Download folder (.zip)" button bundles this task's folder.
<task-id> is the same string the eval framework writes to
task_results[].task_id (e.g., skill-flow-calculator) and equals the
subdir name under <run-id>/<variant-id>/.
A coder_eval experiment that declares variants: runs every task once per arm
and writes each arm to its own subtree (<run-id>/<variant-id>/<task-id>/<NN>/),
stamping variant_id on every run.json row. An experiment with no variants:
still writes one arm, named default — which is why a single-arm run and a
multi-arm run have the same shape on disk.
Evalboard reads that directly:
- The run page's Pass rate tile reports one entry per arm (side by side, so
the tile keeps the height of the single-arm version and the tiles beside it are
not stretched), and reports no pooled rate at all. A blended rate would average configurations that were
deliberately made to differ (and would move when the arms are merely
reordered), so on a variant run it is not a number anyone wants. The tile
states the observed
spreadbetween arms and nothing more. - Total cost and Time keep their pooled totals and carry the per-arm split on their sub-line, in place of p50/p90. A run's spend is a real operational number however many arms produced it, in a way a run's pass rate is not.
- The task grid gains a Variant column and keeps one row per (task, arm). Replicates of one arm still collapse into a single row; arms never collapse into each other — that difference is the measurement. The default ordering ranks a task by its worst arm rather than ranking each row on its own, so failures still sort to the top while a task's arms stay adjacent. Ranking rows independently splits exactly the pairs worth reading: a task whose arms disagree ends up with one row at the top of the grid and the other at the bottom.
- Task detail is addressed by
?v=<variant-id>alongside?r=<replicate>, so each arm opens its own transcript, log, criteria and artifacts. A link that names no arm resolves to the task's first —defaulton an ordinary run, the leading arm on an experiment run, whose arms are all named and which therefore has nodefaultsubtree for a bare link to land in. A link naming an arm the run does not have is a 404, not a fallback: the point of addressing an arm is that you get that arm or nothing.
Every one of those is inert on a run without variants: the column is dropped, the
tiles keep their existing pooled numbers and percentiles, and no link carries
?v=. Cross-run comparison (two separate run ids) is a different feature and is
not what this does.
Statistics are deliberately not computed here. Whether a gap between arms is
real is a question about variance, and coder_eval already answers it in the
experiment report it writes beside run.json (experiment.md — win rates,
per-task comparison, most divergent tasks, paired comparison with Welch t-test
and bootstrap CIs).
A source is one blob container of runs (lib/sources.ts), usually surfaced as
its own tab. One deployment serves all of them — the container is a runtime
dimension threaded through the data layer as a trailing
source: Source = DEFAULT_SOURCE parameter, not a build-time env var:
| Source | Container | Surface |
|---|---|---|
skills (default) |
runs |
Everything not listed below |
scribe |
aria-runs |
/scribe |
gha |
runs-gha |
none — direct links only |
Non-default sources are selected by a ?src=<id> query param, which every
run-scoped page and API route reads. An absent or unrecognised src resolves to
the default source (sourceById coerces rather than throwing, so a stray param
in a shared link degrades to the skills dashboard instead of an error page).
A source need not have a tab. NAV in app/layout.tsx is a hardcoded array
and does not iterate SOURCES, so gha is registered and deliberately unlisted:
ad-hoc runs uploaded by UiPath/skills' run-coder-eval dispatch, reachable only
by the direct link in the GitHub run summary, expiring after 14 days under the
storage account's expire-runs-gha-14d rule. Registration is still mandatory —
sourceById is the only path by which a container becomes reachable at all.
Two invariants worth preserving if you add a source:
- Run ids are only unique within a container. Every suite names runs
YYYY-MM-DD_HH-MM-SS, so the same id routinely exists in two containers. This is why each source gets its own cache dir (runsDirFor), whylib/blob.tsscopes its in-flight dedupe keys by container, and whyunstable_cachekeys inlib/overview.tsandlib/trends.tsall carrysource.id. Drop any one of those and one source starts serving another's data for a colliding id — with no error. - A run whose id is not date-shaped is invisible to the windowed views.
getRunListing,loadRecentRunsInner,getOverviewandlistRunIdsInWindowfilter onparseRunIdDate, so such runs surface only in the ad-hoc section. A new source's page therefore needs its OWNgetAdhocRunListingsection, or ad-hoc uploads to that container land nowhere reachable — unless it is unlisted by design, likegha, where nothing enumerates andapp/runs/[id]reads by id on demand. Note also thatgetAdhocRunListingloads per-run metadata for every non-date-shaped id in the container before truncating to the display limit, so a source expecting a steady stream of ad-hoc uploads needs its own container, not a prefix inruns. - Local mode is per-source too.
listRunIdsresolvesrunsDirFor(RUNS_DIR, source)whenEVALBOARD_LOCAL_RUNS_DIRis set, so/scribereads<local>-scribe. Listing off the bare local dir instead — which is what shipped first — returns the default source's ids for every source while the readers resolve under the sibling, so the listing and the reads disagree about which container they describe.lib/__tests__/source-isolation.test.tspins both halves; it's the only test that exercises the reader layer, where the invariant above actually lives.
/api/file?run=<id>&path=<relpath>[&src=<source>]serves.flow,.uipx, etc. with path-traversal guard (resolveSafePath)./api/download?run=<id>[&task=<id>][&v=<variant>][&src=<source>]streams a zip of a task folder (withtask, from the named arm) or the whole run (without). Files are gathered bycollectTaskFiles/collectRunFiles, which reuse thewalkArtifactsnoise filter, and zipped bylib/zip.ts(a dependency-free DEFLATE writer).- Pass rows render green (
bg-green-50 text-green-700), failures render red (bg-red-50 text-red-700), on a white background.