jevals
Agent evals and guardrails in one typed decision request
TLDR
SYNOPSIS
jevals [--backend url] command [options]
DESCRIPTION
jevals runs agent evaluations and request-path guardrails as typed decision questions instead of a generative LLM judge. Several checks on one trace are packed into a single backend request that returns calibrated probabilities (yes/no, multiple choice, or a scored rubric). Plain code handles splitting, tool-call matching, and regex or Presidio entity detection; the model only answers the judgment questions.Built-in evals live under agent (tool choice, groundedness, scope, loops), security (injection, PII/PHI, secrets, toxicity), and quality (faithfulness, relevancy, completeness, and related Ragas-style metrics). The same class can score a dataset offline, monitor production traces, or sit in a Gate that maps answers to allow, escalate, block, or redact.Backends resolve from the environment unless --backend is set: TypeSafe Jev (`TYPESAFEAPIKEY`), Jev through Vercel AI Gateway (`AIGATEWAYAPIKEY`), a local Kev server (`KEVBASEURL` or `kev://host:port`), in-process Laya (`JEVALSBACKEND=laya`), an emulated chat LLM (`OPENROUTERAPIKEY` / `llm:<model>`), or `mock` for tests.
PARAMETERS
--backend spec
Override backend discovery: `jev`, `vercel`, `kev://host:port`, `laya`, `llm:<model>`, or `mock`.
COMMANDS
run path
Evaluate a JSONL dataset. --evals is a comma-separated list of built-in names or YAML paths. --out writes one result row per trace. --limit, --concurrency, --show-failures N, --json, --quiet.check path
Evaluate one JSON sample (or - for stdin). Prints a table, or JSON with --json. Exit 0 if every eval passed, else 1.gate path
Run evals as a policy gate. Exit 0 allow/modify, 2 escalate, 3 block.list
List built-in evals. --category filters; --json prints machine-readable rows.describe name
Show an eval's required fields, state keys, and questions.validate paths
Parse YAML/JSON eval files and report errors.schema
Print the JSON Schema for declarative YAML evals.docs
Print a compact reference (intended for pasting into an agent context).calibrate path
Fit a threshold table against a label field. Requires --eval and --label. --max-false-pass picks the highest auto-pass rate under that wrong-pass budget.mcp
Run an MCP server, or --install [cursor|claude|vscode|all] to write client config.hook {pre|post}
Claude Code hook: read a PreToolUse/PostToolUse event on stdin and print allow, ask, or deny.bench
Measure requests, tokens, and latency on a fixed dataset. --ragas also runs Ragas on the same rows.
CAVEATS
The project is alpha. Decision-model probabilities need calibration on your own labels before they drive irreversible actions. Chat-LLM backends emulate the typed API and are slower, more expensive, and poorly calibrated (often 0.00 or 1.00). Gates fail open if the backend is still down after retries unless the gate is constructed with `on_error="block"`. PII/PHI detection is stronger with the optional Presidio extra.
HISTORY
jevals is published by Openlayer (2026) as an eval library on top of TypeSafe Jev and compatible local models (Kev, Laya). It reuses Ragas-style metric names while collapsing several LLM-judge round trips into one System One request.
SEE ALSO
pytest(1)
