Skip to content

Latest commit

 

History

History
65 lines (37 loc) · 14.7 KB

File metadata and controls

65 lines (37 loc) · 14.7 KB

Paired evaluation of learning

Historical examples reaching an author is evidence of integration. It does not establish better tests. The original single-task ablation remains visible: both variants caught four mutations, memory produced fewer cases and took longer. Its order and default model were unpinned.

The repeated native run covers two specifications and two repetitions. The no-memory arm caught 12/12 defects; the memory arm caught 9/12, with one author timeout. This result is inconclusive and establishes no learning benefit. Implementation hashes, raw receipts, development-lineage limits and a delayed-timeout caveat remain visible.

scripts/learning-evaluation.js adds a prospectively declared paired protocol. Each independent specification runs with and without the same frozen, genuinely validated historical store. Both arms use the same declared worker argv, requirements, source, three role calls, response-byte request limit, and subprocess timeout. New runs also bind the locally resolved worker launcher and ascertainable script files before invoking either arm. Fixture order is seeded and shuffled; first-arm order is counterbalanced per specification across repetitions after a seeded initial choice. The order seed controls scheduling, not provider randomness. Runs are sequential so competing calls do not confound latency.

The controller keeps evaluation reference tests and planted defect payloads separate from historical memory, agent stdin and generation roots. This establishes exclusion from supplied inputs and trial-root files, not OS read confinement or independently observed zero tool use. The earlier frozen Codex campaign prompted against tools without retaining a structured native tool-event audit; its receipts do not establish observed zero tool use. It first verifies a green reference baseline and actual named-case failure for every labeled defect. Generated tests must pass independent review and TestLore candidate validation before applying in the disposable trial project. Defects are then executed against those generated tests. Import/process errors do not count as detected faults. Review rejection, incomplete execution, and failed trials remain visible and contribute zero recall.

Run with an existing authenticated worker

From a source checkout with installed dependencies:

node scripts/learning-evaluation.js \
  --output /absolute/private/new-learning-evaluation \
  --agent '["/absolute/node","/absolute/testlore/src/adapters/codex.js"]' \
  --identity 'codex/VERSION/default-model-unpinned' \
  --repeat 3 --seed 20260930 --max-calls 54

This explicitly invokes generation and consumes the worker's existing allowance. The controller does not install a worker, authenticate, request provider credentials, or fall back to a separately billed provider. The included Codex adapter uses its normal native authentication and default model. Record the actual installed CLI version in --identity; for a pinned provider, use a wrapper with a fixed model/version and record that identity. --identity remains an operator declaration. workerIdentity.model.requested records an explicit --model/-m worker argument when present; it is null when no request is ascertainable. workerIdentity.model.verified remains null, including when a worker response claims a model name. The protocol does not certify the model/version/seed declaration or recover token billing from workers that do not report it.

The default scope has three independent authored contracts (string normalization, stable deduplication, half-open interval overlap), one separate historical numeric-clamp example and nine planted defects. Three repetitions yield eighteen trials and a maximum of fifty-four worker calls. A smaller local run uses --repeat 1 --max-calls 18. No live run is implied by shipping this script. Protocol-fixture workers are marked --evidence-kind protocol-fixture and never support a learning-quality conclusion.

The fixed three-call budget requires one architect task. A response larger than --max-output-bytes (default 65,536 bytes) stops that trial before another call. The identical declared byte bound is enforced by the worker transport and checked on the parsed response; it is not an enforced provider token or dollar budget. Each call is terminated at --timeout-ms (default 115,000 ms; maximum 120,000 ms). The whole run's worst-case worker time is calls × timeout. There are no automatic retries. More memory necessarily changes input length; output-byte totals and raw payloads disclose this difference.

Bring independently authored specifications

--fixtures /absolute/manifest.json accepts a bounded 2 MB JSON manifest. Export the default schema to inspect it:

node --input-type=module -e "import {defaultDataset} from './scripts/learning-evaluation.js'; console.log(JSON.stringify(defaultDataset(),null,2));" > /private/evaluation-fixtures.json

The manifest contains schemaVersion:1, id, history, and fixtures. Historical entries have distinct id, specificationId, requirements, source files:[{path,content}], and tests:[{path,content}]. Evaluation entries have source files, independent referenceTests, and defects:[{id,files}] that replace existing source files. Paths are bounded canonical src/name.js and test/name.test.js; no manifests, binaries, dependency installation or arbitrary commands appear in fixtures. Maximum scope is eight histories, twelve specifications and eight defects per specification. repeat is 1–5 and total role-call budget cannot exceed 360. IDs, requirement hashes and source hashes must not overlap across specifications. These mechanical checks prevent exact leakage; maintainers must separately audit semantic independence and whether each defect follows the independent contract.

Do not add evaluation specifications, reference tests, defect labels or defect implementations to historical memory. Do not run a training/tuning loop on held-out specifications. Freeze the manifest and worker settings before collecting results; changes require a fresh output directory and a separately named run.

Receipts and interpretation

Every run writes an immutable manifest, raw historical validation, reference/defect executions, per-trial role inputs/outputs, candidate validation, generated-baseline cases, and fault outcomes. Output directories must be new. Directory permissions default to 0700 and files to 0600; raw source and generated tests are private until a maintainer reviews them. Workspace paths are normalized, but source is intentionally retained locally. These are not public pilot exports.

New manifests and summaries include workerIdentity, verificationScope and implementationHashes. Bounded streaming SHA-256 reads bind the resolved launcher (at most 512 MiB), the Node launcher where ascertainable, explicit Node entrypoint/preload files, inline node -e/-p source and existing absolute file operands (at most 8 MiB each). Simple direct Node and /usr/bin/env node shebangs also bind their interpreter/Node files. Resolution follows the child invocation: relative launcher paths, relative PATH entries and relative Node entrypoints use each disposable trial cwd. A repository-relative script that is absent there fails preflight instead of silently resolving against the evaluator checkout. Use absolute script paths for repository adapters. Credential-like file operands are rejected rather than read. No runtime or model probe, shell, authentication or provider call is used for identity collection.

The evaluator hashes its implementation modules and all shipped src/adapters/*.js files, including the Codex adapter, at preflight. Each role checks those files and its launcher/script identities immediately before launch and after receiving a response, before accepting role output. File content, resolved symlink target or permission changes invalidate the trial. The rejected response, normalized bytes, attempt, whether transport was invoked and before/after verification status remain in the role receipt; attempts rejected before launch are counted and have spawned:false. Summary and arm calls/attemptedCalls count attempted roles, including those identity rejections; spawnedCalls counts worker transport invocations, including invocations that subsequently fail. This records invocation of callAgent, not independent proof that the child executed. Identity drift does not retry, rebind or resume an earlier trial. Failed preflight writes preflight.json and controller.json before disposable workspace cleanup. Previous raw experiments remain bound to their original hashes and do not acquire this verification retroactively.

verificationScope.fullRuntimeAttestation is always false. These local checks cannot prevent a file from changing and changing back between observations, prove the interpreter of an arbitrary wrapper, or attest every imported dependency, child process, native provider binary, provider configuration/environment or remote model. A Node launcher recognized only by filename is labeled accordingly and has an unknown version unless its hash matches the running controller's Node binary. The controller Node file is a preflight snapshot; it is rechecked per role only when it is also the resolved worker Node launcher or simple shebang Node. Imported and child code remain explicitly unbound. Paths are normalized in local receipts as described above; raw argv and source are retained locally.

The summary compares defect recall, passing case count, generation time and normalized JSON bytes. More cases are not intrinsically better tests. With accountingVersion:2, the per-trial total clock starts before disposable project/source setup and the frozen memory copy. setupMs records that setup, and retrievalMs separately records recallLessons. Total time includes both, generation, candidate validation, repeated generated baselines and defect execution. Run-level historical validation, independent reference ground truth and final trial receipt writing remain outside this per-trial total. Paired deltas keep rejected trials rather than filtering them out. All raw failure IDs are available to diagnose faults.

Qualification independently repeats the reference baseline, generated baseline and each held-out defect at least twice. --stability-runs 2 is the default (2–5 permitted). Native case identities, statuses and exit codes must agree; incomplete collection, skipped cases or changing outcomes invalidate a trial. Partial fault detections from an invalid trial remain in the raw receipt but contribute zero to the arm's qualified detection count.

Each call's inputBytes counts UTF-8 bytes of JSON.stringify on the role payload before the transport appends transportBudget. outputBytes counts UTF-8 bytes of JSON.stringify on the parsed response. Raw stdout bytes, whitespace and the added transport metadata are excluded. These are normalized JSON payload measurements, not complete wire sizes or tokens; a failed call without a parsed response contributes no measured response bytes. byteAccounting declares these definitions in the manifest, trial and summary. Arm summaries include actual calls, setup and retrieval durations, total trial time and stable-trial counts. billingUSD: null means unknown provider billing; it never means a free call or an estimated dollar saving. The subscription worker consumes the existing allowance.

Each independent reference baseline and held-out fault execution is written immediately to a separate immutable local receipt before validation can throw. A failed reference baseline or undemonstrated fault also gets an incomplete labels-FIXTURE.json record containing the executions and rejection reason. Workspace cleanup removes disposable projects while retaining this ground-truth evidence; no worker is called when ground truth fails.

After manifest creation, controller.json is retained on success and controller failure. Its elapsedMs spans evaluateLearning entry through disposable workspace cleanup. It separately records controller setupMs (validation, output setup and manifest creation), sharedPreparationMs (historical project and candidate validation plus frozen memory loading), and groundTruthMs (independent reference baselines and held-out fault validation). These shared costs are measured once without assigning them to either arm; phases not reached are null. Successful summaries contain the same controller record. Process startup and final summary/controller receipt writes are excluded, and provider tokens or billing are not inferred from elapsed time.

qualification-contracts.js supplies six newly authored contracts and eighteen independently executable defects: CSV fields, duration formatting, first-occurrence binary search, frequency maps, common prefixes and numeric version comparison. Two distinct historical contracts are frozen separately. A live run freezes the manifest before invoking a worker, retains all three repetitions and discloses absent relevant memory. A larger dataset or more repetitions still cannot manufacture a learning gain.

The frozen six-frozen-contracts-20260930-v1 campaign remains bound to its original evaluator and dataset hashes. Its recorded totalMs starts after project setup, memory copying and retrieval, so those costs are omitted and cannot be recovered from the totals. Its input/output counters likewise measure normalized JSON payloads, excluding appended transportBudget and raw stdout bytes. Two frozen fault labels overstate what their implementations exercise: duplicate-last changes the binary search to an upper-bound search and then fails its equality check, returning a miss rather than the last duplicate; lexical-components leaves version components as strings and the existing safe-integer check rejects otherwise valid version strings. Those implementations are held-out conformance faults, but the labels do not establish last-duplicate or lexical-comparison fault coverage. The dataset and prior receipts are preserved, not relabeled or tuned using run outcomes.

Repeated attempts are clustered by specification. A two-sided exact sign test uses the mean recall difference for each independent specification, with ties excluded. A scoped positive/negative difference is labeled only with at least six specifications, three repetitions, complete trials, and p < 0.05. The default three-specification experiment therefore remains inconclusive for learning improvement even when an observed average differs. The test is a coarse exploratory statistic, not a power analysis or production-corpus guarantee. It does not establish superiority over other memory systems, train model weights, or authorize different test selection.