Skip to content

Latest commit

 

History

History

README.md

Reproduce the integrated Nous frontier paper

Run from the repository or extracted research-package root. No model API key is needed. New results use generated states and the MiniGrid simulator, not human or LLM annotation. The production updater is not replaced.

Environment

The validated interpreter is Python 3.10, with NumPy 2.2.6, SciPy 1.15.3, Matplotlib 3.8.4, pytest 9.0.3, MiniGrid 3.0.0 and Gymnasium 1.3.0. The full repository environment additionally has scikit-learn 1.7.2 and pandas 2.3.3 for older external-data diagnostics. requirements-research.txt in the package pins the dependencies for the packaged studies. Set no NOUS_* ablation overrides. macOS/CPU was used; runtime comparisons are descriptive.

Audit existing results

PYTHONPATH=. python benchmark/audit_nous_unified_study.py
PYTHONPATH=. python benchmark/audit_nous_minigrid.py
PYTHONPATH=. python benchmark/audit_nous_compatible_studies.py
PYTHONPATH=. python -m pytest -q

The main repository suite has 109 passing tests and one optional skip. The minimal research package contains only relevant tests, so its count is lower. The mutable-state audit reconstructs policies (refitting with original training seeds when necessary), metrics and all 45,000 ledgers/receipts. The MiniGrid audit reconstructs all 36 old/new certificates and replays both final choices for every one of the 9,000 test seeds. The control audit regenerates all 11,900 fixed-seed calibrations, including 2,000 sequential release tests. These checks are not independent proof review.

Additional scrutiny uses SymPy (pinned in the research package) and the three tests in tests/test_nous_proof_scrutiny.py. The saved result is data/nous_proof_scrutiny_20260907.json; its interpretation and deliberate witness-assumption violation are in NOUS_PROOF_SCRUTINY.md. To regenerate, run PYTHONPATH=. python benchmark/nous_proof_scrutiny.py in a separate copy after moving that copy's saved result aside. The runner refuses overwrite. The conditional-binomial sums exhaust the specified finite sample laws but evaluate probabilities in floating point; they are not Monte Carlo estimates.

Reviewer-response extensions add six bounded-misspecification tests and three witness-power/selection tests. No frozen C2 code or studies are modified. The new modules are nous_misspecification_certificate.py and nous_witness_power.py; their theory notes specify every budget and splitting assumption. The current draft is 19 pages, with main text ending on page 9. The shared-observation theorem is now Theorem 1 with proof in Appendix A; older theory notes call it Theorem U. The remaining theorem numbers shift by one. This restructuring changed presentation, not frozen experiments or proofs.

The fresh continuous-history witness study and post-hoc MiniGrid sensitivity are reproduced with:

PYTHONPATH=. python benchmark/audit_nous_learned_witness.py
PYTHONPATH=. MPLCONFIGDIR=tmp/plot-cache XDG_CACHE_HOME=tmp/cache python benchmark/build_nous_witness_artifacts.py

The audit regenerates 252,000 records across 48 disjoint split seeds and reconstructs every lock, certificate and held-out metric. Its protocol was specified before the first run, but is not an externally preregistered study. The generator run_nous_learned_witness.py and post-hoc analysis nous_sensitivity.py refuse to overwrite existing outputs. For regeneration use a separate copy and move its result directory/file aside first.

Frozen artifacts

The shared-experiment theorem adds exact rational checks, not a synthetic performance benchmark: tests/test_nous_same_experiment.py and data/nous_same_experiment_20260907.json. Run nous_same_experiment.py from the benchmark directory or with PYTHONPATH=. in a separate copy after moving that copy's result aside; it refuses overwrite. The analytic proof covering the full persistence interval is NOUS_SAME_EXPERIMENT_THEORY.md.

  • data/nous_unified_20260906/: original integrated mutable-state experiment and matched-theorem numerical illustration; unchanged.
  • data/nous_minigrid_20260907/: external task, calibration arrays, policy locks, saved actions/rewards and held-out decisions.
  • data/nous_compatible_20260907/: all initial zero-power/control results.
  • data/nous_compatible_phase_20260907/: complete fresh diagnostic grid, explicitly motivated by the preceding negative result.
  • data/nous_compatible_audit_20260907.json: full control regeneration audit.

Every frozen manifest contains protocol and implementation hashes. The control studies preserve sufficient seeds/statistics to regenerate their records rather than storing hundreds of millions of redundant draws. The population identified interval uses generator truth only for diagnostic analysis; the certificate receives only policies, reports and witness masks.

Rerunning generators

The mutable-state runner accepts --output for a new directory. The September 7 runners intentionally refuse to overwrite their fixed study directories. For fresh reproduction, use a separate extracted copy, move its existing result folders to an explicitly named archive within that copy, then run:

PYTHONPATH=. python benchmark/run_nous_compatible_controls.py
PYTHONPATH=. python benchmark/run_nous_compatible_phase.py
PYTHONPATH=. python benchmark/run_nous_minigrid.py

Do not overwrite or tune against the original frozen outputs. Compare numeric results and decisions, not compressed-file timestamp hashes or elapsed times. MiniGrid uses four worker processes; calibration sensor staging is declared in its protocol, while every test episode uses the full controller route.

Paper and figures

PYTHONPATH=. MPLCONFIGDIR=tmp/plot-cache python benchmark/build_nous_frontier_artifacts.py
TEXINPUTS=paper/vendor/iclr2027: pdflatex -interaction=nonstopmode -halt-on-error -output-directory output/pdf paper/nous_iclr2027.tex
TEXINPUTS=paper/vendor/iclr2027: pdflatex -interaction=nonstopmode -halt-on-error -output-directory output/pdf paper/nous_iclr2027.tex
pdftoppm -r 100 -png output/pdf/nous_iclr2027.pdf tmp/pdfs/nous_iclr2027

Create the listed output/cache directories if absent. The package includes the generated matched proof, references and older tables/figure, so it does not require the previous identified-author manuscript to compile. The original official ICLR 2027 style is vendored unmodified. It supplies an automatic review header; no paper has been submitted by this workflow.

Packaging and anonymity

build_nous_research_package.py uses an explicit allowlist, never the entire workspace. It excludes credentials, dialogues, annotation files, virtual environments, Git metadata and unrelated work. It preserves the five actual Nous modules needed for this experiment, with a minimal research-only package initializer instead of the production application entry point. The module contents used by the experiments remain hash-identical. Absolute workspace prefixes in packaged JSON metadata are converted to relative paths for portability and anonymity; the originals are not edited. PACKAGE_MANIFEST.json records original and packaged hashes for each such relocation. Summary-file hashes therefore differ for relocated metadata even though scientific values, seed lists and implementation hashes are unchanged. Downstream JSON hash references are updated consistently within the package and these metadata-only changes are recorded in the same relocation manifest.

The Nous license is retained as third-party attribution for the cited prior implementation. Its copyright holder is not a declaration of this submission's authors. Final author approval of licensing/anonymity treatment is still required. The ZIP is local; it has not been published or sent to reviewers. Older UCI figures are included as historical saved diagnostics, not newly confirmed data or part of the new iid guarantee.