Run from the repository or extracted research-package root. No model API key is needed. New results use generated states and the MiniGrid simulator, not human or LLM annotation. The production updater is not replaced.
The validated interpreter is Python 3.10, with NumPy 2.2.6, SciPy 1.15.3,
Matplotlib 3.8.4, pytest 9.0.3, MiniGrid 3.0.0 and Gymnasium 1.3.0. The full
repository environment additionally has scikit-learn 1.7.2 and pandas 2.3.3
for older external-data diagnostics. requirements-research.txt in the
package pins the dependencies for the packaged studies. Set no NOUS_*
ablation overrides. macOS/CPU was used; runtime comparisons are descriptive.
PYTHONPATH=. python benchmark/audit_nous_unified_study.py
PYTHONPATH=. python benchmark/audit_nous_minigrid.py
PYTHONPATH=. python benchmark/audit_nous_compatible_studies.py
PYTHONPATH=. python -m pytest -qThe main repository suite has 109 passing tests and one optional skip. The minimal research package contains only relevant tests, so its count is lower. The mutable-state audit reconstructs policies (refitting with original training seeds when necessary), metrics and all 45,000 ledgers/receipts. The MiniGrid audit reconstructs all 36 old/new certificates and replays both final choices for every one of the 9,000 test seeds. The control audit regenerates all 11,900 fixed-seed calibrations, including 2,000 sequential release tests. These checks are not independent proof review.
Additional scrutiny uses SymPy (pinned in the research package) and the
three tests in tests/test_nous_proof_scrutiny.py. The saved result is
data/nous_proof_scrutiny_20260907.json; its interpretation and deliberate
witness-assumption violation are in NOUS_PROOF_SCRUTINY.md. To regenerate,
run PYTHONPATH=. python benchmark/nous_proof_scrutiny.py in a separate copy
after moving that copy's saved result aside. The runner refuses overwrite.
The conditional-binomial sums exhaust the specified finite sample laws but
evaluate probabilities in floating point; they are not Monte Carlo estimates.
Reviewer-response extensions add six bounded-misspecification tests and three
witness-power/selection tests. No frozen C2 code or studies are modified.
The new modules are nous_misspecification_certificate.py and
nous_witness_power.py; their theory notes specify every budget and splitting
assumption. The current draft is 19 pages, with main text ending on page 9.
The shared-observation theorem is now Theorem 1 with proof in Appendix A;
older theory notes call it Theorem U. The remaining theorem numbers shift by
one. This restructuring changed presentation, not frozen experiments or proofs.
The fresh continuous-history witness study and post-hoc MiniGrid sensitivity are reproduced with:
PYTHONPATH=. python benchmark/audit_nous_learned_witness.py
PYTHONPATH=. MPLCONFIGDIR=tmp/plot-cache XDG_CACHE_HOME=tmp/cache python benchmark/build_nous_witness_artifacts.pyThe audit regenerates 252,000 records across 48 disjoint split seeds and
reconstructs every lock, certificate and held-out metric. Its protocol was
specified before the first run, but is not an externally preregistered study.
The generator run_nous_learned_witness.py and post-hoc analysis
nous_sensitivity.py refuse to overwrite existing outputs. For regeneration
use a separate copy and move its result directory/file aside first.
The shared-experiment theorem adds exact rational checks, not a synthetic
performance benchmark: tests/test_nous_same_experiment.py and
data/nous_same_experiment_20260907.json. Run nous_same_experiment.py from
the benchmark directory or with PYTHONPATH=. in a separate copy after moving
that copy's result aside; it refuses overwrite. The analytic proof covering
the full persistence interval is NOUS_SAME_EXPERIMENT_THEORY.md.
data/nous_unified_20260906/: original integrated mutable-state experiment and matched-theorem numerical illustration; unchanged.data/nous_minigrid_20260907/: external task, calibration arrays, policy locks, saved actions/rewards and held-out decisions.data/nous_compatible_20260907/: all initial zero-power/control results.data/nous_compatible_phase_20260907/: complete fresh diagnostic grid, explicitly motivated by the preceding negative result.data/nous_compatible_audit_20260907.json: full control regeneration audit.
Every frozen manifest contains protocol and implementation hashes. The control studies preserve sufficient seeds/statistics to regenerate their records rather than storing hundreds of millions of redundant draws. The population identified interval uses generator truth only for diagnostic analysis; the certificate receives only policies, reports and witness masks.
The mutable-state runner accepts --output for a new directory. The September
7 runners intentionally refuse to overwrite their fixed study directories.
For fresh reproduction, use a separate extracted copy, move its existing
result folders to an explicitly named archive within that copy, then run:
PYTHONPATH=. python benchmark/run_nous_compatible_controls.py
PYTHONPATH=. python benchmark/run_nous_compatible_phase.py
PYTHONPATH=. python benchmark/run_nous_minigrid.pyDo not overwrite or tune against the original frozen outputs. Compare numeric results and decisions, not compressed-file timestamp hashes or elapsed times. MiniGrid uses four worker processes; calibration sensor staging is declared in its protocol, while every test episode uses the full controller route.
PYTHONPATH=. MPLCONFIGDIR=tmp/plot-cache python benchmark/build_nous_frontier_artifacts.py
TEXINPUTS=paper/vendor/iclr2027: pdflatex -interaction=nonstopmode -halt-on-error -output-directory output/pdf paper/nous_iclr2027.tex
TEXINPUTS=paper/vendor/iclr2027: pdflatex -interaction=nonstopmode -halt-on-error -output-directory output/pdf paper/nous_iclr2027.tex
pdftoppm -r 100 -png output/pdf/nous_iclr2027.pdf tmp/pdfs/nous_iclr2027Create the listed output/cache directories if absent. The package includes the generated matched proof, references and older tables/figure, so it does not require the previous identified-author manuscript to compile. The original official ICLR 2027 style is vendored unmodified. It supplies an automatic review header; no paper has been submitted by this workflow.
build_nous_research_package.py uses an explicit allowlist, never the entire
workspace. It excludes credentials, dialogues, annotation files, virtual
environments, Git metadata and unrelated work. It preserves the five actual
Nous modules needed for this experiment, with a minimal research-only package
initializer instead of the production application entry point. The module
contents used by the experiments remain hash-identical.
Absolute workspace prefixes in packaged JSON metadata are converted to
relative paths for portability and anonymity; the originals are not edited.
PACKAGE_MANIFEST.json records original and packaged hashes for each such
relocation. Summary-file hashes therefore differ for relocated metadata even
though scientific values, seed lists and implementation hashes are unchanged.
Downstream JSON hash references are updated consistently within the package
and these metadata-only changes are recorded in the same relocation manifest.
The Nous license is retained as third-party attribution for the cited prior implementation. Its copyright holder is not a declaration of this submission's authors. Final author approval of licensing/anonymity treatment is still required. The ZIP is local; it has not been published or sent to reviewers. Older UCI figures are included as historical saved diagnostics, not newly confirmed data or part of the new iid guarantee.