Can you tell how accurate a group of noisy judges was on a test without an answer key — using only their votes? Algebraic evaluation (the NTQR logic) says yes: it builds exact, universal polynomials that link the observed agreements and disagreements of an ensemble to the unknown statistics of its correctness, turning evaluation into an inverse problem solved with algebra rather than probability. Its distinctive safety value is that it can warn you when an ensemble is not good enough to be evaluated at all.
This project tests that promise end to end on real local large language models, and sharpens the warning into a quantitative, label-free diagnostic. For publication metadata, the manuscript cites the public NTQR 0.8 documentation as the package authority.
Three trader "personas" (optimist / neutral / pessimist), instantiated as
versioned system prompts, each make a binary bullish/bearish call on the
same 64 market scenarios. We run the identical trio through six locally-hosted
models via Ollama
and, for each model, recover per-persona, per-label accuracy with
ntqr.r2.ErrorIndependentEvaluation (unsupervised — no answer key), scoring
it against an authored truth used only as a check. The headline, non-obvious
result: inter-judge disagreement does not imply evaluability. What gates a
solution is a per-judge condition — every judge must actually vary — not an
ensemble one. A label-free evaluability diagnostic predicts, before any solve,
exactly which models the algebra can evaluate.
For a single judge the universal evaluation polynomial is
f_a = P_a·P_{i,a} + P_b·(1 − P_{i,b}) — exact for any noisy judge on any finite
test. A trio of error-independent judges closes the system: the prevalence of
label a is a root of a quadratic in the trio's vote-pattern moments, giving
two logically consistent solutions (the truth and its mirror image). When the
judges are in fact error-correlated, the prevalence root can go imaginary — an
iron-clad, answer-key-free alarm that the assumption is violated. See
docs/essays/SimplestExampleOfEvaluationWithAlgebraicGeometry.md
and the manuscript §2.5.
The real recovery ranking is hardened and generalised by deterministic, offline checks:
- A confidence interval on the real number (manuscript §3.1): bootstrapping the authored scenarios puts a 95% CI on the recovery MAE that sits well inside the sampling-noise floor — the headline is not a lucky single draw.
- Convergence across many ensembles (§3.7): over hundreds of synthetic
error-independent ensembles the recovery error falls like sampling noise
(
~1/√Q), with the slope stable across diverse prevalences and accuracies. - Two honest failure boundaries: the imaginary-root alarm is sufficient but not necessary — it catches an anti-correlated judge pair with zero false positives, yet positive common-mode correlation (the shared-pre-training analogue) silently degrades recovery without tripping it (§3.7); and the two-solution tie-break inverts once judges are no longer clearly better than random (§3.8).
- A method trade-off, not a winner (§3.9): the exact error-independent algebra is more accurate than simple majority-voting where its assumption holds, but majority-voting is more robust in that tie-break danger zone — so the harness reports both.
- A validation matrix (§3.10): an offline report reads the LLM, simulation, bootstrap, and figure-sidecar artifacts and checks every load-bearing statistical claim before the manuscript is rendered.
| Path | Contents |
|---|---|
docs/manuscript/ |
Generated manuscript.md, supplemental.md, README.md + template config.yaml/preamble/bibliography |
output/manuscript/, output/figures/ |
template-renderer input generated by scripts/z_generate_manuscript_variables.py |
scripts/manuscript/ |
_values.py (single source of truth for every number), render.py ({{TOKEN}} injection), templates/*.tmpl |
scripts/llm/ |
the 6-model persona sweep + LLM client and collect-votes glue (_llm, _analysis); its computation lives in python/src/ntqr/llm/ |
scripts/sim/ |
deterministic, no-network studies: synthetic_recovery_study.py (convergence/alarm/tie-break/robustness/EI-vs-MV) + bootstrap_recovery_ci.py (CI on the real recovery) |
scripts/validation/ |
offline statistical validation matrix for the manuscript's load-bearing claims |
scripts/ntqr/, scripts/baselines/ |
NTQR method demos and naive baselines |
python/src/ntqr/ |
the vendored ntqr algebraic-evaluation library |
outputs/llm/multi_llm_personas_evaluation/ |
generated data (summary.json, per-model JSON) + figures |
outputs/sim/synthetic_recovery_study/ |
synthetic-study data (summary.json, sweep sidecars) + convergence/alarm/tie-break figures |
outputs/sim/bootstrap_recovery_ci/ |
scenario-bootstrap CI on the real recovery MAE (summary.json + histogram figure + .data.json sidecar) |
outputs/validation/statistical_validation/ |
generated validation report, sidecar, and validation-matrix figure |
docs/essays/ |
conceptual essays on algebraic evaluation |
tests/ |
test_llm_analysis.py, test_recovery.py, test_simulation.py, test_bootstrap.py, test_validation_report.py, test_manuscript_render.py |
Reproduce this project as a standalone checkout. The final PDF uses the public
docxology/template renderer as a sibling checkout or any explicit
TEMPLATE_DIR you provide:
# 1. Regenerate the LLM data (needs a running Ollama with the six models pulled)
.venv/bin/python scripts/llm/multi_llm_personas_evaluation.py --set temperature=0.1 --set timeout=45
# 2. Regenerate the synthetic-validation study (deterministic, no network)
.venv/bin/python scripts/sim/synthetic_recovery_study.py
# 2b. Bootstrap a CI on the real recovery MAE (deterministic, offline)
.venv/bin/python scripts/sim/bootstrap_recovery_ci.py
# 3. Regenerate the offline validation matrix
.venv/bin/python scripts/validation/statistical_validation_report.py
# 4. Render the manuscript from templates + data (numbers are injected, never typed)
.venv/bin/python scripts/manuscript/render.py
.venv/bin/python scripts/z_generate_manuscript_variables.py
(cd "${TEMPLATE_DIR:-../template}" && \
uv run python scripts/03_render_pdf.py --project ntqr_llm_personas)
# 5. Verify the prose has not drifted from the data, and run the suite
.venv/bin/python scripts/manuscript/render.py --check
.venv/bin/python -m pytest tests/ -qThe manuscript is a generated artifact: every numeric claim is a {{TOKEN}}
filled from generated JSON artifacts, so the prose can never silently drift from
the measurements.
- Manuscript:
docs/manuscript/manuscript.md - Supplemental:
docs/manuscript/supplemental.md
Archived on Zenodo — cite the concept DOI (always resolves to the latest
version): 10.5281/zenodo.20498699.
The v1.0.0 record is 10.5281/zenodo.20498700.
Machine-readable metadata lives in CITATION.cff,
.zenodo.json, and codemeta.json.
@software{friedman_ntqr_llm_personas_2026,
author = {Friedman, Daniel Ari},
title = {Recovering LLM-Persona Accuracies from Unlabeled Votes},
year = {2026},
publisher = {Zenodo},
version = {1.0.0},
doi = {10.5281/zenodo.20498699},
url = {https://doi.org/10.5281/zenodo.20498699}
}