Skip to content

Repository files navigation

ntqr_llm_personas — algebraic evaluation of LLM-persona ensembles

DOI License: Apache 2.0

Can you tell how accurate a group of noisy judges was on a test without an answer key — using only their votes? Algebraic evaluation (the NTQR logic) says yes: it builds exact, universal polynomials that link the observed agreements and disagreements of an ensemble to the unknown statistics of its correctness, turning evaluation into an inverse problem solved with algebra rather than probability. Its distinctive safety value is that it can warn you when an ensemble is not good enough to be evaluated at all.

This project tests that promise end to end on real local large language models, and sharpens the warning into a quantitative, label-free diagnostic. For publication metadata, the manuscript cites the public NTQR 0.8 documentation as the package authority.

The study in one paragraph

Three trader "personas" (optimist / neutral / pessimist), instantiated as versioned system prompts, each make a binary bullish/bearish call on the same 64 market scenarios. We run the identical trio through six locally-hosted models via Ollama and, for each model, recover per-persona, per-label accuracy with ntqr.r2.ErrorIndependentEvaluation (unsupervised — no answer key), scoring it against an authored truth used only as a check. The headline, non-obvious result: inter-judge disagreement does not imply evaluability. What gates a solution is a per-judge condition — every judge must actually vary — not an ensemble one. A label-free evaluability diagnostic predicts, before any solve, exactly which models the algebra can evaluate.

The algebra (why it works, and how it fails)

For a single judge the universal evaluation polynomial is f_a = P_a·P_{i,a} + P_b·(1 − P_{i,b}) — exact for any noisy judge on any finite test. A trio of error-independent judges closes the system: the prevalence of label a is a root of a quadratic in the trio's vote-pattern moments, giving two logically consistent solutions (the truth and its mirror image). When the judges are in fact error-correlated, the prevalence root can go imaginary — an iron-clad, answer-key-free alarm that the assumption is violated. See docs/essays/SimplestExampleOfEvaluationWithAlgebraicGeometry.md and the manuscript §2.5.

The real recovery ranking is hardened and generalised by deterministic, offline checks:

  • A confidence interval on the real number (manuscript §3.1): bootstrapping the authored scenarios puts a 95% CI on the recovery MAE that sits well inside the sampling-noise floor — the headline is not a lucky single draw.
  • Convergence across many ensembles (§3.7): over hundreds of synthetic error-independent ensembles the recovery error falls like sampling noise (~1/√Q), with the slope stable across diverse prevalences and accuracies.
  • Two honest failure boundaries: the imaginary-root alarm is sufficient but not necessary — it catches an anti-correlated judge pair with zero false positives, yet positive common-mode correlation (the shared-pre-training analogue) silently degrades recovery without tripping it (§3.7); and the two-solution tie-break inverts once judges are no longer clearly better than random (§3.8).
  • A method trade-off, not a winner (§3.9): the exact error-independent algebra is more accurate than simple majority-voting where its assumption holds, but majority-voting is more robust in that tie-break danger zone — so the harness reports both.
  • A validation matrix (§3.10): an offline report reads the LLM, simulation, bootstrap, and figure-sidecar artifacts and checks every load-bearing statistical claim before the manuscript is rendered.

Layout

Path Contents
docs/manuscript/ Generated manuscript.md, supplemental.md, README.md + template config.yaml/preamble/bibliography
output/manuscript/, output/figures/ template-renderer input generated by scripts/z_generate_manuscript_variables.py
scripts/manuscript/ _values.py (single source of truth for every number), render.py ({{TOKEN}} injection), templates/*.tmpl
scripts/llm/ the 6-model persona sweep + LLM client and collect-votes glue (_llm, _analysis); its computation lives in python/src/ntqr/llm/
scripts/sim/ deterministic, no-network studies: synthetic_recovery_study.py (convergence/alarm/tie-break/robustness/EI-vs-MV) + bootstrap_recovery_ci.py (CI on the real recovery)
scripts/validation/ offline statistical validation matrix for the manuscript's load-bearing claims
scripts/ntqr/, scripts/baselines/ NTQR method demos and naive baselines
python/src/ntqr/ the vendored ntqr algebraic-evaluation library
outputs/llm/multi_llm_personas_evaluation/ generated data (summary.json, per-model JSON) + figures
outputs/sim/synthetic_recovery_study/ synthetic-study data (summary.json, sweep sidecars) + convergence/alarm/tie-break figures
outputs/sim/bootstrap_recovery_ci/ scenario-bootstrap CI on the real recovery MAE (summary.json + histogram figure + .data.json sidecar)
outputs/validation/statistical_validation/ generated validation report, sidecar, and validation-matrix figure
docs/essays/ conceptual essays on algebraic evaluation
tests/ test_llm_analysis.py, test_recovery.py, test_simulation.py, test_bootstrap.py, test_validation_report.py, test_manuscript_render.py

Reproduce

Reproduce this project as a standalone checkout. The final PDF uses the public docxology/template renderer as a sibling checkout or any explicit TEMPLATE_DIR you provide:

# 1. Regenerate the LLM data (needs a running Ollama with the six models pulled)
.venv/bin/python scripts/llm/multi_llm_personas_evaluation.py --set temperature=0.1 --set timeout=45

# 2. Regenerate the synthetic-validation study (deterministic, no network)
.venv/bin/python scripts/sim/synthetic_recovery_study.py

# 2b. Bootstrap a CI on the real recovery MAE (deterministic, offline)
.venv/bin/python scripts/sim/bootstrap_recovery_ci.py

# 3. Regenerate the offline validation matrix
.venv/bin/python scripts/validation/statistical_validation_report.py

# 4. Render the manuscript from templates + data (numbers are injected, never typed)
.venv/bin/python scripts/manuscript/render.py
.venv/bin/python scripts/z_generate_manuscript_variables.py
(cd "${TEMPLATE_DIR:-../template}" && \
  uv run python scripts/03_render_pdf.py --project ntqr_llm_personas)

# 5. Verify the prose has not drifted from the data, and run the suite
.venv/bin/python scripts/manuscript/render.py --check
.venv/bin/python -m pytest tests/ -q

The manuscript is a generated artifact: every numeric claim is a {{TOKEN}} filled from generated JSON artifacts, so the prose can never silently drift from the measurements.

Read

Cite

Archived on Zenodo — cite the concept DOI (always resolves to the latest version): 10.5281/zenodo.20498699. The v1.0.0 record is 10.5281/zenodo.20498700. Machine-readable metadata lives in CITATION.cff, .zenodo.json, and codemeta.json.

@software{friedman_ntqr_llm_personas_2026,
  author    = {Friedman, Daniel Ari},
  title     = {Recovering LLM-Persona Accuracies from Unlabeled Votes},
  year      = {2026},
  publisher = {Zenodo},
  version   = {1.0.0},
  doi       = {10.5281/zenodo.20498699},
  url       = {https://doi.org/10.5281/zenodo.20498699}
}

About

Label-free evaluation of LLM judge ensembles: 3 market personas across 6 local Ollama models vote on 64 scenarios, and the NTQR 0.8 algebraic-geometry solver recovers per-persona accuracy from votes alone — no answer key. Includes error-independence controls, bootstrap 95% CIs, and results-injected manuscript builds.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages