Can a cheap measurement of a pretrained checkpoint, taken before any training, predict how well LoRA will adapt it and how much rank it needs?
If it can, you could decide up front whether to use LoRA at rank 8, LoRA at rank 128, or full fine-tuning, instead of discovering it by running all three.
Six checkpoints, three LoRA ranks each, two learning rates, on mathematical reasoning. Two descriptor tiers extracted and tested against a pre-registered analysis.
The short answer so far is no. Neither descriptor tier predicts rank demand
in a way that survives changing the LoRA learning rate. Both pre-registered
associations were measured at 2e-4, then failed to replicate at 1e-4.
| Pre-registered claim | at 2e-4 |
at 1e-4 |
|---|---|---|
| frozen-weight geometry predicts the gap to full SFT | -0.94, p = 0.017 | -0.49, p = 0.356 |
| activation geometry predicts the rank response | +0.94, p = 0.017 | absent from the top five |
At 1e-4 a randomised control ties the best real feature. Amendment 013
committed in advance to reporting this outcome if it occurred.
What does hold
| Finding | Evidence |
|---|---|
| Rank demand varies by checkpoint | minimum sufficient rank spans 8 to never-reached |
| Minimum sufficient rank is not stable across rates | 8, 8, 128, never x3 at 2e-4; 128, 32, 128, never x3 at 1e-4 |
| Higher rank prefers a lower rate, universally | all six checkpoints, no exception; consistent with alpha / r scaling |
| Minimum sufficient rank is not predictable | no feature beats a randomised control at either rate |
| The MLP is the low-rank bottleneck | attention captured 2-3x better at rank 8 in every checkpoint |
What the descriptors do predict, on targets defined after the pre-registered
analysis failed and therefore exploratory: learning-rate sensitivity
(rho = -1.00, 30/30 held-out orderings) and the rank-rate interaction
(rho = +1.00). A candidate application — predicting whether LoRA will retain
most of full fine-tuning — reaches rho = -0.83 at 27/30. None of this is
confirmatory. Read section 6 of the two-rate report, which states how many tests
stand behind those numbers, before citing any of them.
What is not established: one seed, one task, six checkpoints. Start with
docs/TWO_RATE_FINDINGS_2026-08-16.md,
which supersedes the single-rate reports. The earlier
docs/CHECKPOINT_ANALYSIS_2026-08-11.md
and docs/ADVISOR_BRIEF_2026-08-11.md are
accurate for 2e-4 alone and are kept as the record of what was concluded
before the second rate existed.
configs/study.toml the frozen study configuration: models, ranks, rates, thresholds
manifests/ pinned model, benchmark, and gate manifests with commit SHAs
src/lora_rank_demand/ the package; every stage is a subcommand of `lrd`
data.py dataset preparation, contamination audit
training.py full SFT and LoRA via TRL/PEFT/DeepSpeed
benchmark.py lm-eval harness wrapper, generation-config sanitising
descriptors.py Tier 1 frozen-weight descriptors
activations.py Tier 2 task-conditioned activation descriptors
outcomes.py retained gain, gap, minimum sufficient rank
runs.py immutable run specifications and Slurm submission
scripts/ analysis entry points, run on a compute node
slurm/ core pipeline job scripts
campaigns/ one-off staging jobs for specific experiment blocks
diagnostics/ failure triage and audit jobs, not part of the pipeline
cluster.env site-specific partitions, QoS, node exclusions
protocol/
preregistration/ features, targets, and baselines fixed BEFORE outcomes existed
amendments/ numbered protocol changes, each recorded before execution
incidents/ every failure and its repair, including wrong diagnoses
docs/ analysis reports and the data index
results/*.csv every derived table, readable without a compute allocation
tests/ unit suite, 43 tests
RESEARCH_PLAN.md the scientific protocol and gate definitions
PROGRESS_LOG.md chronological execution record; "Resume Here" is the restart point
Everything generated goes under $LRD_DATA_ROOT, never into the source tree.
The repository holds source, manifests, compact summaries, and checksums only.
scripts/project_env.sh derives the source root from its own location and takes
everything else from the environment, so the only required decision is where
generated data lives:
export LRD_DATA_ROOT=/your/data/path/lora-rank-demand # default: /data/user_data/$USER/lora-rank-demand
source scripts/project_env.shOptional overrides: HF_TOKEN_PATH (default ~/.huggingface/token),
LRD_SCRATCH (default /scratch, used for the per-job Triton cache).
python -m venv "$LRD_DATA_ROOT/.venv"
source "$LRD_DATA_ROOT/.venv/bin/activate"
pip install -r requirements.lock # exact pinned versions
pip install -e .
lrd validate # checks pins and source/config agreementrequirements.lock is the resolved lock file; requirements.in is the
human-edited input. The stack is pinned to PyTorch 2.6 / CUDA 12.4.
Edit slurm/cluster.env for your partitions, QoS, GPU type, and node
exclusions. Slurm parses #SBATCH directives before any shell runs, so those
values are also baked into the job scripts; the command line always wins:
sbatch --partition=YOUR_PARTITION --exclude="" slurm/train.sbatch <index.csv>Every job writes to $LRD_DATA_ROOT/logs/, which the #SBATCH --output lines
reference by absolute path. Update those paths once for a new site.
Each stage writes immutable specifications first, then submits an array over them. Specifications are content-addressed, so re-staging is idempotent and a resubmitted run is byte-identical to the original.
# 0. data: build the frozen 50k/1k/256 split and audit contamination
lrd prepare-data
lrd audit-contamination
# 1. full-SFT references
lrd qualify-sft --learning-rate 2e-5 --submit --max-parallel 4
lrd benchmark-gate0-finals --learning-rate 2e-5 --submit
# 2. LoRA rank grid
lrd stage-gate1a --learning-rate 2e-4 --submit --max-parallel 4
lrd benchmark-gate1a --submit
# 3. descriptors (no training; Tier 2 needs only forward passes)
sbatch slurm/extract_tier1.sbatch "qwen25_15b llama32_1b gemma2_2b olmo2_1b"
sbatch slurm/extract_tier2.sbatch "qwen25_15b llama32_1b gemma2_2b olmo2_1b" 4
sbatch slurm/collect_tier1.sbatch
sbatch slurm/collect_tier2.sbatch
# 4. analysis
sbatch slurm/analyze_gate1a.sbatch --evaluation-name gate1a-4k-i017
sbatch --export=ALL,TIER=1 slurm/analyze_tier1.sbatch --tier 1
sbatch --export=ALL,TIER=2 slurm/analyze_tier1.sbatch --tier 2
sbatch slurm/export_tables.sbatch # refresh results/*.csvlrd --help lists every subcommand. Most accept --dry-run to print the
specifications without submitting.
lrd stage-gate1b --learning-rate 2e-4 --submit --max-parallel 2 # seeds 314, 2718
lrd stage-gate1a --learning-rate 1e-4 --specification-name gate1a-lr0.0001 --submitEvery analysis of a non-default panel needs --suffix. The scripts take
--curves to choose an input and --suffix to name the output; passing the
first without the second overwrites the default panel's results in place. See
incident 020.
sbatch slurm/analyze_gate1a.sbatch --evaluation-name gate1a-4k-lr1e4 --suffix=-lr1e4
sbatch slurm/analyze_tier1.sbatch --tier 1 \
--curves "$LRD_DATA_ROOT/results/gate1a_rank_curves-lr1e4.json" --suffix=-lr1e4
# combined analyses over both rates, once both panels exist
sbatch slurm/analyze_two_rate.sbatch
sbatch slurm/analyze_application.sbatchArgparse reads a leading-dash value as a flag, so --suffix=-lr1e4 must use the
= form.
Start with docs/DATA_INDEX.md. It documents every column of every exported table and points to the raw per-example generations for any claim that needs checking against model output rather than a summary statistic.
Every CSV in results/ is readable from a login node with no allocation.
A -lr1e4 suffix means the 1e-4 panel; no suffix means 2e-4, and the
rate-dependent tables also carry it in a curves_source column.
| File | Contents |
|---|---|
outcomes_by_checkpoint{,-lr1e4}.csv |
the headline table: base, full SFT, LoRA at each rank, gap, retained gain, minimum sufficient rank |
generation_health{,-lr1e4}.csv |
median length, cap rate, repetition rate per checkpoint and rank |
associations_tier{1,2}{,-lr1e4}.csv |
every pre-registered statistical test including baselines and the randomised control |
two_rate_outcomes.csv |
best-of-two-rates outcomes, rate sensitivity, rank-rate interaction |
two_rate_associations.csv |
descriptors against those targets — exploratory |
application_targets.csv |
retention and rank-gain targets per checkpoint |
application_associations.csv |
the full 53-feature search — exploratory, and the reason the p values in it mean little |
descriptors_tier{1,2}.csv |
the frozen descriptor summaries, rate-independent |
matrix_profiles_tier{1,2}.csv |
per-matrix substrate for mechanism analysis, rate-independent |
The two _associations tables marked exploratory are exported in full, with no
top-N cut, precisely so the number of tests behind any single row is visible.
These cost real time to discover. All are documented in protocol/incidents/.
A checkpoint's generation config can silently cap output. lm-eval treats
max_new_tokens as an alias for max_gen_toks, not an override, and passes
only a computed max_length to generate. Transformers then lets the
checkpoint's own generation_config.max_new_tokens win. Qwen ships 2048 and
capped two full evaluation rounds before anyone noticed. benchmark.py now
mirrors affected checkpoints into a sanitised link farm. Verify a generation
budget by measuring output length, never by reading back the command.
See incident 017.
Slurm RUNNING is not evidence a run started. Two distinct hardware faults
appear on this cluster: a hang on the first NCCL collective that burns the full
distributed timeout without logging a step, and No CUDA GPUs are available at
CUDA init seconds after launch. Six nodes have been excluded. Check for a logged
loss before trusting a job. See incidents 015, 016, 018.
Wall-clock is not evidence about a model. Two runs timed out at 24 hours and then completed the identical work in 3-7 hours on a different node. Do not infer degeneration from runtime; read the generation-health columns.
Base checkpoints score near zero for a reason that is not a bug. The
benchmark prompt is a Q:/A: document, Q: is a stop sequence, and an
unadapted model continues the document by emitting the next Q:, which
truncates at position zero. The measurement is faithful. Both headline targets
(gap, rank128 - rank8) avoid the base score entirely. See
docs/DATA_INDEX.md section 4b.
Pre-register before looking. protocol/preregistration/ fixes features,
targets, and baselines before the corresponding numbers exist, so it is
verifiable that nothing was chosen after the fact. Three conclusions drawn at
four checkpoints were overturned at six; the pre-registrations are what make
that a correction rather than a rewrite.
Record amendments and incidents as they happen, including wrong diagnoses. Several incidents contain an explicit "this earlier explanation was wrong" section. That is deliberate: the next person should not retry a ruled-out hypothesis.
Update PROGRESS_LOG.md at every milestone, and before any long wait. Its
Resume Here section is the authoritative restart point.
Never poll Slurm in a loop from an interactive session. Use
scripts/watch_slurm_pipeline.sh <jobid>..., which blocks silently and exits
non-zero on any terminal failure.