Skip to content

Repository files navigation

LoRA Rank Demand

Can a cheap measurement of a pretrained checkpoint, taken before any training, predict how well LoRA will adapt it and how much rank it needs?

If it can, you could decide up front whether to use LoRA at rank 8, LoRA at rank 128, or full fine-tuning, instead of discovering it by running all three.


Status

Six checkpoints, three LoRA ranks each, two learning rates, on mathematical reasoning. Two descriptor tiers extracted and tested against a pre-registered analysis.

The short answer so far is no. Neither descriptor tier predicts rank demand in a way that survives changing the LoRA learning rate. Both pre-registered associations were measured at 2e-4, then failed to replicate at 1e-4.

Pre-registered claim at 2e-4 at 1e-4
frozen-weight geometry predicts the gap to full SFT -0.94, p = 0.017 -0.49, p = 0.356
activation geometry predicts the rank response +0.94, p = 0.017 absent from the top five

At 1e-4 a randomised control ties the best real feature. Amendment 013 committed in advance to reporting this outcome if it occurred.

What does hold

Finding Evidence
Rank demand varies by checkpoint minimum sufficient rank spans 8 to never-reached
Minimum sufficient rank is not stable across rates 8, 8, 128, never x3 at 2e-4; 128, 32, 128, never x3 at 1e-4
Higher rank prefers a lower rate, universally all six checkpoints, no exception; consistent with alpha / r scaling
Minimum sufficient rank is not predictable no feature beats a randomised control at either rate
The MLP is the low-rank bottleneck attention captured 2-3x better at rank 8 in every checkpoint

What the descriptors do predict, on targets defined after the pre-registered analysis failed and therefore exploratory: learning-rate sensitivity (rho = -1.00, 30/30 held-out orderings) and the rank-rate interaction (rho = +1.00). A candidate application — predicting whether LoRA will retain most of full fine-tuning — reaches rho = -0.83 at 27/30. None of this is confirmatory. Read section 6 of the two-rate report, which states how many tests stand behind those numbers, before citing any of them.

What is not established: one seed, one task, six checkpoints. Start with docs/TWO_RATE_FINDINGS_2026-08-16.md, which supersedes the single-rate reports. The earlier docs/CHECKPOINT_ANALYSIS_2026-08-11.md and docs/ADVISOR_BRIEF_2026-08-11.md are accurate for 2e-4 alone and are kept as the record of what was concluded before the second rate existed.


Repository layout

configs/study.toml        the frozen study configuration: models, ranks, rates, thresholds
manifests/                pinned model, benchmark, and gate manifests with commit SHAs
src/lora_rank_demand/     the package; every stage is a subcommand of `lrd`
  data.py                 dataset preparation, contamination audit
  training.py             full SFT and LoRA via TRL/PEFT/DeepSpeed
  benchmark.py            lm-eval harness wrapper, generation-config sanitising
  descriptors.py          Tier 1 frozen-weight descriptors
  activations.py          Tier 2 task-conditioned activation descriptors
  outcomes.py             retained gain, gap, minimum sufficient rank
  runs.py                 immutable run specifications and Slurm submission
scripts/                  analysis entry points, run on a compute node
slurm/                    core pipeline job scripts
  campaigns/              one-off staging jobs for specific experiment blocks
  diagnostics/            failure triage and audit jobs, not part of the pipeline
  cluster.env             site-specific partitions, QoS, node exclusions
protocol/
  preregistration/        features, targets, and baselines fixed BEFORE outcomes existed
  amendments/             numbered protocol changes, each recorded before execution
  incidents/              every failure and its repair, including wrong diagnoses
docs/                     analysis reports and the data index
results/*.csv             every derived table, readable without a compute allocation
tests/                    unit suite, 43 tests
RESEARCH_PLAN.md          the scientific protocol and gate definitions
PROGRESS_LOG.md           chronological execution record; "Resume Here" is the restart point

Everything generated goes under $LRD_DATA_ROOT, never into the source tree. The repository holds source, manifests, compact summaries, and checksums only.


Setup

1. Point at your filesystems

scripts/project_env.sh derives the source root from its own location and takes everything else from the environment, so the only required decision is where generated data lives:

export LRD_DATA_ROOT=/your/data/path/lora-rank-demand    # default: /data/user_data/$USER/lora-rank-demand
source scripts/project_env.sh

Optional overrides: HF_TOKEN_PATH (default ~/.huggingface/token), LRD_SCRATCH (default /scratch, used for the per-job Triton cache).

2. Create the environment

python -m venv "$LRD_DATA_ROOT/.venv"
source "$LRD_DATA_ROOT/.venv/bin/activate"
pip install -r requirements.lock          # exact pinned versions
pip install -e .
lrd validate                              # checks pins and source/config agreement

requirements.lock is the resolved lock file; requirements.in is the human-edited input. The stack is pinned to PyTorch 2.6 / CUDA 12.4.

3. Adjust cluster settings

Edit slurm/cluster.env for your partitions, QoS, GPU type, and node exclusions. Slurm parses #SBATCH directives before any shell runs, so those values are also baked into the job scripts; the command line always wins:

sbatch --partition=YOUR_PARTITION --exclude="" slurm/train.sbatch <index.csv>

Every job writes to $LRD_DATA_ROOT/logs/, which the #SBATCH --output lines reference by absolute path. Update those paths once for a new site.


Running the pipeline

Each stage writes immutable specifications first, then submits an array over them. Specifications are content-addressed, so re-staging is idempotent and a resubmitted run is byte-identical to the original.

# 0. data: build the frozen 50k/1k/256 split and audit contamination
lrd prepare-data
lrd audit-contamination

# 1. full-SFT references
lrd qualify-sft --learning-rate 2e-5 --submit --max-parallel 4
lrd benchmark-gate0-finals --learning-rate 2e-5 --submit

# 2. LoRA rank grid
lrd stage-gate1a --learning-rate 2e-4 --submit --max-parallel 4
lrd benchmark-gate1a --submit

# 3. descriptors (no training; Tier 2 needs only forward passes)
sbatch slurm/extract_tier1.sbatch "qwen25_15b llama32_1b gemma2_2b olmo2_1b"
sbatch slurm/extract_tier2.sbatch "qwen25_15b llama32_1b gemma2_2b olmo2_1b" 4
sbatch slurm/collect_tier1.sbatch
sbatch slurm/collect_tier2.sbatch

# 4. analysis
sbatch slurm/analyze_gate1a.sbatch --evaluation-name gate1a-4k-i017
sbatch --export=ALL,TIER=1 slurm/analyze_tier1.sbatch --tier 1
sbatch --export=ALL,TIER=2 slurm/analyze_tier1.sbatch --tier 2
sbatch slurm/export_tables.sbatch          # refresh results/*.csv

lrd --help lists every subcommand. Most accept --dry-run to print the specifications without submitting.

Reliability replicates and a second rate

lrd stage-gate1b --learning-rate 2e-4 --submit --max-parallel 2   # seeds 314, 2718
lrd stage-gate1a --learning-rate 1e-4 --specification-name gate1a-lr0.0001 --submit

Every analysis of a non-default panel needs --suffix. The scripts take --curves to choose an input and --suffix to name the output; passing the first without the second overwrites the default panel's results in place. See incident 020.

sbatch slurm/analyze_gate1a.sbatch --evaluation-name gate1a-4k-lr1e4 --suffix=-lr1e4
sbatch slurm/analyze_tier1.sbatch --tier 1 \
    --curves "$LRD_DATA_ROOT/results/gate1a_rank_curves-lr1e4.json" --suffix=-lr1e4

# combined analyses over both rates, once both panels exist
sbatch slurm/analyze_two_rate.sbatch
sbatch slurm/analyze_application.sbatch

Argparse reads a leading-dash value as a flag, so --suffix=-lr1e4 must use the = form.


Reading the results

Start with docs/DATA_INDEX.md. It documents every column of every exported table and points to the raw per-example generations for any claim that needs checking against model output rather than a summary statistic.

Every CSV in results/ is readable from a login node with no allocation. A -lr1e4 suffix means the 1e-4 panel; no suffix means 2e-4, and the rate-dependent tables also carry it in a curves_source column.

File Contents
outcomes_by_checkpoint{,-lr1e4}.csv the headline table: base, full SFT, LoRA at each rank, gap, retained gain, minimum sufficient rank
generation_health{,-lr1e4}.csv median length, cap rate, repetition rate per checkpoint and rank
associations_tier{1,2}{,-lr1e4}.csv every pre-registered statistical test including baselines and the randomised control
two_rate_outcomes.csv best-of-two-rates outcomes, rate sensitivity, rank-rate interaction
two_rate_associations.csv descriptors against those targets — exploratory
application_targets.csv retention and rank-gain targets per checkpoint
application_associations.csv the full 53-feature search — exploratory, and the reason the p values in it mean little
descriptors_tier{1,2}.csv the frozen descriptor summaries, rate-independent
matrix_profiles_tier{1,2}.csv per-matrix substrate for mechanism analysis, rate-independent

The two _associations tables marked exploratory are exported in full, with no top-N cut, precisely so the number of tests behind any single row is visible.


Things that will bite you

These cost real time to discover. All are documented in protocol/incidents/.

A checkpoint's generation config can silently cap output. lm-eval treats max_new_tokens as an alias for max_gen_toks, not an override, and passes only a computed max_length to generate. Transformers then lets the checkpoint's own generation_config.max_new_tokens win. Qwen ships 2048 and capped two full evaluation rounds before anyone noticed. benchmark.py now mirrors affected checkpoints into a sanitised link farm. Verify a generation budget by measuring output length, never by reading back the command. See incident 017.

Slurm RUNNING is not evidence a run started. Two distinct hardware faults appear on this cluster: a hang on the first NCCL collective that burns the full distributed timeout without logging a step, and No CUDA GPUs are available at CUDA init seconds after launch. Six nodes have been excluded. Check for a logged loss before trusting a job. See incidents 015, 016, 018.

Wall-clock is not evidence about a model. Two runs timed out at 24 hours and then completed the identical work in 3-7 hours on a different node. Do not infer degeneration from runtime; read the generation-health columns.

Base checkpoints score near zero for a reason that is not a bug. The benchmark prompt is a Q:/A: document, Q: is a stop sequence, and an unadapted model continues the document by emitting the next Q:, which truncates at position zero. The measurement is faithful. Both headline targets (gap, rank128 - rank8) avoid the base score entirely. See docs/DATA_INDEX.md section 4b.


Working conventions

Pre-register before looking. protocol/preregistration/ fixes features, targets, and baselines before the corresponding numbers exist, so it is verifiable that nothing was chosen after the fact. Three conclusions drawn at four checkpoints were overturned at six; the pre-registrations are what make that a correction rather than a rewrite.

Record amendments and incidents as they happen, including wrong diagnoses. Several incidents contain an explicit "this earlier explanation was wrong" section. That is deliberate: the next person should not retry a ruled-out hypothesis.

Update PROGRESS_LOG.md at every milestone, and before any long wait. Its Resume Here section is the authoritative restart point.

Never poll Slurm in a loop from an interactive session. Use scripts/watch_slurm_pipeline.sh <jobid>..., which blocks silently and exits non-zero on any terminal failure.

About

Can cheap pre-training descriptors predict LoRA rank demand? A pre-registered six-checkpoint study on mathematical reasoning.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages