Skip to content

Add LEMUR fixed dimensional encoding - #1164

Merged
davidmezzetti merged 6 commits into
neuml:masterfrom
morgan-coded:feat/lemur
Aug 1, 2026
Merged

davidmezzetti merged 6 commits into
neuml:masterfrom
morgan-coded:feat/lemur

Conversation

@morgan-coded

Copy link
Copy Markdown
Contributor

Adds a LEMUR fixed dimensional encoder for late-interaction models and a separate LemurTrainer pipeline. It uses the requested trainer, artifact, and pooling structure: training writes config.json and model.safetensors, and the pooling module loads either a local directory or Hugging Face Hub repository through the standard loader. The existing late-pooling scaffolding handles LEMUR alongside MUVERA, and lateencoder.py disables muvera and lemur through one encoder dictionary while producing raw multi-vectors.

Benchmark

Measured on one machine with one GPU using colbert-ir/colbertv2.0, torch 2.13.0+cu130, and exact IDMap,Flat. Each cell is NDCG@10 / fixed-vector index KB.

Dataset MUVERA 10,240 MUVERA 2,048 LEMUR ELM 2,048 LEMUR MLP 2,048
nfcorpus 0.23544 / 145,380 0.16299 / 29,124 0.21591 / 29,124 0.25524 / 29,124
scifact 0.50021 / 207,404 0.36757 / 41,548 0.50251 / 41,548 0.54910 / 41,548
arguana 0.34614 / 347,337 0.26280 / 69,769 0.34207 / 69,769 0.42556 / 69,769

At the equal 2,048-dimensional vector budget, LEMUR-MLP improves NDCG@10 over matched-size MUVERA by +56.6% on nfcorpus, +49.4% on scifact, and +61.9% on arguana. Against default 10,240-dimensional MUVERA, the gains are +8.4%, +9.8%, and +22.9% with one-fifth the fixed-vector index size. The same ordering holds on MAP@10, Recall@10, and P@10 for both comparisons on all three datasets. The untrained ELM path needs no training and beats matched-size MUVERA on all three; on scifact and arguana it lands within noise of full-size MUVERA at one-fifth the index size.

This covers three of the five datasets in the benchmark set: arguana, nfcorpus, and scifact. Scidocs and fiqa remain to be run because of a benchmark-machine hardware fault.

Limitation: default IVF above 5,000 rows

These measurements use txtai's normalized single-stage path; fixed vectors are L2-normalized and dense results with non-positive scores are filtered. All values above pin exact search. txtai selects exact IDMap,Flat through 5,000 rows and IVF above that threshold. On scifact, the measured change under default IVF is roughly -43% NDCG@10 for LEMUR versus -25% for MUVERA. LEMUR is materially more sensitive to coarse quantization, so a corpus above 5,000 rows should pin an exact index or tune IVF.

Training and API

The learn distribution was the main training result. In the nfcorpus CPU ablation, encoding learn tokens with the document encoder produced a trained MLP at 0.15870 NDCG@10, below the untrained ELM at 0.19187. The query-encoded path is the one used in the table above, where trained LEMUR is +49.4% to +61.9% over matched-size MUVERA, so learn_category defaults to "query".

epochs has no silent default: callers pass 100 for trained MLP quality or 0 for deterministic ELM features. corpus_subset_size bounds encoding for large corpora by sampling raw corpus texts before either encoding pass, and validation_split enables validation-based epoch selection. Artifacts use config.json and model.safetensors through the standard Hub/local loader.

Compatibility and scope

MUVERA is unchanged. The muvera.py checksum matches the released source, LEMUR-only handling is gated off its path, and saved MUVERA reference encodings remain bit-identical.

A corpus-independent LEMUR model is not addressed here. The paper treats the reduction as corpus-specific, and I have not tested cross-corpus generalization; a general-model experiment is a natural follow-up.

Based on the LEMUR paper and its MIT reference implementation.

Closes #1024

@davidmezzetti

Copy link
Copy Markdown
Member

At first glance this is very impressive! Would you mind testing on a more modern model such as https://huggingface.co/lightonai/LateOn? MUVERA had some issues with other ColBERT models and I'm curious if LEMUR would also have the same issues.

For reference see this conversation: lightonai/pylate#142

Additionally, if you wanted to get rid of the IVF concerns, just use backend: numpy. I often do that to eliminate ANN when testing different vectorization methods. The BEIR datasets are all relatively small.

@davidmezzetti

Copy link
Copy Markdown
Member

Can we also make sure there is test coverage for all the new code: https://coveralls.io/github/neuml/txtai?branch=feat/lemur

@morgan-coded

Copy link
Copy Markdown
Contributor Author

Added the remaining coverage cases; both LEMUR modules now report 100% statement coverage. Pushed as 626cc8fb — the workflow runs there are awaiting approval.

Tested lightonai/LateOn across the same three datasets with both encoders. Both collapsed, with MUVERA generally worse. NDCG@10:

Dataset LEMUR-2,048 ColBERTv2 LEMUR-2,048 LateOn MUVERA-10,240 ColBERTv2 MUVERA-10,240 LateOn
nfcorpus 0.25524 0.00000 0.23544 0.03639
scifact 0.54910 0.04985 0.50021 0.00269
arguana 0.42556 0.18372 0.34614 0.00053

Through LatePooling, LateOn's per-query-token maxsim spread across documents was 0.018 versus 0.370 for ColBERTv2 (~20x lower), while within-document token variance was 0.000186 versus 0.004135. Nearly every document scored ~0.90 against any query token, leaving almost no signal for a fixed-dimensional reduction to preserve. LEMUR's training loss fell to ~0.0025 from 0.0465, consistent with near-constant targets. Token norms were 1.0 for both and shapes were comparable, ruling out normalization and padding.

The cause is unresolved: this may be LateOn's genuine behavior, or LatePooling may be mishandling the newer PyLate format through its dense layer or prefixes.

Thanks for the backend: numpy pointer — all LateOn numbers above used it, so they carry no ANN or IVF caveat.

@davidmezzetti

davidmezzetti commented Jul 31, 2026 •

Copy link
Copy Markdown
Member

Ok, I'll run the tests now.

Regarding accuracy, what about using this model: https://huggingface.co/lightonai/LateOn-hpool-regularized

I recently added changes to late pooling to incorporate multiple dense layers: #1117

Edit: Or this model https://huggingface.co/lightonai/LateOn-regularized

@davidmezzetti davidmezzetti added this to the v9.13.0 milestone Jul 31, 2026
@davidmezzetti

Copy link
Copy Markdown
Member

Some more things to research in this blog: https://huggingface.co/blog/lightonai/lateon-regularization

I'm not sure I've handled mean centering which could be part of the issue in using that regularized model. Perhaps @raphaelsty and/or @NohTow would be willing to share more information on that part.

@morgan-coded

Copy link
Copy Markdown
Contributor Author

Confirmed the centering hunch: on scifact, base LateOn with LEMUR moved from 0.04985 to 0.68624 NDCG@10.

The geometry moved with it:

model cosine raw → centered maxsim spread raw → centered
LateOn 0.9508 → 0.0033 0.0559 → 0.4772
LateOn-regularized 0.9800 → 0.0006 0.0127 → 0.3854
LateOn-hpool-regularized 0.9581 → 0.0007 0.0400 → 0.5580

These are NDCG@10 results on nfcorpus and scifact with backend: numpy exact search; arguana was not run in this pass.

model dataset LEMUR-2048 off → on MUVERA-10240 off → on
LateOn nfcorpus 0.0 → 0.31473 0.03639 → 0.13369
LateOn scifact 0.04985 → 0.68624 0.00269 → 0.29293
LateOn-regularized nfcorpus 0.00262 → 0.08015 0.00973 → 0.02059
LateOn-regularized scifact 0.02108 → 0.12203 0.0 → 0.01573
LateOn-hpool-regularized nfcorpus 0.00063 → 0.33584 0.02778 → 0.20531
LateOn-hpool-regularized scifact 0.11747 → 0.717 0.00382 → 0.40285

Centered hpool-regularized was the strongest model on both datasets, and LEMUR-2048 beat MUVERA-10240 in every centered cell. Plain LateOn-regularized was the exception: even centered, LEMUR reached only 0.08015 on nfcorpus and 0.12203 on scifact. That is a separate unresolved handling issue, not a centering failure.

The recipe was the collection mean over real document-token rows from LatePooling; I subtracted it from document and query tokens, re-L2-normalized, and left padding rows at zero. The center modelargs hook ran after per-token L2 normalization and before the MUVERA/LEMUR transform, beside the existing muvera hook. The cost is a corpus pass to compute μ before indexing.

If that shape fits, I can add center support to LatePooling in this PR or keep it as a follow-up—your call.

@davidmezzetti

davidmezzetti commented Aug 1, 2026 •

Copy link
Copy Markdown
Member

Phenomenal! Let's merge this in and have a separate PR for the centering as that's probably cleaner.

Please make that configurable (i.e. center param). I briefly played around with centering logic and it seemed to hurt some models while helping others. Perhaps a heuristic is to default center to True when there are more than 1 linear layer?

I also saw some centering algorithms was on a per document or per document batch basis?

@davidmezzetti
davidmezzetti merged commit 2ef401b into neuml:master Aug 1, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Feature request: Add LEMUR: Learned Multi-Vector Retrieval

2 participants