Add LEMUR fixed dimensional encoding - #1164
Conversation
|
At first glance this is very impressive! Would you mind testing on a more modern model such as https://huggingface.co/lightonai/LateOn? MUVERA had some issues with other ColBERT models and I'm curious if LEMUR would also have the same issues. For reference see this conversation: lightonai/pylate#142 Additionally, if you wanted to get rid of the IVF concerns, just use |
|
Can we also make sure there is test coverage for all the new code: https://coveralls.io/github/neuml/txtai?branch=feat/lemur |
|
Added the remaining coverage cases; both LEMUR modules now report 100% statement coverage. Pushed as Tested
Through The cause is unresolved: this may be LateOn's genuine behavior, or Thanks for the |
|
Ok, I'll run the tests now. Regarding accuracy, what about using this model: https://huggingface.co/lightonai/LateOn-hpool-regularized I recently added changes to late pooling to incorporate multiple dense layers: #1117 Edit: Or this model https://huggingface.co/lightonai/LateOn-regularized |
|
Some more things to research in this blog: https://huggingface.co/blog/lightonai/lateon-regularization I'm not sure I've handled mean centering which could be part of the issue in using that regularized model. Perhaps @raphaelsty and/or @NohTow would be willing to share more information on that part. |
|
Confirmed the centering hunch: on scifact, base LateOn with LEMUR moved from The geometry moved with it:
These are NDCG@10 results on nfcorpus and scifact with
Centered hpool-regularized was the strongest model on both datasets, and LEMUR-2048 beat MUVERA-10240 in every centered cell. Plain The recipe was the collection mean over real document-token rows from If that shape fits, I can add |
|
Phenomenal! Let's merge this in and have a separate PR for the centering as that's probably cleaner. Please make that configurable (i.e. center param). I briefly played around with centering logic and it seemed to hurt some models while helping others. Perhaps a heuristic is to default I also saw some centering algorithms was on a per document or per document batch basis? |
Adds a LEMUR fixed dimensional encoder for late-interaction models and a separate
LemurTrainerpipeline. It uses the requested trainer, artifact, and pooling structure: training writesconfig.jsonandmodel.safetensors, and the pooling module loads either a local directory or Hugging Face Hub repository through the standard loader. The existing late-pooling scaffolding handles LEMUR alongside MUVERA, andlateencoder.pydisablesmuveraandlemurthrough one encoder dictionary while producing raw multi-vectors.Benchmark
Measured on one machine with one GPU using
colbert-ir/colbertv2.0, torch 2.13.0+cu130, and exactIDMap,Flat. Each cell is NDCG@10 / fixed-vector index KB.At the equal 2,048-dimensional vector budget, LEMUR-MLP improves NDCG@10 over matched-size MUVERA by +56.6% on nfcorpus, +49.4% on scifact, and +61.9% on arguana. Against default 10,240-dimensional MUVERA, the gains are +8.4%, +9.8%, and +22.9% with one-fifth the fixed-vector index size. The same ordering holds on MAP@10, Recall@10, and P@10 for both comparisons on all three datasets. The untrained ELM path needs no training and beats matched-size MUVERA on all three; on scifact and arguana it lands within noise of full-size MUVERA at one-fifth the index size.
This covers three of the five datasets in the benchmark set: arguana, nfcorpus, and scifact. Scidocs and fiqa remain to be run because of a benchmark-machine hardware fault.
Limitation: default IVF above 5,000 rows
These measurements use txtai's normalized single-stage path; fixed vectors are L2-normalized and dense results with non-positive scores are filtered. All values above pin exact search. txtai selects exact
IDMap,Flatthrough 5,000 rows and IVF above that threshold. On scifact, the measured change under default IVF is roughly -43% NDCG@10 for LEMUR versus -25% for MUVERA. LEMUR is materially more sensitive to coarse quantization, so a corpus above 5,000 rows should pin an exact index or tune IVF.Training and API
The learn distribution was the main training result. In the nfcorpus CPU ablation, encoding learn tokens with the document encoder produced a trained MLP at 0.15870 NDCG@10, below the untrained ELM at 0.19187. The query-encoded path is the one used in the table above, where trained LEMUR is +49.4% to +61.9% over matched-size MUVERA, so
learn_categorydefaults to"query".epochshas no silent default: callers pass100for trained MLP quality or0for deterministic ELM features.corpus_subset_sizebounds encoding for large corpora by sampling raw corpus texts before either encoding pass, andvalidation_splitenables validation-based epoch selection. Artifacts useconfig.jsonandmodel.safetensorsthrough the standard Hub/local loader.Compatibility and scope
MUVERA is unchanged. The
muvera.pychecksum matches the released source, LEMUR-only handling is gated off its path, and saved MUVERA reference encodings remain bit-identical.A corpus-independent LEMUR model is not addressed here. The paper treats the reduction as corpus-specific, and I have not tested cross-corpus generalization; a general-model experiment is a natural follow-up.
Based on the LEMUR paper and its MIT reference implementation.
Closes #1024