DenseOn with the LateOn: Fully Open Dense and Late-Interaction Models for Multilingual, Long-Context, and Code Search
Ranking
Overall
87
Content
95
Popularity
67
Observed public metrics from 1 member.
Merged summary
TL;DR - This paper introduces fully open recipes, data, code, and models for dense and late-interaction retrieval. Late interaction notably generalizes better to languages and scripts excluded from translated training data.
- DenseOn and LateOn are 149M-parameter models trained on 665M contrastive pairs and 1.88M supervised pairs with hard negatives.
- They achieve 56.20 and 57.22 average nDCG@10 on BEIR, respectively, setting reported size-class records.
- Multilingual variants use 2.8B pairs spanning eight translated languages and cross-lingual examples.
- Token-level matching helps mLateOn outperform dense retrieval in generalization to unseen languages and scripts.
Sources (1)
DenseOn with the LateOn: Fully Open Dense and Late-Interaction Models for Multilingual, Long-Context, and Code Search
Public signals
Hugging Face upvotes 6
TL;DR - This paper introduces fully open recipes, data, code, and models for dense and late-interaction retrieval. Late interaction notably generalizes better to languages and scripts excluded from translated training data.
- DenseOn and LateOn are 149M-parameter models trained on 665M contrastive pairs and 1.88M supervised pairs with hard negatives.
- They achieve 56.20 and 57.22 average nDCG@10 on BEIR, respectively, setting reported size-class records.
- Multilingual variants use 2.8B pairs spanning eight translated languages and cross-lingual examples.
- Token-level matching helps mLateOn outperform dense retrieval in generalization to unseen languages and scripts.