E5 Large v2
E5 Large v2 is Microsoft Research's large-scale text embedding model from the E5 (EmbEddings from bidirEctional Encoder rEpresentations) family, trained using a weakly-supervised contrastive approach on large text pair datasets followed by fine-tuning on curated retrieval data. It is one of the strongest encoder-based embedding models for English-language retrieval in the base-to-large size class.
Microsoft positions E5 Large v2 as a high-quality general-purpose embedding model for semantic search, question answering retrieval, and NLI tasks. Its dual-encoder architecture makes it particularly well suited for asymmetric retrieval — where short queries are matched against longer document passages — and it has been widely adopted in academic and production RAG systems as a reliable open-weight option.
Key Features
Large encoder architecture delivering strong retrieval and similarity quality
Trained on massive weakly-supervised text pairs for broad domain coverage
Excels at asymmetric query-to-passage retrieval tasks
Competitive on BEIR and MTEB retrieval benchmarks
Open-weight model available on Hugging Face for self-hosting
Supports both symmetric (sentence similarity) and asymmetric (QA) retrieval
Ideal Use Cases
Open-domain question answering retrieval in enterprise search systems
Dense passage retrieval for RAG pipelines over large document sets
Semantic textual similarity for duplicate detection and clustering
Academic literature search and citation recommendation
Information retrieval benchmarking and model comparison
Example Prompts for E5 Large v2
Technical Specifications
| Provider | Microsoft |
| Category | Embedding |
| Modality | Text -> Vector |
Frequently Asked Questions
Try E5 Large v2 now
Start using E5 Large v2 instantly — 100 free credits, no credit card required. Access 750+ AI models through one platform.