BGE-M3, developed by the Beijing Academy of Artificial Intelligence (BAAI), is a multilingual embedding model designed to handle more than 100 languages within a single model. Its key contribution is combining three retrieval methods — dense embedding, sparse lexical retrieval, and multi-vector ColBERT-style retrieval — enabling it to adapt to different retrieval strategies depending on the use case.
BGE-M3 supports long input sequences and excels at cross-lingual retrieval, where queries and documents may be in different languages. It is a practical choice for global search infrastructure, multilingual RAG systems, and any pipeline needing unified retrieval across language boundaries.
Key Features
Hybrid retrieval: dense, sparse (BM25-style), and multi-vector modes
Cross-lingual retrieval across 100+ languages
Long document input support for passage-level embedding
Single model handles multilingual and monolingual workloads
Strong performance on BEIR and MIRACL multilingual benchmarks
Open-weight model available via Hugging Face
Ideal Use Cases
Cross-lingual enterprise search where queries and docs differ in language
Multilingual RAG pipeline retrieval
Hybrid keyword-plus-semantic retrieval for legacy + modern search stacks
Global e-commerce product search across language markets
Academic or scientific search over multilingual paper corpora
Example Prompts for BGE-M3
Technical Specifications
| Provider | BGE |
| Category | Embedding |
| Modality | Text -> Vector |
Frequently Asked Questions
Try BGE-M3 now
Start using BGE-M3 instantly — 100 free credits, no credit card required. Access 750+ AI models through one platform.