Indic
MTEB: Indic est un benchmark public consacré à la qualité des embeddings de texte dans les langues indiennes. Publié en 2025 par MTEB / MMTEB, dans le cadre d’embeddings-benchmark et des travaux de Kenneth Enevoldsen et al., il couvre plusieurs formes d’évaluation multilingue.
MTEB: Indic est un benchmark public consacré à la qualité des embeddings de texte dans les langues indiennes. Publié en 2025 par MTEB / MMTEB, dans le cadre d’embeddings-benchmark et des travaux de Kenneth Enevoldsen et al., il couvre plusieurs formes d’évaluation multilingue.
Il mesure la capacité des modèles à représenter et rapprocher des textes pour le bitext mining, la classification, le clustering, la pair classification, le retrieval, le reranking et la similarité sémantique. Il sert ainsi à comparer les performances des modèles sur un ensemble diversifié de tâches liées aux embeddings.
Carte d'identité
| Caractéristique | Valeur |
|---|---|
| Éditeur du benchmark | MTEB / MMTEB (embeddings-benchmark, Kenneth Enevoldsen et al.) |
| Capacités mesurées | Qualité des embeddings de texte à travers les langues indiennes : bitext mining, classification, clustering, pair classification, retrieval, reranking et similarité sémantique. |
| Modalité | Texte |
| Type de questions | Évaluation d'embeddings de texte multilingues (bitext mining, classification, clustering, pair classification, retrieval, reranking, STS) |
| Métrique d'évaluation | Score dépendant de la tâche (nDCG@10, accuracy, V-measure, Spearman), agrégé |
| Accès | Public |
| Langues | Langues indiennes (indic) |
| Année de publication | 2025 |
| Ressources | Site / dépôt officiel · Article scientifique |
Classement des modèles (118)
| # | Modèle | Éditeur | Licence | Score | Sortie | Fiabilité |
|---|---|---|---|---|---|---|
| 1 | codefuse-ai/F2LLM-v2-14B | codefuse-ai | 🟢 Ouvert | 78,8 % | 9 mars 2026 | ✅ Mesuré |
| 2 | codefuse-ai/F2LLM-v2-8B | codefuse-ai | 🟢 Ouvert | 77,9 % | 9 mars 2026 | ✅ Mesuré |
| 3 | codefuse-ai/F2LLM-v2-4B | codefuse-ai | 🟢 Ouvert | 76,6 % | 9 mars 2026 | ✅ Mesuré |
| 4 | codefuse-ai/F2LLM-v2-1.7B | codefuse-ai | 🟢 Ouvert | 74,2 % | 9 mars 2026 | ✅ Mesuré |
| 5 | LingoIITGN/qwen-indic-v1 | LingoIITGN | 🟢 Ouvert | 73,8 % | 3 juillet 2026 | ✅ Mesuré |
| 6 | intfloat/multilingual-e5-large-instruct | Intfloat | 🟢 Ouvert | 70,2 % | 8 février 2024 | ✅ Mesuré |
| 7 | codefuse-ai/F2LLM-v2-0.6B | codefuse-ai | 🟢 Ouvert | 70,1 % | 9 mars 2026 | ✅ Mesuré |
| 8 | Alibaba-NLP/gte-Qwen2-7B-instruct | Alibaba | 🟢 Ouvert | 69,4 % | 15 juin 2024 | ✅ Mesuré |
| 9 | Cohere/Cohere-embed-multilingual-v3.0 | Cohere | 🟢 Ouvert | 68,7 % | 2 novembre 2023 | ✅ Mesuré |
| 10 | OrdalieTech/Solon-embeddings-large-0.1 | OrdalieTech | 🟢 Ouvert | 68,4 % | 9 décembre 2023 | ✅ Mesuré |
| 11 | Lajavaness/bilingual-embedding-large | Lajavaness | 🟢 Ouvert | 67,4 % | 24 juin 2024 | ✅ Mesuré |
| 12 | codefuse-ai/F2LLM-v2-330M | codefuse-ai | 🟢 Ouvert | 66,9 % | 9 mars 2026 | ✅ Mesuré |
| 13 | Lajavaness/bilingual-embedding-base | Lajavaness | 🟢 Ouvert | 66,6 % | 26 juin 2024 | ✅ Mesuré |
| 14 | Cohere/Cohere-embed-multilingual-light-v3.0 | Cohere | 🟢 Ouvert | 66,1 % | 2 novembre 2023 | ✅ Mesuré |
| 15 | Snowflake/snowflake-arctic-embed-l-v2.0 | Snowflake | 🟢 Ouvert | 65,9 % | 4 décembre 2024 | ✅ Mesuré |
| 16 | Lajavaness/bilingual-embedding-small | Lajavaness | 🟢 Ouvert | 65,0 % | 17 juillet 2024 | ✅ Mesuré |
| 17 | intfloat/multilingual-e5-base | Intfloat | 🟢 Ouvert | 64,6 % | 8 février 2024 | ✅ Mesuré |
| 18 | omarelshehy/arabic-english-sts-matryoshka | omarelshehy | 🟢 Ouvert | 63,6 % | 13 octobre 2024 | ✅ Mesuré |
| 19 | voyageai/voyage-3-lite | Voyage AI | 🟢 Ouvert | 63,2 % | 18 septembre 2024 | ✅ Mesuré |
| 20 | Omartificial-Intelligence-Space/Arabic-labse-Matryoshka | Omartificial-Intelligence-Space | 🟢 Ouvert | 63,0 % | 16 juin 2024 | ✅ Mesuré |
| 21 | codefuse-ai/F2LLM-v2-160M | codefuse-ai | 🟢 Ouvert | 62,1 % | 9 mars 2026 | ✅ Mesuré |
| 22 | Linq-AI-Research/Linq-Embed-Mistral | Linq-AI-Research | 🟢 Ouvert | 61,9 % | 29 mai 2024 | ✅ Mesuré |
| 23 | sentence-transformers/LaBSE | Sentence Transformers | 🟢 Ouvert | 61,7 % | 1 novembre 2019 | ✅ Mesuré |
| 24 | Alibaba-NLP/gte-Qwen2-1.5B-instruct | Alibaba | 🟢 Ouvert | 60,4 % | 29 juillet 2024 | ✅ Mesuré |
| 25 | GritLM/GritLM-7B | GritLM | 🟢 Ouvert | 60,2 % | 15 février 2024 | ✅ Mesuré |
| 26 | ibm-granite/granite-embedding-278m-multilingual | IBM | 🟢 Ouvert | 60,0 % | 18 décembre 2024 | ✅ Mesuré |
| 27 | intfloat/e5-mistral-7b-instruct | Intfloat | 🟢 Ouvert | 60,0 % | 8 février 2024 | ✅ Mesuré |
| 28 | Salesforce/SFR-Embedding-2_R | Salesforce | 🟢 Ouvert | 59,8 % | 14 juin 2024 | ✅ Mesuré |
| 29 | Alibaba-NLP/gte-Qwen1.5-7B-instruct | Alibaba | 🟢 Ouvert | 59,6 % | 20 avril 2024 | ✅ Mesuré |
| 30 | Salesforce/SFR-Embedding-Mistral | Salesforce | 🟢 Ouvert | 59,4 % | 24 janvier 2024 | ✅ Mesuré |
| 31 | Snowflake/snowflake-arctic-embed-m-v2.0 | Snowflake | 🟢 Ouvert | 59,4 % | 4 décembre 2024 | ✅ Mesuré |
| 32 | Omartificial-Intelligence-Space/Arabic-all-nli-triplet-Matryoshka | Omartificial-Intelligence-Space | 🟢 Ouvert | 59,0 % | 14 juin 2024 | ✅ Mesuré |
| 33 | HIT-TMG/KaLM-embedding-multilingual-mini-v1 | HIT-TMG | 🟢 Ouvert | 58,6 % | 27 août 2024 | ✅ Mesuré |
| 34 | sentence-transformers/paraphrase-multilingual-mpnet-base-v2 | Sentence Transformers | 🟢 Ouvert | 58,5 % | 1 novembre 2019 | ✅ Mesuré |
| 35 | codefuse-ai/F2LLM-v2-80M | codefuse-ai | 🟢 Ouvert | 58,4 % | 9 mars 2026 | ✅ Mesuré |
| 36 | Gameselo/STS-multilingual-mpnet-base-v2 | Gameselo | 🟢 Ouvert | 57,8 % | 7 juin 2024 | ✅ Mesuré |
| 37 | HIT-TMG/KaLM-embedding-multilingual-mini-instruct-v1 | HIT-TMG | 🟢 Ouvert | 57,2 % | 23 octobre 2024 | ✅ Mesuré |
| 38 | ibm-granite/granite-embedding-107m-multilingual | IBM | 🟢 Ouvert | 56,8 % | 18 décembre 2024 | ✅ Mesuré |
| 39 | nvidia/NV-Embed-v2 | NVIDIA | 🟢 Ouvert | 56,3 % | 9 septembre 2024 | ✅ Mesuré |
| 40 | Omartificial-Intelligence-Space/Arabic-MiniLM-L12-v2-all-nli-triplet | Omartificial-Intelligence-Space | 🟢 Ouvert | 54,3 % | 25 juin 2024 | ✅ Mesuré |
| 41 | nvidia/NV-Embed-v1 | NVIDIA | 🟢 Ouvert | 52,8 % | 13 septembre 2024 | ✅ Mesuré |
| 42 | NovaSearch/stella_en_1.5B_v5 | NovaSearch | 🟢 Ouvert | 52,3 % | 12 juillet 2024 | ✅ Mesuré |
| 43 | minishlab/potion-multilingual-128M | minishlab | 🟢 Ouvert | 51,6 % | 23 mai 2025 | ✅ Mesuré |
| 44 | sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2 | Sentence Transformers | 🟢 Ouvert | 49,8 % | 1 novembre 2019 | ✅ Mesuré |
| 45 | Haon-Chen/speed-embedding-7b-instruct | Haon-Chen | 🟢 Ouvert | 47,5 % | 31 octobre 2024 | ✅ Mesuré |
| 46 | Sailesh97/Hinvec | Sailesh97 | 🟢 Ouvert | 46,9 % | 19 juin 2025 | ✅ Mesuré |
| 47 | izhx/udever-bloom-3b | izhx | 🟢 Ouvert | 46,8 % | 24 octobre 2023 | ✅ Mesuré |
| 48 | bigscience/sgpt-bloom-7b1-msmarco | bigscience | 🟢 Ouvert | 46,2 % | 26 août 2022 | ✅ Mesuré |
| 49 | izhx/udever-bloom-7b1 | izhx | 🟢 Ouvert | 45,6 % | 24 octobre 2023 | ✅ Mesuré |
| 50 | sentence-transformers/static-similarity-mrl-multilingual-v1 | Sentence Transformers | 🟢 Ouvert | 44,3 % | 15 janvier 2025 | ✅ Mesuré |
| 51 | Jaume/gemma-2b-embeddings | Jaume | 🟢 Ouvert | 44,3 % | 29 juin 2024 | ✅ Mesuré |
| 52 | deepfile/embedder-100p | deepfile | 🟢 Ouvert | 43,6 % | 24 juillet 2023 | ✅ Mesuré |
| 53 | Cohere/Cohere-embed-english-v3.0 | Cohere | 🟢 Ouvert | 37,2 % | 2 novembre 2023 | ✅ Mesuré |
| 54 | NovaSearch/stella_en_400M_v5 | NovaSearch | 🟢 Ouvert | 37,1 % | 12 juillet 2024 | ✅ Mesuré |
| 55 | sdadas/mmlw-e5-base | sdadas | 🟢 Ouvert | 36,9 % | 17 novembre 2023 | ✅ Mesuré |
| 56 | Intfloat: E5-Large-v2 | Intfloat | 🟢 Ouvert | 36,8 % | 18 novembre 2025 | ✅ Mesuré |
| 57 | manu/sentence_croissant_alpha_v0.2 | manu | 🟢 Ouvert | 36,6 % | 15 mars 2024 | ✅ Mesuré |
| 58 | sdadas/mmlw-e5-large | sdadas | 🟢 Ouvert | 36,6 % | 17 novembre 2023 | ✅ Mesuré |
| 59 | Intfloat: E5-Base-v2 | Intfloat | 🟢 Ouvert | 36,3 % | 18 novembre 2025 | ✅ Mesuré |
| 60 | manu/sentence_croissant_alpha_v0.3 | manu | 🟢 Ouvert | 36,3 % | 26 avril 2024 | ✅ Mesuré |
| 61 | nomic-ai/nomic-embed-text-v1-unsupervised | Nomic AI | 🟢 Ouvert | 36,3 % | 15 janvier 2024 | ✅ Mesuré |
| 62 | manu/sentence_croissant_alpha_v0.4 | manu | 🟢 Ouvert | 36,2 % | 27 avril 2024 | ✅ Mesuré |
| 63 | nomic-ai/nomic-embed-text-v1 | Nomic AI | 🟢 Ouvert | 36,2 % | 31 janvier 2024 | ✅ Mesuré |
| 64 | sdadas/mmlw-e5-small | sdadas | 🟢 Ouvert | 36,2 % | 17 novembre 2023 | ✅ Mesuré |
| 65 | nomic-ai/nomic-embed-text-v1.5 | Nomic AI | 🟢 Ouvert | 35,6 % | 10 février 2024 | ✅ Mesuré |
| 66 | Cohere/Cohere-embed-english-light-v3.0 | Cohere | 🟢 Ouvert | 35,0 % | 2 novembre 2023 | ✅ Mesuré |
| 67 | dwzhu/e5-base-4k | dwzhu | 🟢 Ouvert | 34,6 % | 28 mars 2024 | ✅ Mesuré |
| 68 | Omartificial-Intelligence-Space/Arabic-mpnet-base-all-nli-triplet | Omartificial-Intelligence-Space | 🟢 Ouvert | 34,5 % | 15 juin 2024 | ✅ Mesuré |
| 69 | intfloat/e5-small-v2 | Intfloat | 🟢 Ouvert | 34,2 % | 8 février 2024 | ✅ Mesuré |
| 70 | ibm-granite/granite-embedding-125m-english | IBM | 🟢 Ouvert | 34,2 % | 18 décembre 2024 | ✅ Mesuré |
| 71 | Snowflake/snowflake-arctic-embed-m-v1.5 | Snowflake | 🟢 Ouvert | 34,1 % | 8 juillet 2024 | ✅ Mesuré |
| 72 | avsolatorio/GIST-large-Embedding-v0 | avsolatorio | 🟢 Ouvert | 34,0 % | 14 février 2024 | ✅ Mesuré |
| 73 | infgrad/stella-base-en-v2 | infgrad | 🟢 Ouvert | 33,9 % | 19 octobre 2023 | ✅ Mesuré |
| 74 | BAAI: bge-base-en-v1.5 | BAAI | 🟢 Ouvert | 33,9 % | 18 novembre 2025 | ✅ Mesuré |
| 75 | Snowflake/snowflake-arctic-embed-m | Snowflake | 🟢 Ouvert | 33,6 % | 12 avril 2024 | ✅ Mesuré |
| 76 | WhereIsAI/UAE-Large-V1 | WhereIsAI | 🟢 Ouvert | 33,5 % | 4 décembre 2023 | ✅ Mesuré |
| 77 | intfloat/e5-large | Intfloat | 🟢 Ouvert | 33,4 % | 26 décembre 2022 | ✅ Mesuré |
| 78 | avsolatorio/GIST-Embedding-v0 | avsolatorio | 🟢 Ouvert | 33,3 % | 31 janvier 2024 | ✅ Mesuré |
| 79 | mixedbread-ai/mxbai-embed-large-v1 | Mixedbread AI | 🟢 Ouvert | 33,3 % | 7 mars 2024 | ✅ Mesuré |
| 80 | Snowflake/snowflake-arctic-embed-l | Snowflake | 🟢 Ouvert | 33,3 % | 12 avril 2024 | ✅ Mesuré |
| 81 | Snowflake/snowflake-arctic-embed-s | Snowflake | 🟢 Ouvert | 33,0 % | 12 avril 2024 | ✅ Mesuré |
| 82 | minishlab/M2V_base_glove_subword | minishlab | 🟢 Ouvert | 33,0 % | 21 septembre 2024 | ✅ Mesuré |
| 83 | Mihaiii/Ivysaur | Mihaiii | 🟢 Ouvert | 32,9 % | 27 avril 2024 | ✅ Mesuré |
| 84 | Snowflake/snowflake-arctic-embed-xs | Snowflake | 🟢 Ouvert | 32,7 % | 8 juillet 2024 | ✅ Mesuré |
| 85 | minishlab/M2V_base_output | minishlab | 🟢 Ouvert | 32,7 % | 21 septembre 2024 | ✅ Mesuré |
| 86 | Mihaiii/Venusaur | Mihaiii | 🟢 Ouvert | 32,6 % | 29 avril 2024 | ✅ Mesuré |
| 87 | sdadas/mmlw-roberta-large | sdadas | 🟢 Ouvert | 32,6 % | 17 novembre 2023 | ✅ Mesuré |
| 88 | avsolatorio/GIST-all-MiniLM-L6-v2 | avsolatorio | 🟢 Ouvert | 32,4 % | 3 février 2024 | ✅ Mesuré |
| 89 | minishlab/potion-base-8M | minishlab | 🟢 Ouvert | 32,2 % | 29 octobre 2024 | ✅ Mesuré |
| 90 | BAAI/bge-small-en-v1.5 | BAAI | 🟢 Ouvert | 32,1 % | 12 septembre 2023 | ✅ Mesuré |
| 91 | Mihaiii/Squirtle | Mihaiii | 🟢 Ouvert | 32,1 % | 30 avril 2024 | ✅ Mesuré |
| 92 | Mihaiii/Wartortle | Mihaiii | 🟢 Ouvert | 32,1 % | 30 avril 2024 | ✅ Mesuré |
| 93 | intfloat/e5-base | Intfloat | 🟢 Ouvert | 32,0 % | 26 décembre 2022 | ✅ Mesuré |
| 94 | minishlab/potion-base-4M | minishlab | 🟢 Ouvert | 32,0 % | 29 octobre 2024 | ✅ Mesuré |
| 95 | Thenlper: GTE-Base | Alibaba | 🟢 Ouvert | 32,0 % | 18 novembre 2025 | ✅ Mesuré |
| 96 | Mihaiii/Bulbasaur | Mihaiii | 🟢 Ouvert | 31,9 % | 27 avril 2024 | ✅ Mesuré |
| 97 | abhinand/MedEmbed-small-v0.1 | abhinand | 🟢 Ouvert | 31,8 % | 20 octobre 2024 | ✅ Mesuré |
| 98 | Sentence Transformers: all-MiniLM-L6-v2 | Sentence Transformers | 🟢 Ouvert | 31,8 % | 17 novembre 2025 | ✅ Mesuré |
| 99 | brahmairesearch/slx-v0.1 | brahmairesearch | 🟢 Ouvert | 31,8 % | 13 août 2024 | ✅ Mesuré |
| 100 | BAAI: bge-large-en-v1.5 | BAAI | 🟢 Ouvert | 31,8 % | 18 novembre 2025 | ✅ Mesuré |
| 101 | Mihaiii/gte-micro-v4 | Mihaiii | 🟢 Ouvert | 31,8 % | 22 avril 2024 | ✅ Mesuré |
| 102 | intfloat/e5-small | Intfloat | 🟢 Ouvert | 31,7 % | 8 février 2024 | ✅ Mesuré |
| 103 | avsolatorio/GIST-small-Embedding-v0 | avsolatorio | 🟢 Ouvert | 31,6 % | 3 février 2024 | ✅ Mesuré |
| 104 | thenlper/gte-small | Alibaba | 🟢 Ouvert | 31,5 % | 27 juillet 2023 | ✅ Mesuré |
| 105 | minishlab/potion-base-2M | minishlab | 🟢 Ouvert | 31,4 % | 29 octobre 2024 | ✅ Mesuré |
| 106 | silma-ai/silma-embeddding-matryoshka-v0.1 | silma-ai | 🟢 Ouvert | 31,3 % | 12 octobre 2024 | ✅ Mesuré |
| 107 | Thenlper: GTE-Large | Alibaba | 🟢 Ouvert | 31,3 % | 18 novembre 2025 | ✅ Mesuré |
| 108 | ibm-granite/granite-embedding-30m-english | IBM | 🟢 Ouvert | 31,3 % | 18 décembre 2024 | ✅ Mesuré |
| 109 | consciousAI/cai-lunaris-text-embeddings | consciousAI | 🟢 Ouvert | 31,2 % | 22 juin 2023 | ✅ Mesuré |
| 110 | aari1995/German_Semantic_STS_V2 | aari1995 | 🟢 Ouvert | 31,2 % | 17 novembre 2022 | ✅ Mesuré |
| 111 | Omartificial-Intelligence-Space/Arabert-all-nli-triplet-Matryoshka | Omartificial-Intelligence-Space | 🟢 Ouvert | 31,2 % | 16 juin 2024 | ✅ Mesuré |
| 112 | sdadas/mmlw-roberta-base | sdadas | 🟢 Ouvert | 31,0 % | 17 novembre 2023 | ✅ Mesuré |
| 113 | ai-forever/ru-en-RoSBERTa | ai-forever | 🟢 Ouvert | 30,6 % | 29 juillet 2024 | ✅ Mesuré |
| 114 | Mihaiii/gte-micro | Mihaiii | 🟢 Ouvert | 30,3 % | 21 avril 2024 | ✅ Mesuré |
| 115 | Omartificial-Intelligence-Space/Marbert-all-nli-triplet-Matryoshka | Omartificial-Intelligence-Space | 🟢 Ouvert | 30,3 % | 17 juin 2024 | ✅ Mesuré |
| 116 | minishlab/M2V_base_glove | minishlab | 🟢 Ouvert | 28,9 % | 21 septembre 2024 | ✅ Mesuré |
| 117 | malenia1/ternary-weight-embedding | malenia1 | 🟢 Ouvert | 28,9 % | 23 octobre 2024 | ✅ Mesuré |
| 118 | jinaai/jina-embeddings-v2-base-en | Jina AI | 🟢 Ouvert | 28,9 % | 27 septembre 2023 | ✅ Mesuré |
Classement établi sur 118 modèles évalués, dont 2 de grands éditeurs. Score médian de l'ensemble : 36,3 %.
Notre analyse
Un score élevé sur MTEB: Indic traduit une bonne qualité globale des embeddings dans les langues indiennes et sur plusieurs usages, de la recherche d’information à la similarité sémantique. Son interprétation doit toutefois tenir compte de l’agrégation de métriques dépendant des tâches, notamment nDCG@10, accuracy, V-measure et Spearman. Une même valeur synthétise donc des aptitudes différentes et ne décrit pas à elle seule chaque composante du benchmark.
Dans la base, 118 modèles sont évalués. Le score médian de 36 %, comparé aux 79 % de codefuse-ai/F2LLM-v2-14B, fait apparaître un écart important entre le milieu du classement et son meilleur résultat. Cette dispersion ne suggère pas une saturation générale du benchmark et met en évidence une forte différenciation entre modèles. La portée demeure centrée sur les embeddings de texte, les langues indiennes et les sept familles de tâches couvertes.
La lecture du classement appelle enfin une prudence méthodologique, car les scores sont majoritairement auto-déclarés par les éditeurs. Ils ne relèvent donc pas tous d’une mesure indépendante et homogène. Le classement fournit une comparaison utile, mais sa rigueur dépend aussi de la cohérence des évaluations rapportées.
Sources des scores : mteb.