German
MTEB: German est un benchmark communautaire créé par MTEB / MMTEB pour évaluer la qualité des plongements de texte en allemand. Il couvre plusieurs usages complémentaires, notamment la classification, le clustering, la classification de paires, le reranking, la recherche et la similarité…
MTEB: German est un benchmark communautaire créé par MTEB / MMTEB pour évaluer la qualité des plongements de texte en allemand. Il couvre plusieurs usages complémentaires, notamment la classification, le clustering, la classification de paires, le reranking, la recherche et la similarité sémantique.
Cette diversité permet d’apprécier la capacité d’un modèle à produire des représentations utiles dans différents contextes, plutôt que sur une tâche unique. Le score moyen MTEB sert ainsi de synthèse, tandis que les métriques propres à chaque tâche apportent une lecture plus ciblée des performances.
Carte d'identité
| Caractéristique | Valeur |
|---|---|
| Éditeur du benchmark | MTEB / MMTEB (communaute embeddings-benchmark) |
| Capacités mesurées | Evalue la qualite des plongements de texte en allemand sur la classification, le clustering, la classification de paires, le reranking, la recherche et la similarite semantique. |
| Modalité | Texte |
| Type de questions | Taches d'embedding de texte (classification, clustering, pair classification, reranking, retrieval, similarite semantique) |
| Métrique d'évaluation | Variable selon la tache (nDCG@10, accuracy, AP, correlation de Spearman) ; score moyen MTEB |
| Accès | Public |
| Licence | Apache-2.0 (cadre MTEB) ; licences des jeux de donnees variables |
| Langues | Allemand |
| Année de publication | 2025 |
| Ressources | Site / dépôt officiel · Article scientifique |
Classement des modèles (96)
| # | Modèle | Éditeur | Licence | Score | Sortie | Fiabilité |
|---|---|---|---|---|---|---|
| 1 | codefuse-ai/F2LLM-v2-14B | codefuse-ai | 🟢 Ouvert | 67,0 % | 9 mars 2026 | ✅ Mesuré |
| 2 | codefuse-ai/F2LLM-v2-8B | codefuse-ai | 🟢 Ouvert | 66,8 % | 9 mars 2026 | ✅ Mesuré |
| 3 | codefuse-ai/F2LLM-v2-4B | codefuse-ai | 🟢 Ouvert | 66,1 % | 9 mars 2026 | ✅ Mesuré |
| 4 | codefuse-ai/F2LLM-v2-1.7B | codefuse-ai | 🟢 Ouvert | 65,1 % | 9 mars 2026 | ✅ Mesuré |
| 5 | codefuse-ai/F2LLM-v2-0.6B | codefuse-ai | 🟢 Ouvert | 63,1 % | 9 mars 2026 | ✅ Mesuré |
| 6 | codefuse-ai/F2LLM-v2-330M | codefuse-ai | 🟢 Ouvert | 61,6 % | 9 mars 2026 | ✅ Mesuré |
| 7 | jinaai/jina-embeddings-v3 | Jina AI | 🟢 Ouvert | 60,0 % | 18 septembre 2024 | ✅ Mesuré |
| 8 | intfloat/e5-mistral-7b-instruct | Intfloat | 🟢 Ouvert | 58,3 % | 8 février 2024 | ✅ Mesuré |
| 9 | codefuse-ai/F2LLM-v2-160M | codefuse-ai | 🟢 Ouvert | 57,3 % | 9 mars 2026 | ✅ Mesuré |
| 10 | intfloat/multilingual-e5-large-instruct | Intfloat | 🟢 Ouvert | 57,1 % | 8 février 2024 | ✅ Mesuré |
| 11 | Lajavaness/bilingual-embedding-large | Lajavaness | 🟢 Ouvert | 56,0 % | 24 juin 2024 | ✅ Mesuré |
| 12 | Snowflake/snowflake-arctic-embed-l-v2.0 | Snowflake | 🟢 Ouvert | 55,7 % | 4 décembre 2024 | ✅ Mesuré |
| 13 | codefuse-ai/F2LLM-v2-80M | codefuse-ai | 🟢 Ouvert | 55,6 % | 9 mars 2026 | ✅ Mesuré |
| 14 | OrdalieTech/Solon-embeddings-large-0.1 | OrdalieTech | 🟢 Ouvert | 55,0 % | 9 décembre 2023 | ✅ Mesuré |
| 15 | Intfloat: Multilingual-E5-Large | Intfloat | 🟢 Ouvert | 54,9 % | 18 novembre 2025 | ✅ Mesuré |
| 16 | Lajavaness/bilingual-embedding-base | Lajavaness | 🟢 Ouvert | 54,0 % | 26 juin 2024 | ✅ Mesuré |
| 17 | HIT-TMG/KaLM-embedding-multilingual-mini-v1 | HIT-TMG | 🟢 Ouvert | 53,4 % | 27 août 2024 | ✅ Mesuré |
| 18 | aari1995/German_Semantic_STS_V2 | aari1995 | 🟢 Ouvert | 53,2 % | 17 novembre 2022 | ✅ Mesuré |
| 19 | intfloat/multilingual-e5-base | Intfloat | 🟢 Ouvert | 53,0 % | 8 février 2024 | ✅ Mesuré |
| 20 | Lajavaness/bilingual-embedding-small | Lajavaness | 🟢 Ouvert | 52,1 % | 17 juillet 2024 | ✅ Mesuré |
| 21 | ibm-granite/granite-embedding-278m-multilingual | IBM | 🟢 Ouvert | 51,8 % | 18 décembre 2024 | ✅ Mesuré |
| 22 | intfloat/multilingual-e5-small | Intfloat | 🟢 Ouvert | 51,5 % | 8 février 2024 | ✅ Mesuré |
| 23 | ibm-granite/granite-embedding-107m-multilingual | IBM | 🟢 Ouvert | 50,2 % | 18 décembre 2024 | ✅ Mesuré |
| 24 | sentence-transformers/paraphrase-multilingual-mpnet-base-v2 | Sentence Transformers | 🟢 Ouvert | 48,6 % | 1 novembre 2019 | ✅ Mesuré |
| 25 | Omartificial-Intelligence-Space/Arabic-all-nli-triplet-Matryoshka | Omartificial-Intelligence-Space | 🟢 Ouvert | 48,3 % | 14 juin 2024 | ✅ Mesuré |
| 26 | Omartificial-Intelligence-Space/Arabic-labse-Matryoshka | Omartificial-Intelligence-Space | 🟢 Ouvert | 47,2 % | 16 juin 2024 | ✅ Mesuré |
| 27 | manu/sentence_croissant_alpha_v0.2 | manu | 🟢 Ouvert | 46,8 % | 15 mars 2024 | ✅ Mesuré |
| 28 | sentence-transformers/LaBSE | Sentence Transformers | 🟢 Ouvert | 46,5 % | 1 novembre 2019 | ✅ Mesuré |
| 29 | Gameselo/STS-multilingual-mpnet-base-v2 | Gameselo | 🟢 Ouvert | 45,8 % | 7 juin 2024 | ✅ Mesuré |
| 30 | Intfloat: E5-Large-v2 | Intfloat | 🟢 Ouvert | 45,7 % | 18 novembre 2025 | ✅ Mesuré |
| 31 | sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2 | Sentence Transformers | 🟢 Ouvert | 45,5 % | 1 novembre 2019 | ✅ Mesuré |
| 32 | Omartificial-Intelligence-Space/Arabic-MiniLM-L12-v2-all-nli-triplet | Omartificial-Intelligence-Space | 🟢 Ouvert | 44,8 % | 25 juin 2024 | ✅ Mesuré |
| 33 | WhereIsAI/UAE-Large-V1 | WhereIsAI | 🟢 Ouvert | 43,9 % | 4 décembre 2023 | ✅ Mesuré |
| 34 | mixedbread-ai/mxbai-embed-large-v1 | Mixedbread AI | 🟢 Ouvert | 43,7 % | 7 mars 2024 | ✅ Mesuré |
| 35 | manu/sentence_croissant_alpha_v0.4 | manu | 🟢 Ouvert | 43,6 % | 27 avril 2024 | ✅ Mesuré |
| 36 | intfloat/e5-large | Intfloat | 🟢 Ouvert | 43,3 % | 26 décembre 2022 | ✅ Mesuré |
| 37 | Thenlper: GTE-Large | Alibaba | 🟢 Ouvert | 43,1 % | 18 novembre 2025 | ✅ Mesuré |
| 38 | BAAI: bge-large-en-v1.5 | BAAI | 🟢 Ouvert | 43,0 % | 18 novembre 2025 | ✅ Mesuré |
| 39 | shibing624/text2vec-base-multilingual | shibing624 | 🟢 Ouvert | 42,9 % | 22 juin 2023 | ✅ Mesuré |
| 40 | nomic-ai/nomic-embed-text-v1-unsupervised | Nomic AI | 🟢 Ouvert | 42,9 % | 15 janvier 2024 | ✅ Mesuré |
| 41 | sentence-transformers/static-similarity-mrl-multilingual-v1 | Sentence Transformers | 🟢 Ouvert | 42,7 % | 15 janvier 2025 | ✅ Mesuré |
| 42 | avsolatorio/GIST-large-Embedding-v0 | avsolatorio | 🟢 Ouvert | 42,6 % | 14 février 2024 | ✅ Mesuré |
| 43 | Intfloat: E5-Base-v2 | Intfloat | 🟢 Ouvert | 42,4 % | 18 novembre 2025 | ✅ Mesuré |
| 44 | sdadas/mmlw-e5-large | sdadas | 🟢 Ouvert | 42,4 % | 17 novembre 2023 | ✅ Mesuré |
| 45 | Snowflake/snowflake-arctic-embed-l | Snowflake | 🟢 Ouvert | 41,2 % | 12 avril 2024 | ✅ Mesuré |
| 46 | intfloat/e5-base | Intfloat | 🟢 Ouvert | 41,0 % | 26 décembre 2022 | ✅ Mesuré |
| 47 | Thenlper: GTE-Base | Alibaba | 🟢 Ouvert | 40,9 % | 18 novembre 2025 | ✅ Mesuré |
| 48 | dwzhu/e5-base-4k | dwzhu | 🟢 Ouvert | 40,8 % | 28 mars 2024 | ✅ Mesuré |
| 49 | sdadas/mmlw-roberta-large | sdadas | 🟢 Ouvert | 40,5 % | 17 novembre 2023 | ✅ Mesuré |
| 50 | BAAI: bge-base-en-v1.5 | BAAI | 🟢 Ouvert | 40,5 % | 18 novembre 2025 | ✅ Mesuré |
| 51 | avsolatorio/GIST-Embedding-v0 | avsolatorio | 🟢 Ouvert | 40,4 % | 31 janvier 2024 | ✅ Mesuré |
| 52 | intfloat/e5-small-v2 | Intfloat | 🟢 Ouvert | 40,3 % | 8 février 2024 | ✅ Mesuré |
| 53 | ibm-granite/granite-embedding-125m-english | IBM | 🟢 Ouvert | 40,2 % | 18 décembre 2024 | ✅ Mesuré |
| 54 | thenlper/gte-small | Alibaba | 🟢 Ouvert | 40,1 % | 27 juillet 2023 | ✅ Mesuré |
| 55 | Mihaiii/Ivysaur | Mihaiii | 🟢 Ouvert | 40,0 % | 27 avril 2024 | ✅ Mesuré |
| 56 | Snowflake/snowflake-arctic-embed-s | Snowflake | 🟢 Ouvert | 39,7 % | 12 avril 2024 | ✅ Mesuré |
| 57 | intfloat/e5-small | Intfloat | 🟢 Ouvert | 39,5 % | 8 février 2024 | ✅ Mesuré |
| 58 | BAAI/bge-small-en-v1.5 | BAAI | 🟢 Ouvert | 39,4 % | 12 septembre 2023 | ✅ Mesuré |
| 59 | abhinand/MedEmbed-small-v0.1 | abhinand | 🟢 Ouvert | 39,0 % | 20 octobre 2024 | ✅ Mesuré |
| 60 | Sentence Transformers: all-mpnet-base-v2 | Sentence Transformers | 🟢 Ouvert | 39,0 % | 17 novembre 2025 | ✅ Mesuré |
| 61 | avsolatorio/GIST-small-Embedding-v0 | avsolatorio | 🟢 Ouvert | 38,7 % | 3 février 2024 | ✅ Mesuré |
| 62 | sdadas/mmlw-e5-base | sdadas | 🟢 Ouvert | 38,4 % | 17 novembre 2023 | ✅ Mesuré |
| 63 | avsolatorio/GIST-all-MiniLM-L6-v2 | avsolatorio | 🟢 Ouvert | 38,4 % | 3 février 2024 | ✅ Mesuré |
| 64 | sdadas/mmlw-roberta-base | sdadas | 🟢 Ouvert | 37,7 % | 17 novembre 2023 | ✅ Mesuré |
| 65 | Snowflake/snowflake-arctic-embed-m | Snowflake | 🟢 Ouvert | 37,3 % | 12 avril 2024 | ✅ Mesuré |
| 66 | Sentence Transformers: all-MiniLM-L12-v2 | Sentence Transformers | 🟢 Ouvert | 37,1 % | 18 novembre 2025 | ✅ Mesuré |
| 67 | ibm-granite/granite-embedding-30m-english | IBM | 🟢 Ouvert | 37,1 % | 18 décembre 2024 | ✅ Mesuré |
| 68 | Sentence Transformers: all-MiniLM-L6-v2 | Sentence Transformers | 🟢 Ouvert | 36,7 % | 17 novembre 2025 | ✅ Mesuré |
| 69 | Mihaiii/Bulbasaur | Mihaiii | 🟢 Ouvert | 36,7 % | 27 avril 2024 | ✅ Mesuré |
| 70 | Mihaiii/gte-micro-v4 | Mihaiii | 🟢 Ouvert | 36,4 % | 22 avril 2024 | ✅ Mesuré |
| 71 | Mihaiii/Venusaur | Mihaiii | 🟢 Ouvert | 36,4 % | 29 avril 2024 | ✅ Mesuré |
| 72 | sergeyzh/LaBSE-ru-turbo | sergeyzh | 🟢 Ouvert | 36,0 % | 27 juin 2024 | ✅ Mesuré |
| 73 | sdadas/mmlw-e5-small | sdadas | 🟢 Ouvert | 35,6 % | 17 novembre 2023 | ✅ Mesuré |
| 74 | Snowflake/snowflake-arctic-embed-xs | Snowflake | 🟢 Ouvert | 35,5 % | 8 juillet 2024 | ✅ Mesuré |
| 75 | Snowflake/snowflake-arctic-embed-m-v1.5 | Snowflake | 🟢 Ouvert | 35,5 % | 8 juillet 2024 | ✅ Mesuré |
| 76 | Mihaiii/Wartortle | Mihaiii | 🟢 Ouvert | 35,3 % | 30 avril 2024 | ✅ Mesuré |
| 77 | cointegrated/LaBSE-en-ru | cointegrated | 🟢 Ouvert | 34,7 % | 10 juin 2021 | ✅ Mesuré |
| 78 | Mihaiii/Squirtle | Mihaiii | 🟢 Ouvert | 34,3 % | 30 avril 2024 | ✅ Mesuré |
| 79 | ai-forever/ru-en-RoSBERTa | ai-forever | 🟢 Ouvert | 32,0 % | 29 juillet 2024 | ✅ Mesuré |
| 80 | deepvk/USER-base | deepvk | 🟢 Ouvert | 31,3 % | 10 juin 2024 | ✅ Mesuré |
| 81 | Mihaiii/gte-micro | Mihaiii | 🟢 Ouvert | 30,9 % | 21 avril 2024 | ✅ Mesuré |
| 82 | consciousAI/cai-lunaris-text-embeddings | consciousAI | 🟢 Ouvert | 28,5 % | 22 juin 2023 | ✅ Mesuré |
| 83 | cointegrated/rubert-tiny | cointegrated | 🟢 Ouvert | 27,1 % | 24 mai 2021 | ✅ Mesuré |
| 84 | DeepPavlov/rubert-base-cased-sentence | DeepPavlov | 🟢 Ouvert | 26,3 % | 4 mars 2020 | ✅ Mesuré |
| 85 | sergeyzh/rubert-tiny-turbo | sergeyzh | 🟢 Ouvert | 26,2 % | 21 juin 2024 | ✅ Mesuré |
| 86 | Omartificial-Intelligence-Space/Arabic-mpnet-base-all-nli-triplet | Omartificial-Intelligence-Space | 🟢 Ouvert | 26,0 % | 15 juin 2024 | ✅ Mesuré |
| 87 | silma-ai/silma-embeddding-matryoshka-v0.1 | silma-ai | 🟢 Ouvert | 25,7 % | 12 octobre 2024 | ✅ Mesuré |
| 88 | Omartificial-Intelligence-Space/Arabert-all-nli-triplet-Matryoshka | Omartificial-Intelligence-Space | 🟢 Ouvert | 25,3 % | 16 juin 2024 | ✅ Mesuré |
| 89 | DeepPavlov/rubert-base-cased | DeepPavlov | 🟢 Ouvert | 24,9 % | 4 mars 2020 | ✅ Mesuré |
| 90 | DeepPavlov/distilrubert-small-cased-conversational | DeepPavlov | 🟢 Ouvert | 24,7 % | 28 juin 2022 | ✅ Mesuré |
| 91 | cointegrated/rubert-tiny2 | cointegrated | 🟢 Ouvert | 24,4 % | 28 octobre 2021 | ✅ Mesuré |
| 92 | malenia1/ternary-weight-embedding | malenia1 | 🟢 Ouvert | 23,7 % | 23 octobre 2024 | ✅ Mesuré |
| 93 | Omartificial-Intelligence-Space/Marbert-all-nli-triplet-Matryoshka | Omartificial-Intelligence-Space | 🟢 Ouvert | 23,1 % | 17 juin 2024 | ✅ Mesuré |
| 94 | deepvk/deberta-v1-base | deepvk | 🟢 Ouvert | 22,2 % | 7 février 2023 | ✅ Mesuré |
| 95 | ai-forever/sbert_large_mt_nlu_ru | ai-forever | 🟢 Ouvert | 21,8 % | 18 mai 2021 | ✅ Mesuré |
| 96 | ai-forever/sbert_large_nlu_ru | ai-forever | 🟢 Ouvert | 20,3 % | 20 novembre 2020 | ✅ Mesuré |
Classement établi sur 96 modèles évalués. Score médian de l'ensemble : 40,7 %.
Notre analyse
Un score élevé indique qu’un modèle obtient de bons résultats agrégés sur plusieurs familles de tâches en allemand. Son interprétation doit toutefois tenir compte de métriques différentes selon les épreuves, parmi lesquelles nDCG@10, accuracy, AP et corrélation de Spearman. Le score moyen facilite le classement global, mais résume donc des capacités distinctes. La fiabilité appelle aussi à la prudence, car les résultats sont majoritairement auto-déclarés par les éditeurs plutôt que tous mesurés dans un protocole indépendant.
Parmi les 96 modèles évalués dans la base, l’écart entre le score médian de 41 % et les 67 % de codefuse-ai/F2LLM-v2-14B montre une différenciation nette des performances. Le classement ne paraît donc pas uniformément saturé au niveau atteint par son meilleur modèle. Sa portée reste circonscrite à la qualité des embeddings de texte en allemand et aux tâches couvertes. Il ne constitue pas une mesure générale de toutes les capacités d’un modèle. Les résultats auto-déclarés imposent enfin une lecture prudente des comparaisons, notamment lorsque de faibles écarts séparent plusieurs modèles.
Sources des scores : mteb.