Mathematics (Baseline)
Benchable : Mathematics (Baseline) est un benchmark public créé par Benchable pour évaluer les compétences mathématiques des modèles d’intelligence artificielle. Il couvre un spectre allant de l’arithmétique élémentaire au calcul, à l’algèbre linéaire, à la topologie et à la géométrie…
Benchable : Mathematics (Baseline) est un benchmark public créé par Benchable pour évaluer les compétences mathématiques des modèles d’intelligence artificielle. Il couvre un spectre allant de l’arithmétique élémentaire au calcul, à l’algèbre linéaire, à la topologie et à la géométrie algébrique.
Son corpus en anglais comprend également des problèmes inédits de niveau doctoral. Présentées sous forme de QCM à six options, les réponses sont validées par correspondance exacte de la lettre attendue. Le benchmark fournit ainsi une mesure standardisée de la capacité à sélectionner une réponse correcte dans des domaines mathématiques variés.
Carte d'identité
| Caractéristique | Valeur |
|---|---|
| Éditeur du benchmark | Benchable |
| Capacités mesurées | Mathematiques a tous niveaux : de l'arithmetique elementaire au calcul, algebre lineaire, topologie, geometrie algebrique et problemes de niveau doctoral |
| Modalité | Texte |
| Type de questions | QCM (6 options A-F) |
| Métrique d'évaluation | Lettre de la reponse correcte (Exact Match, JSON Path $.answer) |
| Accès | Public |
| Langues | anglais |
| Taille du jeu | 100 questions |
| Ressources | Site / dépôt officiel |
Classement des modèles (141)
| # | Modèle | Éditeur | Licence | Score | Sortie | Fiabilité |
|---|---|---|---|---|---|---|
| 1 | StepFun: Step 3.7 Flash | StepFun | 🟢 Ouvert | 100,0 % | 28 mai 2026 | ✅ Mesuré |
| 2 | Arcee AI: Trinity Large Thinking | Arcee AI | 🟢 Ouvert | 97,5 % | 1 avril 2026 | ✅ Mesuré |
| 3 | Writer: Palmyra X5 | Writer | ▫ n.d. | 97,2 % | 21 janvier 2026 | ✅ Mesuré |
| 4 | gemini-3-pro-image | ▫ n.d. | 97,0 % | — | ✅ Mesuré | |
| 5 | qwen3-235b-a22b-07-25 | Alibaba Cloud / Qwen Team | ▫ n.d. | 97,0 % | — | ✅ Mesuré |
| 6 | kimi-k2.5-0127 | Moonshot AI | ▫ n.d. | 96,8 % | — | ✅ Mesuré |
| 7 | inclusionAI: Ling-2.6-1T | InclusionAI | ▫ n.d. | 96,4 % | 23 avril 2026 | ✅ Mesuré |
| 8 | AionLabs: Aion-2.0 | Aion Labs | ▫ n.d. | 96,0 % | 23 février 2026 | ✅ Mesuré |
| 9 | AionLabs: Aion-3.0 | Aion Labs | ▫ n.d. | 96,0 % | 7 juillet 2026 | ✅ Mesuré |
| 10 | Claude Opus 4.5 | Anthropic | 🔒 Propriétaire | 96,0 % | 24 novembre 2025 | ✅ Mesuré |
| 11 | Deep Cogito: Cogito v2.1 671B | Deep Cogito | ▫ n.d. | 96,0 % | 13 novembre 2025 | ✅ Mesuré |
| 12 | GPT-4.1 | OpenAI | 🔒 Propriétaire | 96,0 % | 14 avril 2025 | ✅ Mesuré |
| 13 | GPT-5.6 Sol | OpenAI | 🔒 Propriétaire | 96,0 % | 9 juillet 2026 | ✅ Mesuré |
| 14 | Seed 1.6 | ByteDance Seed | ▫ n.d. | 96,0 % | 23 décembre 2025 | ✅ Mesuré |
| 15 | Gemini 3.1 Pro Preview | ▫ n.d. | 95,9 % | 19 février 2026 | ✅ Mesuré | |
| 16 | Google: Gemini 3.1 Pro Preview Custom Tools | ▫ n.d. | 95,9 % | 25 février 2026 | ✅ Mesuré | |
| 17 | qwen3-32b-04-28 | Alibaba Cloud / Qwen Team | ▫ n.d. | 95,9 % | — | ✅ Mesuré |
| 18 | Thinking Machines: Inkling | Thinking Machines | 🟢 Ouvert | 95,8 % | 17 juillet 2026 | ✅ Mesuré |
| 19 | gemini-2.5-pro-preview-03-25 | ▫ n.d. | 95,8 % | — | ✅ Mesuré | |
| 20 | Kwaipilot: KAT-Coder-Air V2.5 | kwaipilot | ▫ n.d. | 95,4 % | 10 juillet 2026 | ✅ Mesuré |
| 21 | AionLabs: Aion-3.0-Mini | Aion Labs | ▫ n.d. | 95,0 % | 7 juillet 2026 | ✅ Mesuré |
| 22 | GPT-4.1 mini | OpenAI | 🔒 Propriétaire | 95,0 % | 14 avril 2025 | ✅ Mesuré |
| 23 | Kimi K2 | Moonshot AI | ▫ n.d. | 95,0 % | 6 novembre 2025 | ✅ Mesuré |
| 24 | Z.ai: GLM 5 Turbo | Zhipu AI | ▫ n.d. | 95,0 % | 15 mars 2026 | ✅ Mesuré |
| 25 | deepseek-chat-v3-0324 | DeepSeek | ▫ n.d. | 95,0 % | — | ✅ Mesuré |
| 26 | Hy3 | Tencent | 🟢 Ouvert | 94,7 % | 6 juillet 2026 | ✅ Mesuré |
| 27 | llama-4-maverick-17b-128e-instruct | Meta | ▫ n.d. | 94,5 % | — | ✅ Mesuré |
| 28 | DeepSeek V4 Pro | DeepSeek | ▫ n.d. | 94,4 % | 24 avril 2026 | ✅ Mesuré |
| 29 | Kwaipilot: KAT-Coder-Pro V2.5 | kwaipilot | ▫ n.d. | 94,4 % | 10 juillet 2026 | ✅ Mesuré |
| 30 | Claude Opus 4 | Anthropic | 🔒 Propriétaire | 94,0 % | 22 mai 2025 | ✅ Mesuré |
| 31 | Claude Sonnet 4.5 | Anthropic | 🔒 Propriétaire | 94,0 % | 29 septembre 2025 | ✅ Mesuré |
| 32 | DeepSeek V4 Flash | DeepSeek | ▫ n.d. | 94,0 % | 24 avril 2026 | ✅ Mesuré |
| 33 | GPT-5 mini | OpenAI | 🔒 Propriétaire | 94,0 % | 7 août 2025 | ✅ Mesuré |
| 34 | GPT-5 nano | OpenAI | 🔒 Propriétaire | 94,0 % | 7 août 2025 | ✅ Mesuré |
| 35 | GPT-5.6 Luna | OpenAI | 🔒 Propriétaire | 94,0 % | 9 juillet 2026 | ✅ Mesuré |
| 36 | GPT-5.6 Terra | OpenAI | 🔒 Propriétaire | 94,0 % | 9 juillet 2026 | ✅ Mesuré |
| 37 | OpenAI: GPT-5.2 Chat | OpenAI | 🔒 Propriétaire | 94,0 % | 10 décembre 2025 | ✅ Mesuré |
| 38 | OpenAI: GPT-5.6 Luna Pro | OpenAI | 🔒 Propriétaire | 94,0 % | 9 juillet 2026 | ✅ Mesuré |
| 39 | OpenAI: GPT-5.6 Sol Pro | OpenAI | 🔒 Propriétaire | 94,0 % | 9 juillet 2026 | ✅ Mesuré |
| 40 | Qwen: Qwen3 30B A3B Instruct 2507 | Alibaba Cloud / Qwen Team | 🟢 Ouvert | 94,0 % | 29 juillet 2025 | ✅ Mesuré |
| 41 | Tencent: Hy3 preview | Tencent | 🟢 Ouvert | 94,0 % | 22 avril 2026 | ✅ Mesuré |
| 42 | mistral-large-2512 | mistralai | ▫ n.d. | 94,0 % | — | ✅ Mesuré |
| 43 | Qwen: Qwen3 30B A3B Thinking 2507 | Alibaba Cloud / Qwen Team | 🟢 Ouvert | 93,9 % | 28 août 2025 | ✅ Mesuré |
| 44 | Claude Haiku 4.5 | Anthropic | 🔒 Propriétaire | 93,0 % | 15 octobre 2025 | ✅ Mesuré |
| 45 | Claude Sonnet 4 | Anthropic | 🔒 Propriétaire | 93,0 % | 22 mai 2025 | ✅ Mesuré |
| 46 | Grok 4.5 | xAI | 🔒 Propriétaire | 93,0 % | 16 juillet 2026 | ✅ Mesuré |
| 47 | Kwaipilot: KAT-Coder-Pro V2 | kwaipilot | ▫ n.d. | 93,0 % | 27 mars 2026 | ✅ Mesuré |
| 48 | Mistral Large | Mistral AI | ▫ n.d. | 93,0 % | 26 février 2024 | ✅ Mesuré |
| 49 | Mistral: Mistral Medium 3 | Mistral AI | ▫ n.d. | 93,0 % | 7 mai 2025 | ✅ Mesuré |
| 50 | Mistral: Mistral Medium 3.1 | Mistral AI | ▫ n.d. | 93,0 % | 13 août 2025 | ✅ Mesuré |
| 51 | Nex AGI: Nex-N2-Pro | Nex AGI | 🟢 Ouvert | 93,0 % | 8 juin 2026 | ✅ Mesuré |
| 52 | OpenAI: GPT Chat Latest | OpenAI | 🔒 Propriétaire | 93,0 % | 5 mai 2026 | ✅ Mesuré |
| 53 | OpenAI: GPT-5.1 Chat | OpenAI | 🔒 Propriétaire | 93,0 % | 13 novembre 2025 | ✅ Mesuré |
| 54 | OpenAI: GPT-5.1-Codex-Max | OpenAI | 🔒 Propriétaire | 93,0 % | 4 décembre 2025 | ✅ Mesuré |
| 55 | OpenAI: GPT-5.6 Terra Pro | OpenAI | 🔒 Propriétaire | 93,0 % | 9 juillet 2026 | ✅ Mesuré |
| 56 | qwen3-30b-a3b-04-28 | Alibaba Cloud / Qwen Team | ▫ n.d. | 93,0 % | — | ✅ Mesuré |
| 57 | qwen3-next-80b-a3b-instruct-2509 | Alibaba Cloud / Qwen Team | ▫ n.d. | 93,0 % | — | ✅ Mesuré |
| 58 | xAI: Grok Build 0.1 | xAI | ▫ n.d. | 93,0 % | 20 mai 2026 | ✅ Mesuré |
| 59 | Claude Opus 4.1 | Anthropic | 🔒 Propriétaire | 92,9 % | 5 août 2025 | ✅ Mesuré |
| 60 | GPT-4.1 nano | OpenAI | 🔒 Propriétaire | 92,5 % | 14 avril 2025 | ✅ Mesuré |
| 61 | GPT-5 | OpenAI | 🔒 Propriétaire | 92,0 % | 7 août 2025 | ✅ Mesuré |
| 62 | GPT-5.1 | OpenAI | 🔒 Propriétaire | 92,0 % | 13 novembre 2025 | ✅ Mesuré |
| 63 | GPT-5.2 | OpenAI | 🔒 Propriétaire | 92,0 % | 11 décembre 2025 | ✅ Mesuré |
| 64 | OpenAI: GPT-5.4 Image 2 | OpenAI | 🔒 Propriétaire | 92,0 % | 21 avril 2026 | ✅ Mesuré |
| 65 | Perplexity: Sonar Pro Search | Perplexity | 🔒 Propriétaire | 92,0 % | 30 octobre 2025 | ✅ Mesuré |
| 66 | deepseek-chat-v3 | DeepSeek | ▫ n.d. | 92,0 % | — | ✅ Mesuré |
| 67 | gemini-3.1-flash-image | ▫ n.d. | 92,0 % | — | ✅ Mesuré | |
| 68 | gpt-4o-search-preview-2025-03-11 | OpenAI | 🔒 Propriétaire | 92,0 % | 12 mars 2025 | ✅ Mesuré |
| 69 | laguna-xs-2.1 | poolside | ▫ n.d. | 92,0 % | — | ✅ Mesuré |
| 70 | mistral-small-2603 | mistralai | ▫ n.d. | 92,0 % | — | ✅ Mesuré |
| 71 | inclusionAI: Ring-2.6-1T | InclusionAI | ▫ n.d. | 91,7 % | 8 mai 2026 | ✅ Mesuré |
| 72 | DeepSeek V3.1 Terminus | DeepSeek | 🟢 Ouvert | 91,0 % | 22 septembre 2025 | ✅ Mesuré |
| 73 | gemini-3.1-flash-image-preview | ▫ n.d. | 91,0 % | — | ✅ Mesuré | |
| 74 | laguna-m.1 | poolside | ▫ n.d. | 91,0 % | — | ✅ Mesuré |
| 75 | qwen-plus-2025-01-25 | Alibaba Cloud / Qwen Team | ▫ n.d. | 91,0 % | 8 septembre 2025 | ✅ Mesuré |
| 76 | deepseek-chat-v3.1 | DeepSeek | ▫ n.d. | 90,9 % | — | ✅ Mesuré |
| 77 | qwen3-coder-next-2025-02-03 | Alibaba Cloud / Qwen Team | 🟢 Ouvert | 90,7 % | 4 février 2026 | ✅ Mesuré |
| 78 | GPT-4o | OpenAI | 🔒 Propriétaire | 90,0 % | 27 mars 2025 | ✅ Mesuré |
| 79 | OpenAI: gpt-oss-safeguard-20b | OpenAI | 🟢 Ouvert | 89,0 % | 29 octobre 2025 | ✅ Mesuré |
| 80 | Qwen: Qwen3 Coder 30B A3B Instruct | Alibaba Cloud / Qwen Team | 🟢 Ouvert | 89,0 % | 31 juillet 2025 | ✅ Mesuré |
| 81 | olmo-3-32b-think | allenai | ▫ n.d. | 89,0 % | — | ✅ Mesuré |
| 82 | qwen3-next-80b-a3b-thinking-2509 | Alibaba Cloud / Qwen Team | ▫ n.d. | 88,9 % | — | ✅ Mesuré |
| 83 | Baidu: ERNIE 4.5 VL 424B A47B | Baidu | 🟢 Ouvert | 88,0 % | 30 juin 2025 | ✅ Mesuré |
| 84 | Mistral Large 2407 | Mistral AI | ▫ n.d. | 88,0 % | 19 novembre 2024 | ✅ Mesuré |
| 85 | Sonar | Perplexity | 🔒 Propriétaire | 88,0 % | 29 janvier 2025 | ✅ Mesuré |
| 86 | MiniMax M1 | MiniMax | 🟢 Ouvert | 87,5 % | 17 juin 2025 | ✅ Mesuré |
| 87 | ByteDance Seed: Seed-2.0-Mini | ByteDance Seed | ▫ n.d. | 87,0 % | 26 février 2026 | ✅ Mesuré |
| 88 | Muse Spark 1.1 | Meta | 🔒 Propriétaire | 87,0 % | 9 juillet 2026 | ✅ Mesuré |
| 89 | o1 | OpenAI | 🔒 Propriétaire | 87,0 % | 17 décembre 2024 | ✅ Mesuré |
| 90 | Mistral: Mixtral 8x22B Instruct | Mistral AI | 🟢 Ouvert | 86,0 % | 17 avril 2024 | ✅ Mesuré |
| 91 | Sonar Pro | Perplexity | 🔒 Propriétaire | 86,0 % | 7 mars 2025 | ✅ Mesuré |
| 92 | WizardLM-2 8x22B | Microsoft | 🟢 Ouvert | 86,0 % | 16 avril 2024 | ✅ Mesuré |
| 93 | Command A+ | Cohere | 🟢 Ouvert | 85,0 % | 20 mai 2026 | ✅ Mesuré |
| 94 | Relace: Relace Search | relace | ▫ n.d. | 85,0 % | 8 décembre 2025 | ✅ Mesuré |
| 95 | Upstage: Solar Pro 3 | Upstage | ▫ n.d. | 85,0 % | 27 janvier 2026 | ✅ Mesuré |
| 96 | qwen3-8b-04-28 | Alibaba Cloud / Qwen Team | ▫ n.d. | 85,0 % | — | ✅ Mesuré |
| 97 | Qwen: Qwen3 Coder Flash | Alibaba Cloud / Qwen Team | ▫ n.d. | 84,0 % | 17 septembre 2025 | ✅ Mesuré |
| 98 | gpt-4o-mini-search-preview-2025-03-11 | OpenAI | 🔒 Propriétaire | 84,0 % | 12 mars 2025 | ✅ Mesuré |
| 99 | nova-2-lite-v1 | Amazon | 🔒 Propriétaire | 83,0 % | 2 décembre 2025 | ✅ Mesuré |
| 100 | ministral-3b-2512 | mistralai | ▫ n.d. | 82,0 % | — | ✅ Mesuré |
| 101 | ministral-14b-2512 | mistralai | ▫ n.d. | 81,0 % | — | ✅ Mesuré |
| 102 | uncensored | venice | ▫ n.d. | 81,0 % | — | ✅ Mesuré |
| 103 | ministral-8b-2512 | mistralai | ▫ n.d. | 80,0 % | — | ✅ Mesuré |
| 104 | Arcee AI: Virtuoso Large | Arcee AI | ▫ n.d. | 79,0 % | 5 mai 2025 | ✅ Mesuré |
| 105 | Nous: Hermes 4 405B | Nous Research | 🟢 Ouvert | 79,0 % | 26 août 2025 | ✅ Mesuré |
| 106 | ui-tars-1.5-7b | ByteDance | ▫ n.d. | 78,0 % | — | ✅ Mesuré |
| 107 | qwen3-coder-480b-a35b-07-25 | Alibaba Cloud / Qwen Team | ▫ n.d. | 77,8 % | — | ✅ Mesuré |
| 108 | ByteDance Seed: Seed 1.6 Flash | ByteDance Seed | ▫ n.d. | 77,0 % | 23 décembre 2025 | ✅ Mesuré |
| 109 | hermes-3-llama-3.1-405b | nousresearch | ▫ n.d. | 77,0 % | — | ✅ Mesuré |
| 110 | Nex AGI: Nex-N2-Mini | Nex AGI | 🟢 Ouvert | 76,5 % | 24 juin 2026 | ✅ Mesuré |
| 111 | cydonia-24b-v4.1 | thedrummer | ▫ n.d. | 73,9 % | — | ✅ Mesuré |
| 112 | GPT-4o mini | OpenAI | 🔒 Propriétaire | 71,0 % | 18 juillet 2024 | ✅ Mesuré |
| 113 | IBM: Granite 4.1 8B | IBM | 🟢 Ouvert | 71,0 % | 30 avril 2026 | ✅ Mesuré |
| 114 | Cohere: Command R (08-2024) | Cohere | 🟢 Ouvert | 70,0 % | 30 août 2024 | ✅ Mesuré |
| 115 | command-r-plus-08-2024 | Cohere | ▫ n.d. | 70,0 % | — | ✅ Mesuré |
| 116 | Nous: Hermes 4 70B | Nous Research | 🟢 Ouvert | 69,0 % | 26 août 2025 | ✅ Mesuré |
| 117 | hermes-3-llama-3.1-70b | nousresearch | ▫ n.d. | 68,0 % | — | ✅ Mesuré |
| 118 | Cohere: Command R7B (12-2024) | Cohere | ▫ n.d. | 62,0 % | 14 décembre 2024 | ✅ Mesuré |
| 119 | Inflection: Inflection 3 Pi | Inflection | 🔒 Propriétaire | 59,0 % | 11 octobre 2024 | ✅ Mesuré |
| 120 | Tencent: Hunyuan A13B Instruct | Tencent | 🟢 Ouvert | 59,0 % | 8 juillet 2025 | ✅ Mesuré |
| 121 | Claude 3 Haiku | Anthropic | 🔒 Propriétaire | 58,0 % | 13 mars 2024 | ✅ Mesuré |
| 122 | Inflection: Inflection 3 Productivity | Inflection | 🔒 Propriétaire | 57,0 % | 11 octobre 2024 | ✅ Mesuré |
| 123 | nemotron-nano-12b-v2-vl | NVIDIA | 🟢 Ouvert | 52,6 % | 28 octobre 2025 | ✅ Mesuré |
| 124 | nova-premier-v1 | Amazon | 🔒 Propriétaire | 48,0 % | — | ✅ Mesuré |
| 125 | nova-pro-v1 | Amazon | 🔒 Propriétaire | 42,0 % | — | ✅ Mesuré |
| 126 | AI21: Jamba Large 1.7 | AI21 Labs | 🟢 Ouvert | 40,0 % | 8 août 2025 | ✅ Mesuré |
| 127 | llama-4-scout-17b-16e-instruct | Meta | ▫ n.d. | 39,0 % | — | ✅ Mesuré |
| 128 | nova-micro-v1 | Amazon | 🔒 Propriétaire | 39,0 % | — | ✅ Mesuré |
| 129 | inclusionAI: Ling-2.6-flash | InclusionAI | ▫ n.d. | 37,0 % | 21 avril 2026 | ✅ Mesuré |
| 130 | l3-lunaris-8b | sao10k | ▫ n.d. | 19,0 % | — | ✅ Mesuré |
| 131 | Mistral NeMo | Mistral AI | ▫ n.d. | 15,0 % | 18 juillet 2024 | ✅ Mesuré |
| 132 | l3.3-euryale-70b-v2.3 | sao10k | ▫ n.d. | 13,0 % | — | ✅ Mesuré |
| 133 | Mistral: Codestral 2508 | Mistral AI | ▫ n.d. | 3,0 % | 1 août 2025 | ✅ Mesuré |
| 134 | Meta: Llama 3.2 1B Instruct | Meta | 🟢 Ouvert | 0,0 % | 25 septembre 2024 | ✅ Mesuré |
| 135 | MiniMax: MiniMax-01 | MiniMax | 🟢 Ouvert | 0,0 % | 15 janvier 2025 | ✅ Mesuré |
| 136 | Qwen 3.5 Plus | Qwen | ▫ n.d. | 0,0 % | 16 février 2026 | ✅ Mesuré |
| 137 | Qwen: Qwen3.5-Flash | Alibaba Cloud / Qwen Team | ▫ n.d. | 0,0 % | 25 février 2026 | ✅ Mesuré |
| 138 | Qwen: Qwen3.6 Flash | Alibaba Cloud / Qwen Team | ▫ n.d. | 0,0 % | 27 avril 2026 | ✅ Mesuré |
| 139 | granite-4.0-h-micro | IBM | 🟢 Ouvert | 0,0 % | 20 octobre 2025 | ✅ Mesuré |
| 140 | qwen3.6-plus-04-02 | Alibaba Cloud / Qwen Team | ▫ n.d. | 0,0 % | 2 avril 2026 | ✅ Mesuré |
| 141 | weaver | mancer | ▫ n.d. | 0,0 % | — | ✅ Mesuré |
Classement établi sur 141 modèles évalués, dont 64 de grands éditeurs. Score médian de l'ensemble : 91,7 %.
Notre analyse
Un score élevé indique une forte aptitude à identifier la bonne option parmi six réponses sur ce corpus, dans des disciplines et à des niveaux de difficulté très variés. La validation par Exact Match impose une sortie strictement conforme, mais elle mesure la lettre finale plutôt que la qualité du raisonnement. La fiabilité est renforcée par des scores au moins partiellement mesurés par un tiers, ce qui apporte davantage de rigueur qu’un ensemble reposant uniquement sur des résultats auto-déclarés.
Avec une médiane de 92 % parmi 225 modèles et un meilleur score de 100 % obtenu par Qwen3 VL 32B Instruct, le classement montre un niveau global très élevé. Cette concentration près du maximum signale une saturation partielle et réduit le pouvoir de discrimination entre les modèles les plus performants. Le caractère public du benchmark implique aussi un risque d’exposition préalable aux questions, donc de contamination. Enfin, ses 100 QCM évaluent une portée mathématique large, mais restent centrés sur la sélection d’une réponse en anglais, sans évaluation directe d’une démonstration ou d’un raisonnement développé.
Sources des scores : benchable.