Reasoning (Baseline)
Benchable : Reasoning (Baseline) est un benchmark public créé par Benchable pour évaluer le raisonnement complexe des modèles d’intelligence artificielle. Il s’appuie sur des questions ouvertes couvrant notamment la logique, la reconnaissance de motifs, l’inférence déductive, le…
Benchable : Reasoning (Baseline) est un benchmark public créé par Benchable pour évaluer le raisonnement complexe des modèles d’intelligence artificielle. Il s’appuie sur des questions ouvertes couvrant notamment la logique, la reconnaissance de motifs, l’inférence déductive, le raisonnement mathématique et la résolution abstraite.
Ses tâches prennent la forme d’énigmes logiques, de suites, de cryptarithmes, de problèmes spatiaux et d’exercices de déduction. Le benchmark sert ainsi de référence de base pour comparer la capacité des modèles à produire une solution logique précise plutôt qu’une simple réponse à choix multiple.
Carte d'identité
| Caractéristique | Valeur |
|---|---|
| Éditeur du benchmark | Benchable |
| Capacités mesurées | Raisonnement complexe : logique, reconnaissance de motifs, inference deductive, raisonnement mathematique, resolution abstraite |
| Modalité | Texte |
| Type de questions | Questions ouvertes (enigmes logiques, suites, cryptarithmes, raisonnement spatial, deduction) |
| Métrique d'évaluation | Solution logique precise, validation par equivalence semantique (verification par IA sur JSON Path $.answer) |
| Accès | Public |
| Langues | anglais |
| Taille du jeu | 50 questions |
| Ressources | Site / dépôt officiel |
Classement des modèles (156)
| # | Modèle | Éditeur | Licence | Score | Sortie | Fiabilité |
|---|---|---|---|---|---|---|
| 1 | AionLabs: Aion-2.0 | Aion Labs | ▫ n.d. | 100,0 % | 23 février 2026 | ✅ Mesuré |
| 2 | Arcee AI: Trinity Large Thinking | Arcee AI | 🟢 Ouvert | 100,0 % | 1 avril 2026 | ✅ Mesuré |
| 3 | DeepSeek V4 Flash | DeepSeek | ▫ n.d. | 100,0 % | 24 avril 2026 | ✅ Mesuré |
| 4 | DeepSeek V4 Pro | DeepSeek | ▫ n.d. | 100,0 % | 24 avril 2026 | ✅ Mesuré |
| 5 | Gemini 3.1 Pro Preview | ▫ n.d. | 100,0 % | 19 février 2026 | ✅ Mesuré | |
| 6 | Google: Gemini 3.1 Pro Preview Custom Tools | ▫ n.d. | 100,0 % | 25 février 2026 | ✅ Mesuré | |
| 7 | Grok 4.5 | xAI | 🔒 Propriétaire | 100,0 % | 16 juillet 2026 | ✅ Mesuré |
| 8 | Kwaipilot: KAT-Coder-Pro V2.5 | kwaipilot | ▫ n.d. | 100,0 % | 10 juillet 2026 | ✅ Mesuré |
| 9 | Seed 1.6 | ByteDance Seed | ▫ n.d. | 100,0 % | 23 décembre 2025 | ✅ Mesuré |
| 10 | gemini-3-pro-image | ▫ n.d. | 100,0 % | — | ✅ Mesuré | |
| 11 | inclusionAI: Ring-2.6-1T | InclusionAI | ▫ n.d. | 100,0 % | 8 mai 2026 | ✅ Mesuré |
| 12 | kimi-k2.5-0127 | Moonshot AI | ▫ n.d. | 100,0 % | — | ✅ Mesuré |
| 13 | o1 | OpenAI | 🔒 Propriétaire | 100,0 % | 17 décembre 2024 | ✅ Mesuré |
| 14 | o3 | OpenAI | 🔒 Propriétaire | 100,0 % | 16 avril 2025 | ✅ Mesuré |
| 15 | qwen3-32b-04-28 | Alibaba Cloud / Qwen Team | ▫ n.d. | 100,0 % | — | ✅ Mesuré |
| 16 | AionLabs: Aion-3.0 | Aion Labs | ▫ n.d. | 98,0 % | 7 juillet 2026 | ✅ Mesuré |
| 17 | AionLabs: Aion-3.0-Mini | Aion Labs | ▫ n.d. | 98,0 % | 7 juillet 2026 | ✅ Mesuré |
| 18 | Claude Opus 4.5 | Anthropic | 🔒 Propriétaire | 98,0 % | 24 novembre 2025 | ✅ Mesuré |
| 19 | GPT-5 | OpenAI | 🔒 Propriétaire | 98,0 % | 7 août 2025 | ✅ Mesuré |
| 20 | GPT-5.6 Sol | OpenAI | 🔒 Propriétaire | 98,0 % | 9 juillet 2026 | ✅ Mesuré |
| 21 | GPT-5.6 Terra | OpenAI | 🔒 Propriétaire | 98,0 % | 9 juillet 2026 | ✅ Mesuré |
| 22 | OpenAI: GPT-5.6 Sol Pro | OpenAI | 🔒 Propriétaire | 98,0 % | 9 juillet 2026 | ✅ Mesuré |
| 23 | OpenAI: GPT-5.6 Terra Pro | OpenAI | 🔒 Propriétaire | 98,0 % | 9 juillet 2026 | ✅ Mesuré |
| 24 | Qwen: Qwen3 30B A3B Thinking 2507 | Alibaba Cloud / Qwen Team | 🟢 Ouvert | 98,0 % | 28 août 2025 | ✅ Mesuré |
| 25 | Sakana: Fugu Ultra | Sakana AI | ▫ n.d. | 98,0 % | 24 juin 2026 | ✅ Mesuré |
| 26 | Tencent: Hy3 preview | Tencent | 🟢 Ouvert | 98,0 % | 22 avril 2026 | ✅ Mesuré |
| 27 | gemini-3.1-flash-image-preview | ▫ n.d. | 98,0 % | — | ✅ Mesuré | |
| 28 | laguna-m.1 | poolside | ▫ n.d. | 98,0 % | — | ✅ Mesuré |
| 29 | qwen3-14b-04-28 | Alibaba Cloud / Qwen Team | ▫ n.d. | 98,0 % | — | ✅ Mesuré |
| 30 | qwen3-235b-a22b-04-28 | Alibaba Cloud / Qwen Team | ▫ n.d. | 98,0 % | — | ✅ Mesuré |
| 31 | Thinking Machines: Inkling | Thinking Machines | 🟢 Ouvert | 97,9 % | 17 juillet 2026 | ✅ Mesuré |
| 32 | ByteDance Seed: Seed-2.0-Mini | ByteDance Seed | ▫ n.d. | 96,0 % | 26 février 2026 | ✅ Mesuré |
| 33 | GPT-5 mini | OpenAI | 🔒 Propriétaire | 96,0 % | 7 août 2025 | ✅ Mesuré |
| 34 | GPT-5 nano | OpenAI | 🔒 Propriétaire | 96,0 % | 7 août 2025 | ✅ Mesuré |
| 35 | GPT-5.1 | OpenAI | 🔒 Propriétaire | 96,0 % | 13 novembre 2025 | ✅ Mesuré |
| 36 | Nex AGI: Nex-N2-Pro | Nex AGI | 🟢 Ouvert | 96,0 % | 8 juin 2026 | ✅ Mesuré |
| 37 | Tencent: Hunyuan A13B Instruct | Tencent | 🟢 Ouvert | 96,0 % | 8 juillet 2025 | ✅ Mesuré |
| 38 | gemini-3.1-flash-image | ▫ n.d. | 96,0 % | — | ✅ Mesuré | |
| 39 | qwen3-30b-a3b-04-28 | Alibaba Cloud / Qwen Team | ▫ n.d. | 96,0 % | — | ✅ Mesuré |
| 40 | qwen3-next-80b-a3b-thinking-2509 | Alibaba Cloud / Qwen Team | ▫ n.d. | 96,0 % | — | ✅ Mesuré |
| 41 | xAI: Grok Build 0.1 | xAI | ▫ n.d. | 96,0 % | 20 mai 2026 | ✅ Mesuré |
| 42 | ByteDance Seed: Seed 1.6 Flash | ByteDance Seed | ▫ n.d. | 94,0 % | 23 décembre 2025 | ✅ Mesuré |
| 43 | OpenAI: GPT Chat Latest | OpenAI | 🔒 Propriétaire | 94,0 % | 5 mai 2026 | ✅ Mesuré |
| 44 | OpenAI: GPT-5.6 Luna Pro | OpenAI | 🔒 Propriétaire | 94,0 % | 9 juillet 2026 | ✅ Mesuré |
| 45 | OpenAI: gpt-oss-safeguard-20b | OpenAI | 🟢 Ouvert | 94,0 % | 29 octobre 2025 | ✅ Mesuré |
| 46 | Z.ai: GLM 5 Turbo | Zhipu AI | ▫ n.d. | 94,0 % | 15 mars 2026 | ✅ Mesuré |
| 47 | nemotron-nano-12b-v2-vl | NVIDIA | 🟢 Ouvert | 94,0 % | 28 octobre 2025 | ✅ Mesuré |
| 48 | MiniMax M1 | MiniMax | 🟢 Ouvert | 93,9 % | 17 juin 2025 | ✅ Mesuré |
| 49 | OpenAI: GPT-5.1 Chat | OpenAI | 🔒 Propriétaire | 93,9 % | 13 novembre 2025 | ✅ Mesuré |
| 50 | Claude Opus 4.1 | Anthropic | 🔒 Propriétaire | 92,0 % | 5 août 2025 | ✅ Mesuré |
| 51 | Kwaipilot: KAT-Coder-Air V2.5 | kwaipilot | ▫ n.d. | 92,0 % | 10 juillet 2026 | ✅ Mesuré |
| 52 | laguna-xs-2.1 | poolside | ▫ n.d. | 92,0 % | — | ✅ Mesuré |
| 53 | Claude Opus 4 | Anthropic | 🔒 Propriétaire | 90,0 % | 22 mai 2025 | ✅ Mesuré |
| 54 | DeepSeek V3.1 Terminus | DeepSeek | 🟢 Ouvert | 90,0 % | 22 septembre 2025 | ✅ Mesuré |
| 55 | GPT-5.6 Luna | OpenAI | 🔒 Propriétaire | 90,0 % | 9 juillet 2026 | ✅ Mesuré |
| 56 | Claude Sonnet 4 | Anthropic | 🔒 Propriétaire | 88,0 % | 22 mai 2025 | ✅ Mesuré |
| 57 | Claude Sonnet 4.5 | Anthropic | 🔒 Propriétaire | 88,0 % | 29 septembre 2025 | ✅ Mesuré |
| 58 | deepseek-chat-v3-0324 | DeepSeek | ▫ n.d. | 88,0 % | — | ✅ Mesuré |
| 59 | olmo-3-32b-think | allenai | ▫ n.d. | 88,0 % | — | ✅ Mesuré |
| 60 | qwen3-next-80b-a3b-instruct-2509 | Alibaba Cloud / Qwen Team | ▫ n.d. | 88,0 % | — | ✅ Mesuré |
| 61 | Kimi K2 | Moonshot AI | ▫ n.d. | 86,0 % | 6 novembre 2025 | ✅ Mesuré |
| 62 | OpenAI: GPT-5.4 Image 2 | OpenAI | 🔒 Propriétaire | 86,0 % | 21 avril 2026 | ✅ Mesuré |
| 63 | qwen-plus-2025-01-25 | Alibaba Cloud / Qwen Team | ▫ n.d. | 86,0 % | 8 septembre 2025 | ✅ Mesuré |
| 64 | GPT-4o | OpenAI | 🔒 Propriétaire | 84,0 % | 27 mars 2025 | ✅ Mesuré |
| 65 | Kwaipilot: KAT-Coder-Pro V2 | kwaipilot | ▫ n.d. | 84,0 % | 27 mars 2026 | ✅ Mesuré |
| 66 | qwen3-235b-a22b-07-25 | Alibaba Cloud / Qwen Team | ▫ n.d. | 84,0 % | — | ✅ Mesuré |
| 67 | qwen3-coder-480b-a35b-07-25 | Alibaba Cloud / Qwen Team | ▫ n.d. | 83,3 % | — | ✅ Mesuré |
| 68 | GPT-4.1 | OpenAI | 🔒 Propriétaire | 82,0 % | 14 avril 2025 | ✅ Mesuré |
| 69 | Hy3 | Tencent | 🟢 Ouvert | 81,2 % | 6 juillet 2026 | ✅ Mesuré |
| 70 | Writer: Palmyra X5 | Writer | ▫ n.d. | 80,6 % | 21 janvier 2026 | ✅ Mesuré |
| 71 | OpenAI: GPT-5.1-Codex-Max | OpenAI | 🔒 Propriétaire | 80,0 % | 4 décembre 2025 | ✅ Mesuré |
| 72 | deepseek-chat-v3 | DeepSeek | ▫ n.d. | 80,0 % | — | ✅ Mesuré |
| 73 | deepseek-chat-v3.1 | DeepSeek | ▫ n.d. | 80,0 % | — | ✅ Mesuré |
| 74 | llama-4-maverick-17b-128e-instruct | Meta | ▫ n.d. | 80,0 % | — | ✅ Mesuré |
| 75 | mistral-large-2512 | mistralai | ▫ n.d. | 80,0 % | — | ✅ Mesuré |
| 76 | qwen3-coder-next-2025-02-03 | Alibaba Cloud / Qwen Team | 🟢 Ouvert | 79,6 % | 4 février 2026 | ✅ Mesuré |
| 77 | Deep Cogito: Cogito v2.1 671B | Deep Cogito | ▫ n.d. | 78,0 % | 13 novembre 2025 | ✅ Mesuré |
| 78 | OpenAI: GPT-5.2 Chat | OpenAI | 🔒 Propriétaire | 78,0 % | 10 décembre 2025 | ✅ Mesuré |
| 79 | Qwen: Qwen3 30B A3B Instruct 2507 | Alibaba Cloud / Qwen Team | 🟢 Ouvert | 78,0 % | 29 juillet 2025 | ✅ Mesuré |
| 80 | Qwen: Qwen3 Coder 30B A3B Instruct | Alibaba Cloud / Qwen Team | 🟢 Ouvert | 78,0 % | 31 juillet 2025 | ✅ Mesuré |
| 81 | Sonar Pro | Perplexity | 🔒 Propriétaire | 78,0 % | 7 mars 2025 | ✅ Mesuré |
| 82 | Claude Haiku 4.5 | Anthropic | 🔒 Propriétaire | 76,0 % | 15 octobre 2025 | ✅ Mesuré |
| 83 | GPT-5.2 | OpenAI | 🔒 Propriétaire | 76,0 % | 11 décembre 2025 | ✅ Mesuré |
| 84 | Qwen: Qwen3 Coder Flash | Alibaba Cloud / Qwen Team | ▫ n.d. | 76,0 % | 17 septembre 2025 | ✅ Mesuré |
| 85 | Arcee AI: Virtuoso Large | Arcee AI | ▫ n.d. | 74,0 % | 5 mai 2025 | ✅ Mesuré |
| 86 | MiniMax: MiniMax-01 | MiniMax | 🟢 Ouvert | 74,0 % | 15 janvier 2025 | ✅ Mesuré |
| 87 | Perplexity: Sonar Pro Search | Perplexity | 🔒 Propriétaire | 74,0 % | 30 octobre 2025 | ✅ Mesuré |
| 88 | Sonar | Perplexity | 🔒 Propriétaire | 74,0 % | 29 janvier 2025 | ✅ Mesuré |
| 89 | mistral-small-2603 | mistralai | ▫ n.d. | 74,0 % | — | ✅ Mesuré |
| 90 | GPT-4.1 mini | OpenAI | 🔒 Propriétaire | 72,0 % | 14 avril 2025 | ✅ Mesuré |
| 91 | Mistral: Mistral Medium 3 | Mistral AI | ▫ n.d. | 72,0 % | 7 mai 2025 | ✅ Mesuré |
| 92 | Mistral: Mistral Medium 3.1 | Mistral AI | ▫ n.d. | 72,0 % | 13 août 2025 | ✅ Mesuré |
| 93 | Nous: Hermes 4 405B | Nous Research | 🟢 Ouvert | 72,0 % | 26 août 2025 | ✅ Mesuré |
| 94 | hermes-3-llama-3.1-405b | nousresearch | ▫ n.d. | 70,0 % | — | ✅ Mesuré |
| 95 | Magnum v4 72B | anthracite-org | 🟢 Ouvert | 68,0 % | 22 octobre 2024 | ✅ Mesuré |
| 96 | nova-premier-v1 | Amazon | 🔒 Propriétaire | 68,0 % | — | ✅ Mesuré |
| 97 | Relace: Relace Search | relace | ▫ n.d. | 66,0 % | 8 décembre 2025 | ✅ Mesuré |
| 98 | gpt-4o-mini-search-preview-2025-03-11 | OpenAI | 🔒 Propriétaire | 66,0 % | 12 mars 2025 | ✅ Mesuré |
| 99 | Mistral Large | Mistral AI | ▫ n.d. | 64,0 % | 26 février 2024 | ✅ Mesuré |
| 100 | Mistral: Voxtral Small 24B 2507 | Mistral AI | 🟢 Ouvert | 64,0 % | 30 octobre 2025 | ✅ Mesuré |
| 101 | Nex AGI: Nex-N2-Mini | Nex AGI | 🟢 Ouvert | 62,0 % | 24 juin 2026 | ✅ Mesuré |
| 102 | mistral-small-24b-instruct-2501 | mistralai | ▫ n.d. | 62,0 % | — | ✅ Mesuré |
| 103 | Command A+ | Cohere | 🟢 Ouvert | 60,0 % | 20 mai 2026 | ✅ Mesuré |
| 104 | mistral-saba-2502 | mistralai | ▫ n.d. | 60,0 % | — | ✅ Mesuré |
| 105 | skyfall-36b-v2 | thedrummer | ▫ n.d. | 60,0 % | — | ✅ Mesuré |
| 106 | Inflection: Inflection 3 Productivity | Inflection | 🔒 Propriétaire | 58,0 % | 11 octobre 2024 | ✅ Mesuré |
| 107 | Mistral Large 2407 | Mistral AI | ▫ n.d. | 58,0 % | 19 novembre 2024 | ✅ Mesuré |
| 108 | llama-4-scout-17b-16e-instruct | Meta | ▫ n.d. | 58,0 % | — | ✅ Mesuré |
| 109 | uncensored | venice | ▫ n.d. | 57,1 % | — | ✅ Mesuré |
| 110 | GPT-4.1 nano | OpenAI | 🔒 Propriétaire | 56,0 % | 14 avril 2025 | ✅ Mesuré |
| 111 | GPT-4o mini | OpenAI | 🔒 Propriétaire | 56,0 % | 18 juillet 2024 | ✅ Mesuré |
| 112 | cydonia-24b-v4.1 | thedrummer | ▫ n.d. | 56,0 % | — | ✅ Mesuré |
| 113 | qwen3-8b-04-28 | Alibaba Cloud / Qwen Team | ▫ n.d. | 56,0 % | — | ✅ Mesuré |
| 114 | Perceptron: Perceptron Mk1 | Perceptron | ▫ n.d. | 55,1 % | 12 mai 2026 | ✅ Mesuré |
| 115 | AI21: Jamba Large 1.7 | AI21 Labs | 🟢 Ouvert | 54,0 % | 8 août 2025 | ✅ Mesuré |
| 116 | Mistral: Mixtral 8x22B Instruct | Mistral AI | 🟢 Ouvert | 54,0 % | 17 avril 2024 | ✅ Mesuré |
| 117 | Nous: Hermes 4 70B | Nous Research | 🟢 Ouvert | 54,0 % | 26 août 2025 | ✅ Mesuré |
| 118 | Baidu: ERNIE 4.5 VL 424B A47B | Baidu | 🟢 Ouvert | 52,0 % | 30 juin 2025 | ✅ Mesuré |
| 119 | Inflection: Inflection 3 Pi | Inflection | 🔒 Propriétaire | 52,0 % | 11 octobre 2024 | ✅ Mesuré |
| 120 | Upstage: Solar Pro 3 | Upstage | ▫ n.d. | 52,0 % | 27 janvier 2026 | ✅ Mesuré |
| 121 | hermes-3-llama-3.1-70b | nousresearch | ▫ n.d. | 52,0 % | — | ✅ Mesuré |
| 122 | ministral-14b-2512 | mistralai | ▫ n.d. | 52,0 % | — | ✅ Mesuré |
| 123 | ministral-8b-2512 | mistralai | ▫ n.d. | 52,0 % | — | ✅ Mesuré |
| 124 | MiniMax: MiniMax M2-her | MiniMax | ▫ n.d. | 50,0 % | 23 janvier 2026 | ✅ Mesuré |
| 125 | StepFun: Step 3.7 Flash | StepFun | 🟢 Ouvert | 50,0 % | 28 mai 2026 | ✅ Mesuré |
| 126 | WizardLM-2 8x22B | Microsoft | 🟢 Ouvert | 50,0 % | 16 avril 2024 | ✅ Mesuré |
| 127 | command-r-plus-08-2024 | Cohere | ▫ n.d. | 48,0 % | — | ✅ Mesuré |
| 128 | nova-pro-v1 | Amazon | 🔒 Propriétaire | 48,0 % | — | ✅ Mesuré |
| 129 | IBM: Granite 4.1 8B | IBM | 🟢 Ouvert | 46,0 % | 30 avril 2026 | ✅ Mesuré |
| 130 | Mistral: Codestral 2508 | Mistral AI | ▫ n.d. | 46,0 % | 1 août 2025 | ✅ Mesuré |
| 131 | granite-4.0-h-micro | IBM | 🟢 Ouvert | 44,7 % | 20 octobre 2025 | ✅ Mesuré |
| 132 | nova-2-lite-v1 | Amazon | 🔒 Propriétaire | 42,0 % | 2 décembre 2025 | ✅ Mesuré |
| 133 | inclusionAI: Ling-2.6-flash | InclusionAI | ▫ n.d. | 38,0 % | 21 avril 2026 | ✅ Mesuré |
| 134 | Claude 3 Haiku | Anthropic | 🔒 Propriétaire | 36,0 % | 13 mars 2024 | ✅ Mesuré |
| 135 | ui-tars-1.5-7b | ByteDance | ▫ n.d. | 36,0 % | — | ✅ Mesuré |
| 136 | Cohere: Command R (08-2024) | Cohere | 🟢 Ouvert | 34,0 % | 30 août 2024 | ✅ Mesuré |
| 137 | inclusionAI: Ling-2.6-1T | InclusionAI | ▫ n.d. | 33,3 % | 23 avril 2026 | ✅ Mesuré |
| 138 | l3.1-euryale-70b | sao10k | ▫ n.d. | 30,0 % | — | ✅ Mesuré |
| 139 | Mistral NeMo | Mistral AI | ▫ n.d. | 28,0 % | 18 juillet 2024 | ✅ Mesuré |
| 140 | rocinante-12b | thedrummer | ▫ n.d. | 26,0 % | — | ✅ Mesuré |
| 141 | Cohere: Command R7B (12-2024) | Cohere | ▫ n.d. | 24,5 % | 14 décembre 2024 | ✅ Mesuré |
| 142 | Meta: Llama 3.2 1B Instruct | Meta | 🟢 Ouvert | 22,0 % | 25 septembre 2024 | ✅ Mesuré |
| 143 | l3.3-euryale-70b-v2.3 | sao10k | ▫ n.d. | 22,0 % | — | ✅ Mesuré |
| 144 | ministral-3b-2512 | mistralai | ▫ n.d. | 22,0 % | — | ✅ Mesuré |
| 145 | unslopnemo-12b | thedrummer | ▫ n.d. | 20,4 % | — | ✅ Mesuré |
| 146 | l3-lunaris-8b | sao10k | ▫ n.d. | 20,0 % | — | ✅ Mesuré |
| 147 | reka-edge-2603 | Reka AI | ▫ n.d. | 14,0 % | — | ✅ Mesuré |
| 148 | aion-rp-llama-3.1-8b | Aion Labs | ▫ n.d. | 12,0 % | — | ✅ Mesuré |
| 149 | mythomax-l2-13b | gryphe | ▫ n.d. | 10,0 % | — | ✅ Mesuré |
| 150 | remm-slerp-l2-13b | undi95 | ▫ n.d. | 6,0 % | — | ✅ Mesuré |
| 151 | weaver | mancer | ▫ n.d. | 4,0 % | — | ✅ Mesuré |
| 152 | nova-micro-v1 | Amazon | 🔒 Propriétaire | 2,0 % | — | ✅ Mesuré |
| 153 | Qwen 3.5 Plus | Qwen | ▫ n.d. | 0,0 % | 16 février 2026 | ✅ Mesuré |
| 154 | Qwen: Qwen3.5-Flash | Alibaba Cloud / Qwen Team | ▫ n.d. | 0,0 % | 25 février 2026 | ✅ Mesuré |
| 155 | Qwen: Qwen3.6 Flash | Alibaba Cloud / Qwen Team | ▫ n.d. | 0,0 % | 27 avril 2026 | ✅ Mesuré |
| 156 | qwen3.6-plus-04-02 | Alibaba Cloud / Qwen Team | ▫ n.d. | 0,0 % | 2 avril 2026 | ✅ Mesuré |
Classement établi sur 156 modèles évalués, dont 63 de grands éditeurs. Score médian de l'ensemble : 78,0 %.
Notre analyse
Un score élevé indique qu’un modèle résout correctement une grande proportion de ces tâches ouvertes et restitue la réponse attendue avec une équivalence sémantique suffisante. La validation est contrôlée par une IA sur le champ JSON Path $.answer, ce qui permet de reconnaître des formulations équivalentes sans imposer une correspondance textuelle stricte. Les scores sont au moins partiellement mesurés par un tiers, leur portée ne repose donc pas uniquement sur des résultats auto-déclarés.
Le niveau médian de 80 % parmi 246 modèles montre que l’ensemble est déjà bien maîtrisé par une part importante du classement. Le score de 100 % obtenu par AionLabs: Aion-2.0 signale une saturation au sommet, où ce benchmark ne distingue plus les modèles ayant résolu toutes les questions. Son format limité à 50 tâches en anglais circonscrit aussi l’interprétation aux formes de raisonnement représentées. Enfin, son accès public crée un risque de contamination lors de l’entraînement ou de l’évaluation. Le classement révèle donc surtout des écarts sur ce socle ciblé de logique et de résolution abstraite.
Sources des scores : benchable.