General Knowledge (Baseline)
Benchable : General Knowledge (Baseline) est un benchmark créé par Benchable pour évaluer les connaissances générales des modèles d’IA. Il repose sur des questions à choix multiple en anglais, dont la difficulté progresse de sujets accessibles vers des connaissances très obscures.
Benchable : General Knowledge (Baseline) est un benchmark créé par Benchable pour évaluer les connaissances générales des modèles d’IA. Il repose sur des questions à choix multiple en anglais, dont la difficulté progresse de sujets accessibles vers des connaissances très obscures.
Son périmètre couvre notamment l’histoire, les sciences, la géographie, les arts, la littérature, l’actualité et des domaines académiques spécialisés. Il sert de test de référence pour comparer la capacité des modèles à identifier précisément une réponse correcte parmi quatre options.
Carte d'identité
| Caractéristique | Valeur |
|---|---|
| Éditeur du benchmark | Benchable |
| Capacités mesurées | Connaissances generales (histoire, science, geographie, arts, litterature, actualite, domaines academiques specialises) du niveau facile a tres obscur |
| Modalité | Texte |
| Type de questions | QCM (4 options A/B/C/D) |
| Métrique d'évaluation | Option correcte selectionnee (Exact Match, JSON Path $.answer) |
| Accès | Public |
| Langues | anglais |
| Taille du jeu | 200 questions |
| Ressources | Site / dépôt officiel |
Classement des modèles (162)
| # | Modèle | Éditeur | Licence | Score | Sortie | Fiabilité |
|---|---|---|---|---|---|---|
| 1 | AionLabs: Aion-3.0 | Aion Labs | ▫ n.d. | 100,0 % | 7 juillet 2026 | ✅ Mesuré |
| 2 | AionLabs: Aion-3.0-Mini | Aion Labs | ▫ n.d. | 100,0 % | 7 juillet 2026 | ✅ Mesuré |
| 3 | ByteDance Seed: Seed-2.0-Mini | ByteDance Seed | ▫ n.d. | 100,0 % | 26 février 2026 | ✅ Mesuré |
| 4 | Claude Opus 4 | Anthropic | 🔒 Propriétaire | 100,0 % | 22 mai 2025 | ✅ Mesuré |
| 5 | Claude Opus 4.1 | Anthropic | 🔒 Propriétaire | 100,0 % | 5 août 2025 | ✅ Mesuré |
| 6 | Claude Opus 4.5 | Anthropic | 🔒 Propriétaire | 100,0 % | 24 novembre 2025 | ✅ Mesuré |
| 7 | Claude Sonnet 4.5 | Anthropic | 🔒 Propriétaire | 100,0 % | 29 septembre 2025 | ✅ Mesuré |
| 8 | GPT-5 | OpenAI | 🔒 Propriétaire | 100,0 % | 7 août 2025 | ✅ Mesuré |
| 9 | GPT-5 mini | OpenAI | 🔒 Propriétaire | 100,0 % | 7 août 2025 | ✅ Mesuré |
| 10 | GPT-5 nano | OpenAI | 🔒 Propriétaire | 100,0 % | 7 août 2025 | ✅ Mesuré |
| 11 | GPT-5.1 | OpenAI | 🔒 Propriétaire | 100,0 % | 13 novembre 2025 | ✅ Mesuré |
| 12 | GPT-5.2 | OpenAI | 🔒 Propriétaire | 100,0 % | 11 décembre 2025 | ✅ Mesuré |
| 13 | GPT-5.6 Luna | OpenAI | 🔒 Propriétaire | 100,0 % | 9 juillet 2026 | ✅ Mesuré |
| 14 | GPT-5.6 Sol | OpenAI | 🔒 Propriétaire | 100,0 % | 9 juillet 2026 | ✅ Mesuré |
| 15 | GPT-5.6 Terra | OpenAI | 🔒 Propriétaire | 100,0 % | 9 juillet 2026 | ✅ Mesuré |
| 16 | Kwaipilot: KAT-Coder-Air V2.5 | kwaipilot | ▫ n.d. | 100,0 % | 10 juillet 2026 | ✅ Mesuré |
| 17 | Kwaipilot: KAT-Coder-Pro V2 | kwaipilot | ▫ n.d. | 100,0 % | 27 mars 2026 | ✅ Mesuré |
| 18 | Kwaipilot: KAT-Coder-Pro V2.5 | kwaipilot | ▫ n.d. | 100,0 % | 10 juillet 2026 | ✅ Mesuré |
| 19 | Mistral: Mistral Medium 3 | Mistral AI | ▫ n.d. | 100,0 % | 7 mai 2025 | ✅ Mesuré |
| 20 | Mistral: Mistral Medium 3.1 | Mistral AI | ▫ n.d. | 100,0 % | 13 août 2025 | ✅ Mesuré |
| 21 | Nex AGI: Nex-N2-Pro | Nex AGI | 🟢 Ouvert | 100,0 % | 8 juin 2026 | ✅ Mesuré |
| 22 | OpenAI: GPT Chat Latest | OpenAI | 🔒 Propriétaire | 100,0 % | 5 mai 2026 | ✅ Mesuré |
| 23 | OpenAI: GPT-5.1 Chat | OpenAI | 🔒 Propriétaire | 100,0 % | 13 novembre 2025 | ✅ Mesuré |
| 24 | OpenAI: GPT-5.1-Codex-Max | OpenAI | 🔒 Propriétaire | 100,0 % | 4 décembre 2025 | ✅ Mesuré |
| 25 | OpenAI: GPT-5.2 Chat | OpenAI | 🔒 Propriétaire | 100,0 % | 10 décembre 2025 | ✅ Mesuré |
| 26 | OpenAI: GPT-5.4 Image 2 | OpenAI | 🔒 Propriétaire | 100,0 % | 21 avril 2026 | ✅ Mesuré |
| 27 | OpenAI: GPT-5.6 Luna Pro | OpenAI | 🔒 Propriétaire | 100,0 % | 9 juillet 2026 | ✅ Mesuré |
| 28 | OpenAI: GPT-5.6 Sol Pro | OpenAI | 🔒 Propriétaire | 100,0 % | 9 juillet 2026 | ✅ Mesuré |
| 29 | OpenAI: GPT-5.6 Terra Pro | OpenAI | 🔒 Propriétaire | 100,0 % | 9 juillet 2026 | ✅ Mesuré |
| 30 | Sonar | Perplexity | 🔒 Propriétaire | 100,0 % | 29 janvier 2025 | ✅ Mesuré |
| 31 | Sonar Pro | Perplexity | 🔒 Propriétaire | 100,0 % | 7 mars 2025 | ✅ Mesuré |
| 32 | StepFun: Step 3.7 Flash | StepFun | 🟢 Ouvert | 100,0 % | 28 mai 2026 | ✅ Mesuré |
| 33 | Tencent: Hy3 preview | Tencent | 🟢 Ouvert | 100,0 % | 22 avril 2026 | ✅ Mesuré |
| 34 | Thinking Machines: Inkling | Thinking Machines | 🟢 Ouvert | 100,0 % | 17 juillet 2026 | ✅ Mesuré |
| 35 | Z.ai: GLM 5 Turbo | Zhipu AI | ▫ n.d. | 100,0 % | 15 mars 2026 | ✅ Mesuré |
| 36 | gemini-3-pro-image | ▫ n.d. | 100,0 % | — | ✅ Mesuré | |
| 37 | kimi-k2.5-0127 | Moonshot AI | ▫ n.d. | 100,0 % | — | ✅ Mesuré |
| 38 | laguna-m.1 | poolside | ▫ n.d. | 100,0 % | — | ✅ Mesuré |
| 39 | mistral-large-2512 | mistralai | ▫ n.d. | 100,0 % | — | ✅ Mesuré |
| 40 | qwen3-235b-a22b-04-28 | Alibaba Cloud / Qwen Team | ▫ n.d. | 100,0 % | — | ✅ Mesuré |
| 41 | qwen3-30b-a3b-04-28 | Alibaba Cloud / Qwen Team | ▫ n.d. | 100,0 % | — | ✅ Mesuré |
| 42 | qwen3-32b-04-28 | Alibaba Cloud / Qwen Team | ▫ n.d. | 100,0 % | — | ✅ Mesuré |
| 43 | qwen3-next-80b-a3b-instruct-2509 | Alibaba Cloud / Qwen Team | ▫ n.d. | 100,0 % | — | ✅ Mesuré |
| 44 | Claude Sonnet 4 | Anthropic | 🔒 Propriétaire | 99,8 % | 22 mai 2025 | ✅ Mesuré |
| 45 | MiniMax M1 | MiniMax | 🟢 Ouvert | 99,8 % | 17 juin 2025 | ✅ Mesuré |
| 46 | gemini-2.5-pro-preview-03-25 | ▫ n.d. | 99,8 % | — | ✅ Mesuré | |
| 47 | o1 | OpenAI | 🔒 Propriétaire | 99,8 % | 17 décembre 2024 | ✅ Mesuré |
| 48 | AionLabs: Aion-2.0 | Aion Labs | ▫ n.d. | 99,5 % | 23 février 2026 | ✅ Mesuré |
| 49 | Arcee AI: Virtuoso Large | Arcee AI | ▫ n.d. | 99,5 % | 5 mai 2025 | ✅ Mesuré |
| 50 | Baidu: ERNIE 4.5 VL 424B A47B | Baidu | 🟢 Ouvert | 99,5 % | 30 juin 2025 | ✅ Mesuré |
| 51 | DeepSeek V3.1 Terminus | DeepSeek | 🟢 Ouvert | 99,5 % | 22 septembre 2025 | ✅ Mesuré |
| 52 | DeepSeek V4 Flash | DeepSeek | ▫ n.d. | 99,5 % | 24 avril 2026 | ✅ Mesuré |
| 53 | GPT-4.1 | OpenAI | 🔒 Propriétaire | 99,5 % | 14 avril 2025 | ✅ Mesuré |
| 54 | GPT-4o | OpenAI | 🔒 Propriétaire | 99,5 % | 27 mars 2025 | ✅ Mesuré |
| 55 | GPT-4o mini | OpenAI | 🔒 Propriétaire | 99,5 % | 18 juillet 2024 | ✅ Mesuré |
| 56 | Gemini 3.1 Pro Preview | ▫ n.d. | 99,5 % | 19 février 2026 | ✅ Mesuré | |
| 57 | Google: Gemini 3.1 Pro Preview Custom Tools | ▫ n.d. | 99,5 % | 25 février 2026 | ✅ Mesuré | |
| 58 | Kimi K2 | Moonshot AI | ▫ n.d. | 99,5 % | 6 novembre 2025 | ✅ Mesuré |
| 59 | Mistral Large 2407 | Mistral AI | ▫ n.d. | 99,5 % | 19 novembre 2024 | ✅ Mesuré |
| 60 | Nous: Hermes 4 405B | Nous Research | 🟢 Ouvert | 99,5 % | 26 août 2025 | ✅ Mesuré |
| 61 | Qwen: Qwen3 30B A3B Thinking 2507 | Alibaba Cloud / Qwen Team | 🟢 Ouvert | 99,5 % | 28 août 2025 | ✅ Mesuré |
| 62 | Seed 1.6 | ByteDance Seed | ▫ n.d. | 99,5 % | 23 décembre 2025 | ✅ Mesuré |
| 63 | deepseek-chat-v3 | DeepSeek | ▫ n.d. | 99,5 % | — | ✅ Mesuré |
| 64 | deepseek-chat-v3-0324 | DeepSeek | ▫ n.d. | 99,5 % | — | ✅ Mesuré |
| 65 | deepseek-chat-v3.1 | DeepSeek | ▫ n.d. | 99,5 % | — | ✅ Mesuré |
| 66 | gpt-5-chat-2025-08-07 | OpenAI | 🔒 Propriétaire | 99,5 % | 7 août 2025 | ✅ Mesuré |
| 67 | laguna-xs-2.1 | poolside | ▫ n.d. | 99,5 % | — | ✅ Mesuré |
| 68 | nemotron-nano-12b-v2-vl | NVIDIA | 🟢 Ouvert | 99,5 % | 28 octobre 2025 | ✅ Mesuré |
| 69 | qwen-plus-2025-01-25 | Alibaba Cloud / Qwen Team | ▫ n.d. | 99,5 % | 8 septembre 2025 | ✅ Mesuré |
| 70 | qwen3-235b-a22b-07-25 | Alibaba Cloud / Qwen Team | ▫ n.d. | 99,5 % | — | ✅ Mesuré |
| 71 | qwen3-next-80b-a3b-thinking-2509 | Alibaba Cloud / Qwen Team | ▫ n.d. | 99,5 % | — | ✅ Mesuré |
| 72 | xAI: Grok Build 0.1 | xAI | ▫ n.d. | 99,5 % | 20 mai 2026 | ✅ Mesuré |
| 73 | Writer: Palmyra X5 | Writer | ▫ n.d. | 99,4 % | 21 janvier 2026 | ✅ Mesuré |
| 74 | llama-4-maverick-17b-128e-instruct | Meta | ▫ n.d. | 99,2 % | — | ✅ Mesuré |
| 75 | inclusionAI: Ling-2.6-1T | InclusionAI | ▫ n.d. | 99,1 % | 23 avril 2026 | ✅ Mesuré |
| 76 | Claude Haiku 4.5 | Anthropic | 🔒 Propriétaire | 99,0 % | 15 octobre 2025 | ✅ Mesuré |
| 77 | GPT-4.1 mini | OpenAI | 🔒 Propriétaire | 99,0 % | 14 avril 2025 | ✅ Mesuré |
| 78 | MiniMax: MiniMax-01 | MiniMax | 🟢 Ouvert | 99,0 % | 15 janvier 2025 | ✅ Mesuré |
| 79 | gemini-3.1-flash-image-preview | ▫ n.d. | 99,0 % | — | ✅ Mesuré | |
| 80 | gpt-4o-mini-search-preview-2025-03-11 | OpenAI | 🔒 Propriétaire | 99,0 % | 12 mars 2025 | ✅ Mesuré |
| 81 | olmo-3-32b-think | allenai | ▫ n.d. | 99,0 % | — | ✅ Mesuré |
| 82 | qwen3-coder-next-2025-02-03 | Alibaba Cloud / Qwen Team | 🟢 Ouvert | 99,0 % | 4 février 2026 | ✅ Mesuré |
| 83 | uncensored | venice | ▫ n.d. | 99,0 % | — | ✅ Mesuré |
| 84 | qwen3-14b-04-28 | Alibaba Cloud / Qwen Team | ▫ n.d. | 98,8 % | — | ✅ Mesuré |
| 85 | Magnum v4 72B | anthracite-org | 🟢 Ouvert | 98,5 % | 22 octobre 2024 | ✅ Mesuré |
| 86 | Mistral Large | Mistral AI | ▫ n.d. | 98,5 % | 26 février 2024 | ✅ Mesuré |
| 87 | Mistral: Voxtral Small 24B 2507 | Mistral AI | 🟢 Ouvert | 98,5 % | 30 octobre 2025 | ✅ Mesuré |
| 88 | devstral-2512 | mistralai | ▫ n.d. | 98,5 % | — | ✅ Mesuré |
| 89 | hermes-3-llama-3.1-405b | nousresearch | ▫ n.d. | 98,5 % | — | ✅ Mesuré |
| 90 | ministral-14b-2512 | mistralai | ▫ n.d. | 98,5 % | — | ✅ Mesuré |
| 91 | DeepSeek V4 Pro | DeepSeek | ▫ n.d. | 98,5 % | 24 avril 2026 | ✅ Mesuré |
| 92 | Deep Cogito: Cogito v2.1 671B | Deep Cogito | ▫ n.d. | 98,0 % | 13 novembre 2025 | ✅ Mesuré |
| 93 | Perplexity: Sonar Pro Search | Perplexity | 🔒 Propriétaire | 98,0 % | 30 octobre 2025 | ✅ Mesuré |
| 94 | mistral-small-24b-instruct-2501 | mistralai | ▫ n.d. | 98,0 % | — | ✅ Mesuré |
| 95 | nova-pro-v1 | Amazon | 🔒 Propriétaire | 98,0 % | — | ✅ Mesuré |
| 96 | mistral-saba-2502 | mistralai | ▫ n.d. | 97,8 % | — | ✅ Mesuré |
| 97 | inclusionAI: Ring-2.6-1T | InclusionAI | ▫ n.d. | 97,6 % | 8 mai 2026 | ✅ Mesuré |
| 98 | Tencent: Hunyuan A13B Instruct | Tencent | 🟢 Ouvert | 97,5 % | 8 juillet 2025 | ✅ Mesuré |
| 99 | nova-2-lite-v1 | Amazon | 🔒 Propriétaire | 97,5 % | 2 décembre 2025 | ✅ Mesuré |
| 100 | qwen3-8b-04-28 | Alibaba Cloud / Qwen Team | ▫ n.d. | 97,5 % | — | ✅ Mesuré |
| 101 | skyfall-36b-v2 | thedrummer | ▫ n.d. | 97,5 % | — | ✅ Mesuré |
| 102 | hermes-3-llama-3.1-70b | nousresearch | ▫ n.d. | 97,2 % | — | ✅ Mesuré |
| 103 | OpenAI: gpt-oss-safeguard-20b | OpenAI | 🟢 Ouvert | 97,0 % | 29 octobre 2025 | ✅ Mesuré |
| 104 | Qwen: Qwen3 30B A3B Instruct 2507 | Alibaba Cloud / Qwen Team | 🟢 Ouvert | 97,0 % | 29 juillet 2025 | ✅ Mesuré |
| 105 | llama-4-scout-17b-16e-instruct | Meta | ▫ n.d. | 97,0 % | — | ✅ Mesuré |
| 106 | mistral-small-2603 | mistralai | ▫ n.d. | 97,0 % | — | ✅ Mesuré |
| 107 | Arcee AI: Trinity Large Thinking | Arcee AI | 🟢 Ouvert | 96,5 % | 1 avril 2026 | ✅ Mesuré |
| 108 | Claude 3 Haiku | Anthropic | 🔒 Propriétaire | 96,5 % | 13 mars 2024 | ✅ Mesuré |
| 109 | Command A+ | Cohere | 🟢 Ouvert | 96,5 % | 20 mai 2026 | ✅ Mesuré |
| 110 | GPT-4.1 nano | OpenAI | 🔒 Propriétaire | 96,5 % | 14 avril 2025 | ✅ Mesuré |
| 111 | Inflection: Inflection 3 Productivity | Inflection | 🔒 Propriétaire | 96,5 % | 11 octobre 2024 | ✅ Mesuré |
| 112 | Nous: Hermes 4 70B | Nous Research | 🟢 Ouvert | 96,5 % | 26 août 2025 | ✅ Mesuré |
| 113 | Reka Flash 3 | Reka AI | 🟢 Ouvert | 96,0 % | 12 mars 2025 | ✅ Mesuré |
| 114 | gemini-3.1-flash-image | ▫ n.d. | 96,0 % | — | ✅ Mesuré | |
| 115 | gpt-4o-search-preview-2025-03-11 | OpenAI | 🔒 Propriétaire | 96,0 % | 12 mars 2025 | ✅ Mesuré |
| 116 | ministral-8b-2512 | mistralai | ▫ n.d. | 96,0 % | — | ✅ Mesuré |
| 117 | cydonia-24b-v4.1 | thedrummer | ▫ n.d. | 95,7 % | — | ✅ Mesuré |
| 118 | Mistral: Mixtral 8x22B Instruct | Mistral AI | 🟢 Ouvert | 95,5 % | 17 avril 2024 | ✅ Mesuré |
| 119 | OpenAI: GPT-3.5 Turbo 16k | OpenAI | 🔒 Propriétaire | 95,5 % | 28 août 2023 | ✅ Mesuré |
| 120 | Upstage: Solar Pro 3 | Upstage | ▫ n.d. | 95,5 % | 27 janvier 2026 | ✅ Mesuré |
| 121 | WizardLM-2 8x22B | Microsoft | 🟢 Ouvert | 95,5 % | 16 avril 2024 | ✅ Mesuré |
| 122 | nova-micro-v1 | Amazon | 🔒 Propriétaire | 95,0 % | — | ✅ Mesuré |
| 123 | command-r-plus-08-2024 | Cohere | ▫ n.d. | 94,5 % | — | ✅ Mesuré |
| 124 | ministral-3b-2512 | mistralai | ▫ n.d. | 94,0 % | — | ✅ Mesuré |
| 125 | OpenAI: GPT-3.5 Turbo Instruct | OpenAI | 🔒 Propriétaire | 93,5 % | 28 septembre 2023 | ✅ Mesuré |
| 126 | AI21: Jamba Large 1.7 | AI21 Labs | 🟢 Ouvert | 93,0 % | 8 août 2025 | ✅ Mesuré |
| 127 | Cohere: Command R (08-2024) | Cohere | 🟢 Ouvert | 93,0 % | 30 août 2024 | ✅ Mesuré |
| 128 | IBM: Granite 4.1 8B | IBM | 🟢 Ouvert | 93,0 % | 30 avril 2026 | ✅ Mesuré |
| 129 | l3.1-euryale-70b | sao10k | ▫ n.d. | 93,0 % | — | ✅ Mesuré |
| 130 | ui-tars-1.5-7b | ByteDance | ▫ n.d. | 93,0 % | — | ✅ Mesuré |
| 131 | Nex AGI: Nex-N2-Mini | Nex AGI | 🟢 Ouvert | 92,5 % | 24 juin 2026 | ✅ Mesuré |
| 132 | rocinante-12b | thedrummer | ▫ n.d. | 92,0 % | — | ✅ Mesuré |
| 133 | granite-4.0-h-micro | IBM | 🟢 Ouvert | 91,4 % | 20 octobre 2025 | ✅ Mesuré |
| 134 | ByteDance Seed: Seed 1.6 Flash | ByteDance Seed | ▫ n.d. | 90,5 % | 23 décembre 2025 | ✅ Mesuré |
| 135 | Inflection: Inflection 3 Pi | Inflection | 🔒 Propriétaire | 88,5 % | 11 octobre 2024 | ✅ Mesuré |
| 136 | Cohere: Command R7B (12-2024) | Cohere | ▫ n.d. | 87,8 % | 14 décembre 2024 | ✅ Mesuré |
| 137 | unslopnemo-12b | thedrummer | ▫ n.d. | 87,5 % | — | ✅ Mesuré |
| 138 | Mistral NeMo | Mistral AI | ▫ n.d. | 84,2 % | 18 juillet 2024 | ✅ Mesuré |
| 139 | l3.3-euryale-70b-v2.3 | sao10k | ▫ n.d. | 84,0 % | — | ✅ Mesuré |
| 140 | Perceptron: Perceptron Mk1 | Perceptron | ▫ n.d. | 83,8 % | 12 mai 2026 | ✅ Mesuré |
| 141 | Qwen: Qwen3 Coder Flash | Alibaba Cloud / Qwen Team | ▫ n.d. | 81,5 % | 17 septembre 2025 | ✅ Mesuré |
| 142 | Relace: Relace Search | relace | ▫ n.d. | 79,9 % | 8 décembre 2025 | ✅ Mesuré |
| 143 | reka-edge-2603 | Reka AI | ▫ n.d. | 70,0 % | — | ✅ Mesuré |
| 144 | l3-lunaris-8b | sao10k | ▫ n.d. | 69,0 % | — | ✅ Mesuré |
| 145 | MiniMax: MiniMax M2-her | MiniMax | ▫ n.d. | 65,5 % | 23 janvier 2026 | ✅ Mesuré |
| 146 | aion-rp-llama-3.1-8b | Aion Labs | ▫ n.d. | 58,5 % | — | ✅ Mesuré |
| 147 | inclusionAI: Ling-2.6-flash | InclusionAI | ▫ n.d. | 52,5 % | 21 avril 2026 | ✅ Mesuré |
| 148 | Mistral: Codestral 2508 | Mistral AI | ▫ n.d. | 49,5 % | 1 août 2025 | ✅ Mesuré |
| 149 | Hy3 | Tencent | 🟢 Ouvert | 36,0 % | 6 juillet 2026 | ✅ Mesuré |
| 150 | Qwen: Qwen3 Coder 30B A3B Instruct | Alibaba Cloud / Qwen Team | 🟢 Ouvert | 32,2 % | 31 juillet 2025 | ✅ Mesuré |
| 151 | remm-slerp-l2-13b | undi95 | ▫ n.d. | 0,5 % | — | ✅ Mesuré |
| 152 | weaver | mancer | ▫ n.d. | 0,5 % | — | ✅ Mesuré |
| 153 | mythomax-l2-13b | gryphe | ▫ n.d. | 0,2 % | — | ✅ Mesuré |
| 154 | Meta: Llama 3.2 1B Instruct | Meta | 🟢 Ouvert | 0,0 % | 25 septembre 2024 | ✅ Mesuré |
| 155 | Meta: Llama Guard 4 12B | Meta | 🟢 Ouvert | 0,0 % | 30 avril 2025 | ✅ Mesuré |
| 156 | OpenAI: GPT-4 Turbo Preview | OpenAI | 🔒 Propriétaire | 0,0 % | 25 janvier 2024 | ✅ Mesuré |
| 157 | Qwen 3.5 Plus | Qwen | ▫ n.d. | 0,0 % | 16 février 2026 | ✅ Mesuré |
| 158 | Qwen: Qwen3.5-Flash | Alibaba Cloud / Qwen Team | ▫ n.d. | 0,0 % | 25 février 2026 | ✅ Mesuré |
| 159 | Qwen: Qwen3.6 Flash | Alibaba Cloud / Qwen Team | ▫ n.d. | 0,0 % | 27 avril 2026 | ✅ Mesuré |
| 160 | gpt-3.5-turbo-0613 | OpenAI | 🔒 Propriétaire | 0,0 % | — | ✅ Mesuré |
| 161 | nova-premier-v1 | Amazon | 🔒 Propriétaire | 0,0 % | — | ✅ Mesuré |
| 162 | qwen3.6-plus-04-02 | Alibaba Cloud / Qwen Team | ▫ n.d. | 0,0 % | 2 avril 2026 | ✅ Mesuré |
Classement établi sur 162 modèles évalués, dont 69 de grands éditeurs. Score médian de l'ensemble : 99,0 %.
Notre analyse
Un score élevé indique qu’un modèle sélectionne très fréquemment l’option attendue dans ce format de QCM. La validation par Exact Match sur le champ JSON $.answer rend la notation stricte et reproductible, sans appréciation subjective de la formulation. Les scores étant au moins partiellement mesurés par un tiers, ils bénéficient d’une base de vérification externe plutôt que de reposer uniquement sur des déclarations des fournisseurs.
Le score médian de 99 % parmi 255 modèles et le résultat de 100 % obtenu par AionLabs: Aion-3.0 montrent toutefois une forte saturation. Le classement distingue donc difficilement les modèles les plus performants, quelques réponses seulement pouvant séparer des résultats presque identiques. Le caractère public du jeu invite aussi à considérer un risque de contamination lors de l’interprétation. Enfin, ce benchmark mesure avant tout la restitution de connaissances générales en anglais dans un cadre fermé à quatre choix. Un excellent résultat ne suffit donc pas à établir des performances équivalentes sur des tâches ouvertes ou sur d’autres capacités. Le classement révèle surtout une maîtrise désormais très élevée de ce format par une large partie des modèles évalués.
Sources des scores : benchable.