Coding (Baseline)
Benchable : Coding (Baseline) est un benchmark public créé par Benchable pour tester les connaissances générales en programmation. Il prend la forme de questions à choix multiple en anglais, avec six réponses possibles, et couvre plusieurs langages ainsi que différents niveaux de…
Benchable : Coding (Baseline) est un benchmark public créé par Benchable pour tester les connaissances générales en programmation. Il prend la forme de questions à choix multiple en anglais, avec six réponses possibles, et couvre plusieurs langages ainsi que différents niveaux de difficulté.
Son contenu va de la syntaxe élémentaire aux algorithmes, structures de données, bases de données, réseaux, design patterns et mécanismes de concurrence. Il sert à comparer la capacité des modèles à reconnaître précisément une réponse correcte dans un large éventail de sujets techniques.
Carte d'identité
| Caractéristique | Valeur |
|---|---|
| Éditeur du benchmark | Benchable |
| Capacités mesurées | Connaissances en programmation : de la syntaxe de base aux algorithmes, structures de donnees, concurrence, bases de donnees, design patterns, reseaux |
| Modalité | Texte |
| Type de questions | QCM (6 options A-F) |
| Métrique d'évaluation | Lettre de l'option correcte (Exact Match, JSON Path $.answer) |
| Accès | Public |
| Langues | anglais (enonces) ; couvre Python, JavaScript, Java, C++, Go, SQL, PHP, Ruby, Rust, TypeScript, C#, Scala, Haskell |
| Taille du jeu | 100 questions |
| Ressources | Site / dépôt officiel |
Classement des modèles (163)
| # | Modèle | Éditeur | Licence | Score | Sortie | Fiabilité |
|---|---|---|---|---|---|---|
| 1 | StepFun: Step 3.7 Flash | StepFun | 🟢 Ouvert | 100,0 % | 28 mai 2026 | ✅ Mesuré |
| 2 | cydonia-24b-v4.1 | thedrummer | ▫ n.d. | 100,0 % | — | ✅ Mesuré |
| 3 | gemini-3-pro-image | ▫ n.d. | 97,0 % | — | ✅ Mesuré | |
| 4 | qwen3-235b-a22b-04-28 | Alibaba Cloud / Qwen Team | ▫ n.d. | 97,0 % | — | ✅ Mesuré |
| 5 | Google: Gemini 3.1 Pro Preview Custom Tools | ▫ n.d. | 96,9 % | 25 février 2026 | ✅ Mesuré | |
| 6 | Kwaipilot: KAT-Coder-Pro V2.5 | kwaipilot | ▫ n.d. | 96,1 % | 10 juillet 2026 | ✅ Mesuré |
| 7 | GPT-5.2 | OpenAI | 🔒 Propriétaire | 96,0 % | 11 décembre 2025 | ✅ Mesuré |
| 8 | Grok 4.5 | xAI | 🔒 Propriétaire | 96,0 % | 16 juillet 2026 | ✅ Mesuré |
| 9 | OpenAI: GPT Chat Latest | OpenAI | 🔒 Propriétaire | 96,0 % | 5 mai 2026 | ✅ Mesuré |
| 10 | OpenAI: GPT-5.2 Chat | OpenAI | 🔒 Propriétaire | 96,0 % | 10 décembre 2025 | ✅ Mesuré |
| 11 | gemini-2.5-pro-preview-03-25 | ▫ n.d. | 96,0 % | — | ✅ Mesuré | |
| 12 | gemini-3.1-flash-image-preview | ▫ n.d. | 96,0 % | — | ✅ Mesuré | |
| 13 | AionLabs: Aion-3.0 | Aion Labs | ▫ n.d. | 96,0 % | 7 juillet 2026 | ✅ Mesuré |
| 14 | Gemini 3.1 Pro Preview | ▫ n.d. | 96,0 % | 19 février 2026 | ✅ Mesuré | |
| 15 | Z.ai: GLM 5 Turbo | Zhipu AI | ▫ n.d. | 95,5 % | 15 mars 2026 | ✅ Mesuré |
| 16 | DeepSeek V4 Pro | DeepSeek | ▫ n.d. | 95,5 % | 24 avril 2026 | ✅ Mesuré |
| 17 | AionLabs: Aion-3.0-Mini | Aion Labs | ▫ n.d. | 95,0 % | 7 juillet 2026 | ✅ Mesuré |
| 18 | Claude Opus 4.5 | Anthropic | 🔒 Propriétaire | 95,0 % | 24 novembre 2025 | ✅ Mesuré |
| 19 | Claude Sonnet 4.5 | Anthropic | 🔒 Propriétaire | 95,0 % | 29 septembre 2025 | ✅ Mesuré |
| 20 | GPT-5 mini | OpenAI | 🔒 Propriétaire | 95,0 % | 7 août 2025 | ✅ Mesuré |
| 21 | GPT-5.6 Luna | OpenAI | 🔒 Propriétaire | 95,0 % | 9 juillet 2026 | ✅ Mesuré |
| 22 | GPT-5.6 Sol | OpenAI | 🔒 Propriétaire | 95,0 % | 9 juillet 2026 | ✅ Mesuré |
| 23 | GPT-5.6 Terra | OpenAI | 🔒 Propriétaire | 95,0 % | 9 juillet 2026 | ✅ Mesuré |
| 24 | OpenAI: GPT-5.1-Codex-Max | OpenAI | 🔒 Propriétaire | 95,0 % | 4 décembre 2025 | ✅ Mesuré |
| 25 | OpenAI: GPT-5.6 Luna Pro | OpenAI | 🔒 Propriétaire | 95,0 % | 9 juillet 2026 | ✅ Mesuré |
| 26 | OpenAI: GPT-5.6 Sol Pro | OpenAI | 🔒 Propriétaire | 95,0 % | 9 juillet 2026 | ✅ Mesuré |
| 27 | OpenAI: GPT-5.6 Terra Pro | OpenAI | 🔒 Propriétaire | 95,0 % | 9 juillet 2026 | ✅ Mesuré |
| 28 | OpenAI: gpt-oss-safeguard-20b | OpenAI | 🟢 Ouvert | 95,0 % | 29 octobre 2025 | ✅ Mesuré |
| 29 | qwen3-32b-04-28 | Alibaba Cloud / Qwen Team | ▫ n.d. | 95,0 % | — | ✅ Mesuré |
| 30 | xAI: Grok Build 0.1 | xAI | ▫ n.d. | 95,0 % | 20 mai 2026 | ✅ Mesuré |
| 31 | qwen3-next-80b-a3b-thinking-2509 | Alibaba Cloud / Qwen Team | ▫ n.d. | 94,9 % | — | ✅ Mesuré |
| 32 | ByteDance Seed: Seed-2.0-Mini | ByteDance Seed | ▫ n.d. | 94,0 % | 26 février 2026 | ✅ Mesuré |
| 33 | Claude Opus 4 | Anthropic | 🔒 Propriétaire | 94,0 % | 22 mai 2025 | ✅ Mesuré |
| 34 | Claude Opus 4.1 | Anthropic | 🔒 Propriétaire | 94,0 % | 5 août 2025 | ✅ Mesuré |
| 35 | DeepSeek V4 Flash | DeepSeek | ▫ n.d. | 94,0 % | 24 avril 2026 | ✅ Mesuré |
| 36 | GPT-5.1 | OpenAI | 🔒 Propriétaire | 94,0 % | 13 novembre 2025 | ✅ Mesuré |
| 37 | OpenAI: GPT-5.1 Chat | OpenAI | 🔒 Propriétaire | 94,0 % | 13 novembre 2025 | ✅ Mesuré |
| 38 | Seed 1.6 | ByteDance Seed | ▫ n.d. | 94,0 % | 23 décembre 2025 | ✅ Mesuré |
| 39 | gemini-3.1-flash-image | ▫ n.d. | 94,0 % | — | ✅ Mesuré | |
| 40 | qwen3-30b-a3b-04-28 | Alibaba Cloud / Qwen Team | ▫ n.d. | 94,0 % | — | ✅ Mesuré |
| 41 | AionLabs: Aion-2.0 | Aion Labs | ▫ n.d. | 93,5 % | 23 février 2026 | ✅ Mesuré |
| 42 | Claude Sonnet 4 | Anthropic | 🔒 Propriétaire | 93,0 % | 22 mai 2025 | ✅ Mesuré |
| 43 | GPT-4o | OpenAI | 🔒 Propriétaire | 93,0 % | 27 mars 2025 | ✅ Mesuré |
| 44 | GPT-5 | OpenAI | 🔒 Propriétaire | 93,0 % | 7 août 2025 | ✅ Mesuré |
| 45 | GPT-5 nano | OpenAI | 🔒 Propriétaire | 93,0 % | 7 août 2025 | ✅ Mesuré |
| 46 | Kimi K2 | Moonshot AI | ▫ n.d. | 93,0 % | 6 novembre 2025 | ✅ Mesuré |
| 47 | Nex AGI: Nex-N2-Pro | Nex AGI | 🟢 Ouvert | 93,0 % | 8 juin 2026 | ✅ Mesuré |
| 48 | Qwen: Qwen3 30B A3B Instruct 2507 | Alibaba Cloud / Qwen Team | 🟢 Ouvert | 93,0 % | 29 juillet 2025 | ✅ Mesuré |
| 49 | Tencent: Hy3 preview | Tencent | 🟢 Ouvert | 93,0 % | 22 avril 2026 | ✅ Mesuré |
| 50 | gpt-5-chat-2025-08-07 | OpenAI | 🔒 Propriétaire | 93,0 % | 7 août 2025 | ✅ Mesuré |
| 51 | mistral-small-2603 | mistralai | ▫ n.d. | 93,0 % | — | ✅ Mesuré |
| 52 | Arcee AI: Trinity Large Thinking | Arcee AI | 🟢 Ouvert | 92,9 % | 1 avril 2026 | ✅ Mesuré |
| 53 | Qwen: Qwen3 30B A3B Thinking 2507 | Alibaba Cloud / Qwen Team | 🟢 Ouvert | 92,9 % | 28 août 2025 | ✅ Mesuré |
| 54 | Hy3 | Tencent | 🟢 Ouvert | 92,9 % | 6 juillet 2026 | ✅ Mesuré |
| 55 | Thinking Machines: Inkling | Thinking Machines | 🟢 Ouvert | 92,2 % | 17 juillet 2026 | ✅ Mesuré |
| 56 | GPT-4.1 mini | OpenAI | 🔒 Propriétaire | 92,0 % | 14 avril 2025 | ✅ Mesuré |
| 57 | Mistral: Mistral Medium 3.1 | Mistral AI | ▫ n.d. | 92,0 % | 13 août 2025 | ✅ Mesuré |
| 58 | Perplexity: Sonar Pro Search | Perplexity | 🔒 Propriétaire | 92,0 % | 30 octobre 2025 | ✅ Mesuré |
| 59 | Qwen: Qwen3 Coder Flash | Alibaba Cloud / Qwen Team | ▫ n.d. | 92,0 % | 17 septembre 2025 | ✅ Mesuré |
| 60 | deepseek-chat-v3-0324 | DeepSeek | ▫ n.d. | 92,0 % | — | ✅ Mesuré |
| 61 | laguna-xs-2.1 | poolside | ▫ n.d. | 92,0 % | — | ✅ Mesuré |
| 62 | o1 | OpenAI | 🔒 Propriétaire | 92,0 % | 17 décembre 2024 | ✅ Mesuré |
| 63 | olmo-3-32b-think | allenai | ▫ n.d. | 92,0 % | — | ✅ Mesuré |
| 64 | qwen3-next-80b-a3b-instruct-2509 | Alibaba Cloud / Qwen Team | ▫ n.d. | 92,0 % | — | ✅ Mesuré |
| 65 | inclusionAI: Ring-2.6-1T | InclusionAI | ▫ n.d. | 91,7 % | 8 mai 2026 | ✅ Mesuré |
| 66 | Claude Haiku 4.5 | Anthropic | 🔒 Propriétaire | 91,0 % | 15 octobre 2025 | ✅ Mesuré |
| 67 | Deep Cogito: Cogito v2.1 671B | Deep Cogito | ▫ n.d. | 91,0 % | 13 novembre 2025 | ✅ Mesuré |
| 68 | GPT-4.1 | OpenAI | 🔒 Propriétaire | 91,0 % | 14 avril 2025 | ✅ Mesuré |
| 69 | MiniMax M1 | MiniMax | 🟢 Ouvert | 91,0 % | 17 juin 2025 | ✅ Mesuré |
| 70 | Relace: Relace Search | relace | ▫ n.d. | 91,0 % | 8 décembre 2025 | ✅ Mesuré |
| 71 | laguna-m.1 | poolside | ▫ n.d. | 91,0 % | — | ✅ Mesuré |
| 72 | mistral-large-2512 | mistralai | ▫ n.d. | 91,0 % | — | ✅ Mesuré |
| 73 | qwen3-14b-04-28 | Alibaba Cloud / Qwen Team | ▫ n.d. | 91,0 % | — | ✅ Mesuré |
| 74 | DeepSeek V3.1 Terminus | DeepSeek | 🟢 Ouvert | 90,0 % | 22 septembre 2025 | ✅ Mesuré |
| 75 | Tencent: Hunyuan A13B Instruct | Tencent | 🟢 Ouvert | 90,0 % | 8 juillet 2025 | ✅ Mesuré |
| 76 | qwen3-coder-next-2025-02-03 | Alibaba Cloud / Qwen Team | 🟢 Ouvert | 89,5 % | 4 février 2026 | ✅ Mesuré |
| 77 | Baidu: ERNIE 4.5 VL 424B A47B | Baidu | 🟢 Ouvert | 89,0 % | 30 juin 2025 | ✅ Mesuré |
| 78 | Qwen: Qwen3 Coder 30B A3B Instruct | Alibaba Cloud / Qwen Team | 🟢 Ouvert | 89,0 % | 31 juillet 2025 | ✅ Mesuré |
| 79 | Reka Flash 3 | Reka AI | 🟢 Ouvert | 89,0 % | 12 mars 2025 | ✅ Mesuré |
| 80 | deepseek-chat-v3 | DeepSeek | ▫ n.d. | 89,0 % | — | ✅ Mesuré |
| 81 | deepseek-chat-v3.1 | DeepSeek | ▫ n.d. | 89,0 % | — | ✅ Mesuré |
| 82 | gpt-4o-search-preview-2025-03-11 | OpenAI | 🔒 Propriétaire | 89,0 % | 12 mars 2025 | ✅ Mesuré |
| 83 | Command A+ | Cohere | 🟢 Ouvert | 88,0 % | 20 mai 2026 | ✅ Mesuré |
| 84 | devstral-2512 | mistralai | ▫ n.d. | 88,0 % | — | ✅ Mesuré |
| 85 | nova-2-lite-v1 | Amazon | 🔒 Propriétaire | 88,0 % | 2 décembre 2025 | ✅ Mesuré |
| 86 | qwen3-coder-480b-a35b-07-25 | Alibaba Cloud / Qwen Team | ▫ n.d. | 88,0 % | — | ✅ Mesuré |
| 87 | ByteDance Seed: Seed 1.6 Flash | ByteDance Seed | ▫ n.d. | 87,0 % | 23 décembre 2025 | ✅ Mesuré |
| 88 | GPT-4o mini | OpenAI | 🔒 Propriétaire | 87,0 % | 18 juillet 2024 | ✅ Mesuré |
| 89 | Mistral: Mistral Medium 3 | Mistral AI | ▫ n.d. | 87,0 % | 7 mai 2025 | ✅ Mesuré |
| 90 | OpenAI: GPT-5.4 Image 2 | OpenAI | 🔒 Propriétaire | 87,0 % | 21 avril 2026 | ✅ Mesuré |
| 91 | ministral-14b-2512 | mistralai | ▫ n.d. | 87,0 % | — | ✅ Mesuré |
| 92 | ministral-8b-2512 | mistralai | ▫ n.d. | 87,0 % | — | ✅ Mesuré |
| 93 | Muse Spark 1.1 | Meta | 🔒 Propriétaire | 86,9 % | 9 juillet 2026 | ✅ Mesuré |
| 94 | qwen3-235b-a22b-07-25 | Alibaba Cloud / Qwen Team | ▫ n.d. | 85,9 % | — | ✅ Mesuré |
| 95 | kimi-k2.5-0127 | Moonshot AI | ▫ n.d. | 85,0 % | — | ✅ Mesuré |
| 96 | hermes-3-llama-3.1-405b | nousresearch | ▫ n.d. | 85,0 % | — | ✅ Mesuré |
| 97 | Arcee AI: Virtuoso Large | Arcee AI | ▫ n.d. | 84,0 % | 5 mai 2025 | ✅ Mesuré |
| 98 | GPT-4.1 nano | OpenAI | 🔒 Propriétaire | 84,0 % | 14 avril 2025 | ✅ Mesuré |
| 99 | Kwaipilot: KAT-Coder-Air V2.5 | kwaipilot | ▫ n.d. | 84,0 % | 10 juillet 2026 | ✅ Mesuré |
| 100 | nova-pro-v1 | Amazon | 🔒 Propriétaire | 84,0 % | — | ✅ Mesuré |
| 101 | qwen-plus-2025-01-25 | Alibaba Cloud / Qwen Team | ▫ n.d. | 84,0 % | 8 septembre 2025 | ✅ Mesuré |
| 102 | Writer: Palmyra X5 | Writer | ▫ n.d. | 83,3 % | 21 janvier 2026 | ✅ Mesuré |
| 103 | Mistral: Mixtral 8x22B Instruct | Mistral AI | 🟢 Ouvert | 83,0 % | 17 avril 2024 | ✅ Mesuré |
| 104 | Nous: Hermes 4 405B | Nous Research | 🟢 Ouvert | 83,0 % | 26 août 2025 | ✅ Mesuré |
| 105 | Mistral Large | Mistral AI | ▫ n.d. | 82,0 % | 26 février 2024 | ✅ Mesuré |
| 106 | Mistral Large 2407 | Mistral AI | ▫ n.d. | 82,0 % | 19 novembre 2024 | ✅ Mesuré |
| 107 | Sonar | Perplexity | 🔒 Propriétaire | 82,0 % | 29 janvier 2025 | ✅ Mesuré |
| 108 | Upstage: Solar Pro 3 | Upstage | ▫ n.d. | 82,0 % | 27 janvier 2026 | ✅ Mesuré |
| 109 | nova-micro-v1 | Amazon | 🔒 Propriétaire | 82,0 % | — | ✅ Mesuré |
| 110 | uncensored | venice | ▫ n.d. | 81,8 % | — | ✅ Mesuré |
| 111 | Kwaipilot: KAT-Coder-Pro V2 | kwaipilot | ▫ n.d. | 81,0 % | 27 mars 2026 | ✅ Mesuré |
| 112 | MiniMax: MiniMax-01 | MiniMax | 🟢 Ouvert | 81,0 % | 15 janvier 2025 | ✅ Mesuré |
| 113 | hermes-3-llama-3.1-70b | nousresearch | ▫ n.d. | 81,0 % | — | ✅ Mesuré |
| 114 | Mistral: Voxtral Small 24B 2507 | Mistral AI | 🟢 Ouvert | 80,0 % | 30 octobre 2025 | ✅ Mesuré |
| 115 | Sonar Pro | Perplexity | 🔒 Propriétaire | 80,0 % | 7 mars 2025 | ✅ Mesuré |
| 116 | gpt-4o-mini-search-preview-2025-03-11 | OpenAI | 🔒 Propriétaire | 80,0 % | 12 mars 2025 | ✅ Mesuré |
| 117 | llama-4-maverick-17b-128e-instruct | Meta | ▫ n.d. | 79,5 % | — | ✅ Mesuré |
| 118 | llama-4-scout-17b-16e-instruct | Meta | ▫ n.d. | 79,5 % | — | ✅ Mesuré |
| 119 | Mistral NeMo | Mistral AI | ▫ n.d. | 79,0 % | 18 juillet 2024 | ✅ Mesuré |
| 120 | mistral-saba-2502 | mistralai | ▫ n.d. | 79,0 % | — | ✅ Mesuré |
| 121 | mistral-small-24b-instruct-2501 | mistralai | ▫ n.d. | 79,0 % | — | ✅ Mesuré |
| 122 | IBM: Granite 4.1 8B | IBM | 🟢 Ouvert | 78,0 % | 30 avril 2026 | ✅ Mesuré |
| 123 | ministral-3b-2512 | mistralai | ▫ n.d. | 78,0 % | — | ✅ Mesuré |
| 124 | Perceptron: Perceptron Mk1 | Perceptron | ▫ n.d. | 77,8 % | 12 mai 2026 | ✅ Mesuré |
| 125 | Cohere: Command R (08-2024) | Cohere | 🟢 Ouvert | 77,0 % | 30 août 2024 | ✅ Mesuré |
| 126 | command-r-plus-08-2024 | Cohere | ▫ n.d. | 77,0 % | — | ✅ Mesuré |
| 127 | skyfall-36b-v2 | thedrummer | ▫ n.d. | 77,0 % | — | ✅ Mesuré |
| 128 | inclusionAI: Ling-2.6-1T | InclusionAI | ▫ n.d. | 76,9 % | 23 avril 2026 | ✅ Mesuré |
| 129 | rocinante-12b | thedrummer | ▫ n.d. | 76,0 % | — | ✅ Mesuré |
| 130 | Cohere: Command R7B (12-2024) | Cohere | ▫ n.d. | 74,0 % | 14 décembre 2024 | ✅ Mesuré |
| 131 | Claude 3 Haiku | Anthropic | 🔒 Propriétaire | 72,0 % | 13 mars 2024 | ✅ Mesuré |
| 132 | Nex AGI: Nex-N2-Mini | Nex AGI | 🟢 Ouvert | 72,0 % | 24 juin 2026 | ✅ Mesuré |
| 133 | OpenAI: GPT-3.5 Turbo Instruct | OpenAI | 🔒 Propriétaire | 71,0 % | 28 septembre 2023 | ✅ Mesuré |
| 134 | ui-tars-1.5-7b | ByteDance | ▫ n.d. | 68,0 % | — | ✅ Mesuré |
| 135 | AI21: Jamba Large 1.7 | AI21 Labs | 🟢 Ouvert | 67,0 % | 8 août 2025 | ✅ Mesuré |
| 136 | unslopnemo-12b | thedrummer | ▫ n.d. | 67,0 % | — | ✅ Mesuré |
| 137 | nemotron-nano-12b-v2-vl | NVIDIA | 🟢 Ouvert | 65,6 % | 28 octobre 2025 | ✅ Mesuré |
| 138 | WizardLM-2 8x22B | Microsoft | 🟢 Ouvert | 62,0 % | 16 avril 2024 | ✅ Mesuré |
| 139 | Inflection: Inflection 3 Productivity | Inflection | 🔒 Propriétaire | 57,0 % | 11 octobre 2024 | ✅ Mesuré |
| 140 | Inflection: Inflection 3 Pi | Inflection | 🔒 Propriétaire | 56,0 % | 11 octobre 2024 | ✅ Mesuré |
| 141 | Nous: Hermes 4 70B | Nous Research | 🟢 Ouvert | 55,0 % | 26 août 2025 | ✅ Mesuré |
| 142 | Mistral: Codestral 2508 | Mistral AI | ▫ n.d. | 48,0 % | 1 août 2025 | ✅ Mesuré |
| 143 | aion-rp-llama-3.1-8b | Aion Labs | ▫ n.d. | 47,0 % | — | ✅ Mesuré |
| 144 | nova-premier-v1 | Amazon | 🔒 Propriétaire | 43,0 % | — | ✅ Mesuré |
| 145 | inclusionAI: Ling-2.6-flash | InclusionAI | ▫ n.d. | 37,0 % | 21 avril 2026 | ✅ Mesuré |
| 146 | qwen3-8b-04-28 | Alibaba Cloud / Qwen Team | ▫ n.d. | 34,0 % | — | ✅ Mesuré |
| 147 | MiniMax: MiniMax M2-her | MiniMax | ▫ n.d. | 25,0 % | 23 janvier 2026 | ✅ Mesuré |
| 148 | granite-4.0-h-micro | IBM | 🟢 Ouvert | 22,7 % | 20 octobre 2025 | ✅ Mesuré |
| 149 | l3.1-euryale-70b | sao10k | ▫ n.d. | 22,0 % | — | ✅ Mesuré |
| 150 | l3.3-euryale-70b-v2.3 | sao10k | ▫ n.d. | 21,0 % | — | ✅ Mesuré |
| 151 | GPT-4 | OpenAI | 🔒 Propriétaire | 7,0 % | 28 août 2023 | ✅ Mesuré |
| 152 | l3-lunaris-8b | sao10k | ▫ n.d. | 4,0 % | — | ✅ Mesuré |
| 153 | mythomax-l2-13b | gryphe | ▫ n.d. | 3,0 % | — | ✅ Mesuré |
| 154 | weaver | mancer | ▫ n.d. | 3,0 % | — | ✅ Mesuré |
| 155 | remm-slerp-l2-13b | undi95 | ▫ n.d. | 1,0 % | — | ✅ Mesuré |
| 156 | Meta: Llama 3.2 1B Instruct | Meta | 🟢 Ouvert | 0,0 % | 25 septembre 2024 | ✅ Mesuré |
| 157 | Meta: Llama Guard 4 12B | Meta | 🟢 Ouvert | 0,0 % | 30 avril 2025 | ✅ Mesuré |
| 158 | Morph: Morph V3 Fast | morph | ▫ n.d. | 0,0 % | 7 juillet 2025 | ✅ Mesuré |
| 159 | Morph: Morph V3 Large | morph | ▫ n.d. | 0,0 % | 7 juillet 2025 | ✅ Mesuré |
| 160 | Qwen 3.5 Plus | Qwen | ▫ n.d. | 0,0 % | 16 février 2026 | ✅ Mesuré |
| 161 | Qwen: Qwen3.5-Flash | Alibaba Cloud / Qwen Team | ▫ n.d. | 0,0 % | 25 février 2026 | ✅ Mesuré |
| 162 | Qwen: Qwen3.6 Flash | Alibaba Cloud / Qwen Team | ▫ n.d. | 0,0 % | 27 avril 2026 | ✅ Mesuré |
| 163 | qwen3.6-plus-04-02 | Alibaba Cloud / Qwen Team | ▫ n.d. | 0,0 % | 2 avril 2026 | ✅ Mesuré |
Classement établi sur 163 modèles évalués, dont 69 de grands éditeurs. Score médian de l'ensemble : 89,0 %.
Notre analyse
Un score élevé indique une forte maîtrise des connaissances de programmation couvertes et une bonne capacité à sélectionner la réponse attendue parmi six options. La validation par Exact Match sur la lettre fournie rend la notation stricte et reproductible, sans crédit partiel. Les scores sont au moins partiellement mesurés par un tiers, ce qui leur confère une assise plus robuste que des résultats uniquement auto-déclarés.
Le score médian de 89 % parmi 254 modèles et le résultat parfait de Qwen3 VL 32B Instruct suggèrent toutefois une saturation importante du benchmark. Les écarts en tête du classement peuvent donc devenir faibles et moins discriminants. Son accès public crée aussi un risque de contamination, les questions pouvant avoir été intégrées à des corpus d'entraînement ou d'évaluation. Enfin, sa portée reste celle d'un QCM de connaissances : il évalue la reconnaissance d'une option correcte dans plusieurs langages et domaines, plutôt que la production directe de code. Le classement révèle ainsi un niveau globalement élevé, avec au moins un modèle atteignant le plafond du test.
Sources des scores : benchable.