Instruction Following (Baseline)
Benchable : Instruction Following (Baseline) est un benchmark public créé par Benchable pour évaluer la capacité des modèles d’IA à suivre précisément des consignes en anglais. Les tâches progressent de directives simples vers des instructions conditionnelles à plusieurs niveaux, pouvant…
Benchable : Instruction Following (Baseline) est un benchmark public créé par Benchable pour évaluer la capacité des modèles d’IA à suivre précisément des consignes en anglais. Les tâches progressent de directives simples vers des instructions conditionnelles à plusieurs niveaux, pouvant cumuler de nombreuses exigences.
Le test examine notamment le respect du formatage, de l’ordre du contenu, des calculs et de la logique conditionnelle. Sa notation par correspondance exacte de l’ensemble du texte en fait un indicateur strict de conformité, utile pour distinguer les modèles capables d’exécuter fidèlement des instructions complexes.
Carte d'identité
| Caractéristique | Valeur |
|---|---|
| Éditeur du benchmark | Benchable |
| Capacités mesurées | Suivi precis d'instructions (formatage, ordre du contenu, calculs, logique conditionnelle) sur une gradation de complexite |
| Modalité | Texte |
| Type de questions | Taches de suivi d'instructions a complexite croissante |
| Métrique d'évaluation | Conformite exacte aux instructions (Exact Match, tout le texte) |
| Accès | Public |
| Langues | anglais |
| Taille du jeu | 100 etapes |
| Ressources | Site / dépôt officiel |
Classement des modèles (164)
| # | Modèle | Éditeur | Licence | Score | Sortie | Fiabilité |
|---|---|---|---|---|---|---|
| 1 | StepFun: Step 3.7 Flash | StepFun | 🟢 Ouvert | 100,0 % | 28 mai 2026 | ✅ Mesuré |
| 2 | inclusionAI: Ling-2.6-1T | InclusionAI | ▫ n.d. | 100,0 % | 23 avril 2026 | ✅ Mesuré |
| 3 | inclusionAI: Ring-2.6-1T | InclusionAI | ▫ n.d. | 100,0 % | 8 mai 2026 | ✅ Mesuré |
| 4 | Google: Gemini 3.1 Pro Preview Custom Tools | ▫ n.d. | 94,9 % | 25 février 2026 | ✅ Mesuré | |
| 5 | Gemini 3.1 Pro Preview | ▫ n.d. | 93,9 % | 19 février 2026 | ✅ Mesuré | |
| 6 | OpenAI: GPT-5.6 Sol Pro | OpenAI | 🔒 Propriétaire | 92,0 % | 9 juillet 2026 | ✅ Mesuré |
| 7 | Perceptron: Perceptron Mk1 | Perceptron | ▫ n.d. | 91,4 % | 12 mai 2026 | ✅ Mesuré |
| 8 | GPT-5 | OpenAI | 🔒 Propriétaire | 91,0 % | 7 août 2025 | ✅ Mesuré |
| 9 | GPT-5.6 Sol | OpenAI | 🔒 Propriétaire | 91,0 % | 9 juillet 2026 | ✅ Mesuré |
| 10 | OpenAI: GPT-5.2 Chat | OpenAI | 🔒 Propriétaire | 91,0 % | 10 décembre 2025 | ✅ Mesuré |
| 11 | OpenAI: GPT Chat Latest | OpenAI | 🔒 Propriétaire | 88,0 % | 5 mai 2026 | ✅ Mesuré |
| 12 | Sakana: Fugu Ultra | Sakana AI | ▫ n.d. | 88,0 % | 24 juin 2026 | ✅ Mesuré |
| 13 | OpenAI: GPT-5.6 Luna Pro | OpenAI | 🔒 Propriétaire | 86,9 % | 9 juillet 2026 | ✅ Mesuré |
| 14 | Muse Spark 1.1 | Meta | 🔒 Propriétaire | 86,2 % | 9 juillet 2026 | ✅ Mesuré |
| 15 | GPT-5.2 | OpenAI | 🔒 Propriétaire | 86,0 % | 11 décembre 2025 | ✅ Mesuré |
| 16 | gemini-3-pro-image | ▫ n.d. | 86,0 % | — | ✅ Mesuré | |
| 17 | OpenAI: GPT-5.1-Codex-Max | OpenAI | 🔒 Propriétaire | 85,0 % | 4 décembre 2025 | ✅ Mesuré |
| 18 | Qwen: Qwen3.6 Max Preview | Alibaba Cloud / Qwen Team | ▫ n.d. | 84,8 % | 27 avril 2026 | ✅ Mesuré |
| 19 | DeepSeek V4 Flash | DeepSeek | ▫ n.d. | 84,0 % | 24 avril 2026 | ✅ Mesuré |
| 20 | Z.ai: GLM 5 Turbo | Zhipu AI | ▫ n.d. | 84,0 % | 15 mars 2026 | ✅ Mesuré |
| 21 | gemini-2.5-pro-preview-03-25 | ▫ n.d. | 84,0 % | — | ✅ Mesuré | |
| 22 | kimi-k2.5-0127 | Moonshot AI | ▫ n.d. | 84,0 % | — | ✅ Mesuré |
| 23 | GPT-5.1 | OpenAI | 🔒 Propriétaire | 83,0 % | 13 novembre 2025 | ✅ Mesuré |
| 24 | OpenAI: GPT-5.6 Terra Pro | OpenAI | 🔒 Propriétaire | 83,0 % | 9 juillet 2026 | ✅ Mesuré |
| 25 | OpenAI: GPT-5.1 Chat | OpenAI | 🔒 Propriétaire | 83,0 % | 13 novembre 2025 | ✅ Mesuré |
| 26 | OpenAI: GPT-5.4 Image 2 | OpenAI | 🔒 Propriétaire | 82,0 % | 21 avril 2026 | ✅ Mesuré |
| 27 | Seed 1.6 | ByteDance Seed | ▫ n.d. | 82,0 % | 23 décembre 2025 | ✅ Mesuré |
| 28 | qwen3.6-plus-04-02 | Alibaba Cloud / Qwen Team | ▫ n.d. | 82,0 % | 2 avril 2026 | ✅ Mesuré |
| 29 | GPT-5.6 Luna | OpenAI | 🔒 Propriétaire | 81,0 % | 9 juillet 2026 | ✅ Mesuré |
| 30 | GPT-5.6 Terra | OpenAI | 🔒 Propriétaire | 81,0 % | 9 juillet 2026 | ✅ Mesuré |
| 31 | cydonia-24b-v4.1 | thedrummer | ▫ n.d. | 80,4 % | — | ✅ Mesuré |
| 32 | DeepSeek V4 Pro | DeepSeek | ▫ n.d. | 80,0 % | 24 avril 2026 | ✅ Mesuré |
| 33 | Tencent: Hy3 preview | Tencent | 🟢 Ouvert | 79,0 % | 22 avril 2026 | ✅ Mesuré |
| 34 | gemini-3.1-flash-image | ▫ n.d. | 78,0 % | — | ✅ Mesuré | |
| 35 | Claude Opus 4.5 | Anthropic | 🔒 Propriétaire | 77,0 % | 24 novembre 2025 | ✅ Mesuré |
| 36 | Kimi K2 | Moonshot AI | ▫ n.d. | 77,0 % | 6 novembre 2025 | ✅ Mesuré |
| 37 | gemini-3.1-flash-image-preview | ▫ n.d. | 77,0 % | — | ✅ Mesuré | |
| 38 | o1 | OpenAI | 🔒 Propriétaire | 77,0 % | 17 décembre 2024 | ✅ Mesuré |
| 39 | GPT-4.1 mini | OpenAI | 🔒 Propriétaire | 76,4 % | 14 avril 2025 | ✅ Mesuré |
| 40 | GPT-4.1 | OpenAI | 🔒 Propriétaire | 76,0 % | 14 avril 2025 | ✅ Mesuré |
| 41 | Qwen 3.5 Plus | Qwen | ▫ n.d. | 76,0 % | 16 février 2026 | ✅ Mesuré |
| 42 | AionLabs: Aion-3.0-Mini | Aion Labs | ▫ n.d. | 75,8 % | 7 juillet 2026 | ✅ Mesuré |
| 43 | GPT-5 mini | OpenAI | 🔒 Propriétaire | 75,0 % | 7 août 2025 | ✅ Mesuré |
| 44 | Nex AGI: Nex-N2-Pro | Nex AGI | 🟢 Ouvert | 75,0 % | 8 juin 2026 | ✅ Mesuré |
| 45 | Thinking Machines: Inkling | Thinking Machines | 🟢 Ouvert | 75,0 % | 17 juillet 2026 | ✅ Mesuré |
| 46 | gpt-5-chat-2025-08-07 | OpenAI | 🔒 Propriétaire | 75,0 % | 7 août 2025 | ✅ Mesuré |
| 47 | Hy3 | Tencent | 🟢 Ouvert | 74,0 % | 6 juillet 2026 | ✅ Mesuré |
| 48 | Claude Opus 4 | Anthropic | 🔒 Propriétaire | 71,0 % | 22 mai 2025 | ✅ Mesuré |
| 49 | Qwen: Qwen3 Coder Plus | Alibaba Cloud / Qwen Team | ▫ n.d. | 71,0 % | 23 septembre 2025 | ✅ Mesuré |
| 50 | deepseek-chat-v3 | DeepSeek | ▫ n.d. | 71,0 % | — | ✅ Mesuré |
| 51 | deepseek-chat-v3-0324 | DeepSeek | ▫ n.d. | 71,0 % | — | ✅ Mesuré |
| 52 | mistral-large-2512 | mistralai | ▫ n.d. | 71,0 % | — | ✅ Mesuré |
| 53 | ByteDance Seed: Seed 1.6 Flash | ByteDance Seed | ▫ n.d. | 70,0 % | 23 décembre 2025 | ✅ Mesuré |
| 54 | Claude Haiku 4.5 | Anthropic | 🔒 Propriétaire | 70,0 % | 15 octobre 2025 | ✅ Mesuré |
| 55 | Deep Cogito: Cogito v2.1 671B | Deep Cogito | ▫ n.d. | 70,0 % | 13 novembre 2025 | ✅ Mesuré |
| 56 | Qwen: Qwen3.6 Flash | Alibaba Cloud / Qwen Team | ▫ n.d. | 70,0 % | 27 avril 2026 | ✅ Mesuré |
| 57 | deepseek-chat-v3.1 | DeepSeek | ▫ n.d. | 70,0 % | — | ✅ Mesuré |
| 58 | Claude Opus 4.1 | Anthropic | 🔒 Propriétaire | 69,0 % | 5 août 2025 | ✅ Mesuré |
| 59 | GPT-4o | OpenAI | 🔒 Propriétaire | 69,0 % | 27 mars 2025 | ✅ Mesuré |
| 60 | GPT-5 nano | OpenAI | 🔒 Propriétaire | 69,0 % | 7 août 2025 | ✅ Mesuré |
| 61 | OpenAI: GPT-4 Turbo Preview | OpenAI | 🔒 Propriétaire | 69,0 % | 25 janvier 2024 | ✅ Mesuré |
| 62 | OpenAI: gpt-oss-safeguard-20b | OpenAI | 🟢 Ouvert | 69,0 % | 29 octobre 2025 | ✅ Mesuré |
| 63 | laguna-xs-2.1 | poolside | ▫ n.d. | 68,0 % | — | ✅ Mesuré |
| 64 | AionLabs: Aion-3.0 | Aion Labs | ▫ n.d. | 67,7 % | 7 juillet 2026 | ✅ Mesuré |
| 65 | Claude Sonnet 4.5 | Anthropic | 🔒 Propriétaire | 67,7 % | 29 septembre 2025 | ✅ Mesuré |
| 66 | GPT-4 | OpenAI | 🔒 Propriétaire | 67,0 % | 28 août 2023 | ✅ Mesuré |
| 67 | Claude Sonnet 4 | Anthropic | 🔒 Propriétaire | 66,0 % | 22 mai 2025 | ✅ Mesuré |
| 68 | Mistral: Mistral Medium 3.1 | Mistral AI | ▫ n.d. | 66,0 % | 13 août 2025 | ✅ Mesuré |
| 69 | qwen3-coder-480b-a35b-07-25 | Alibaba Cloud / Qwen Team | ▫ n.d. | 65,4 % | — | ✅ Mesuré |
| 70 | Kwaipilot: KAT-Coder-Pro V2 | kwaipilot | ▫ n.d. | 65,0 % | 27 mars 2026 | ✅ Mesuré |
| 71 | MiniMax: MiniMax-01 | MiniMax | 🟢 Ouvert | 65,0 % | 15 janvier 2025 | ✅ Mesuré |
| 72 | Mistral: Mistral Medium 3 | Mistral AI | ▫ n.d. | 64,0 % | 7 mai 2025 | ✅ Mesuré |
| 73 | Nex AGI: Nex-N2-Mini | Nex AGI | 🟢 Ouvert | 64,0 % | 24 juin 2026 | ✅ Mesuré |
| 74 | Qwen: Qwen3.5-Flash | Alibaba Cloud / Qwen Team | ▫ n.d. | 63,6 % | 25 février 2026 | ✅ Mesuré |
| 75 | Arcee AI: Virtuoso Large | Arcee AI | ▫ n.d. | 63,0 % | 5 mai 2025 | ✅ Mesuré |
| 76 | Mistral Large | Mistral AI | ▫ n.d. | 63,0 % | 26 février 2024 | ✅ Mesuré |
| 77 | hermes-3-llama-3.1-405b | nousresearch | ▫ n.d. | 63,0 % | — | ✅ Mesuré |
| 78 | mistral-saba-2502 | mistralai | ▫ n.d. | 63,0 % | — | ✅ Mesuré |
| 79 | qwen3-next-80b-a3b-instruct-2509 | Alibaba Cloud / Qwen Team | ▫ n.d. | 63,0 % | — | ✅ Mesuré |
| 80 | MiniMax M1 | MiniMax | 🟢 Ouvert | 62,6 % | 17 juin 2025 | ✅ Mesuré |
| 81 | Baidu: ERNIE 4.5 VL 424B A47B | Baidu | 🟢 Ouvert | 62,0 % | 30 juin 2025 | ✅ Mesuré |
| 82 | Mistral Large 2407 | Mistral AI | ▫ n.d. | 62,0 % | 19 novembre 2024 | ✅ Mesuré |
| 83 | devstral-2512 | mistralai | ▫ n.d. | 62,0 % | — | ✅ Mesuré |
| 84 | Claude 3 Haiku | Anthropic | 🔒 Propriétaire | 61,0 % | 13 mars 2024 | ✅ Mesuré |
| 85 | Nous: Hermes 4 405B | Nous Research | 🟢 Ouvert | 61,0 % | 26 août 2025 | ✅ Mesuré |
| 86 | qwen-plus-2025-01-25 | Alibaba Cloud / Qwen Team | ▫ n.d. | 61,0 % | 8 septembre 2025 | ✅ Mesuré |
| 87 | qwen3-30b-a3b-04-28 | Alibaba Cloud / Qwen Team | ▫ n.d. | 60,8 % | — | ✅ Mesuré |
| 88 | GPT-4o mini | OpenAI | 🔒 Propriétaire | 60,5 % | 18 juillet 2024 | ✅ Mesuré |
| 89 | Perplexity: Sonar Pro Search | Perplexity | 🔒 Propriétaire | 60,0 % | 30 octobre 2025 | ✅ Mesuré |
| 90 | qwen3-8b-04-28 | Alibaba Cloud / Qwen Team | ▫ n.d. | 60,0 % | — | ✅ Mesuré |
| 91 | xAI: Grok Build 0.1 | xAI | ▫ n.d. | 60,0 % | 20 mai 2026 | ✅ Mesuré |
| 92 | Qwen: Qwen3 30B A3B Instruct 2507 | Alibaba Cloud / Qwen Team | 🟢 Ouvert | 59,8 % | 29 juillet 2025 | ✅ Mesuré |
| 93 | qwen3-coder-next-2025-02-03 | Alibaba Cloud / Qwen Team | 🟢 Ouvert | 59,6 % | 4 février 2026 | ✅ Mesuré |
| 94 | uncensored | venice | ▫ n.d. | 59,1 % | — | ✅ Mesuré |
| 95 | l3.3-euryale-70b-v2.3 | sao10k | ▫ n.d. | 59,0 % | — | ✅ Mesuré |
| 96 | Mistral: Codestral 2508 | Mistral AI | ▫ n.d. | 57,0 % | 1 août 2025 | ✅ Mesuré |
| 97 | Qwen: Qwen3 Coder Flash | Alibaba Cloud / Qwen Team | ▫ n.d. | 56,0 % | 17 septembre 2025 | ✅ Mesuré |
| 98 | qwen3-32b-04-28 | Alibaba Cloud / Qwen Team | ▫ n.d. | 55,7 % | — | ✅ Mesuré |
| 99 | hermes-3-llama-3.1-70b | nousresearch | ▫ n.d. | 55,6 % | — | ✅ Mesuré |
| 100 | GPT-4.1 nano | OpenAI | 🔒 Propriétaire | 55,5 % | 14 avril 2025 | ✅ Mesuré |
| 101 | Nous: Hermes 4 70B | Nous Research | 🟢 Ouvert | 55,1 % | 26 août 2025 | ✅ Mesuré |
| 102 | AionLabs: Aion-2.0 | Aion Labs | ▫ n.d. | 55,0 % | 23 février 2026 | ✅ Mesuré |
| 103 | nova-pro-v1 | Amazon | 🔒 Propriétaire | 55,0 % | — | ✅ Mesuré |
| 104 | Relace: Relace Search | relace | ▫ n.d. | 53,0 % | 8 décembre 2025 | ✅ Mesuré |
| 105 | ministral-14b-2512 | mistralai | ▫ n.d. | 52,0 % | — | ✅ Mesuré |
| 106 | mistral-small-24b-instruct-2501 | mistralai | ▫ n.d. | 52,0 % | — | ✅ Mesuré |
| 107 | mistral-small-2603 | mistralai | ▫ n.d. | 52,0 % | — | ✅ Mesuré |
| 108 | Qwen: Qwen3 Coder 30B A3B Instruct | Alibaba Cloud / Qwen Team | 🟢 Ouvert | 51,5 % | 31 juillet 2025 | ✅ Mesuré |
| 109 | Mistral: Voxtral Small 24B 2507 | Mistral AI | 🟢 Ouvert | 51,0 % | 30 octobre 2025 | ✅ Mesuré |
| 110 | qwen3-14b-04-28 | Alibaba Cloud / Qwen Team | ▫ n.d. | 50,5 % | — | ✅ Mesuré |
| 111 | Cohere: Command R (08-2024) | Cohere | 🟢 Ouvert | 50,0 % | 30 août 2024 | ✅ Mesuré |
| 112 | Mistral: Mixtral 8x22B Instruct | Mistral AI | 🟢 Ouvert | 50,0 % | 17 avril 2024 | ✅ Mesuré |
| 113 | Upstage: Solar Pro 3 | Upstage | ▫ n.d. | 50,0 % | 27 janvier 2026 | ✅ Mesuré |
| 114 | ByteDance Seed: Seed-2.0-Mini | ByteDance Seed | ▫ n.d. | 49,0 % | 26 février 2026 | ✅ Mesuré |
| 115 | gpt-3.5-turbo-0613 | OpenAI | 🔒 Propriétaire | 49,0 % | — | ✅ Mesuré |
| 116 | command-r-plus-08-2024 | Cohere | ▫ n.d. | 48,0 % | — | ✅ Mesuré |
| 117 | nova-2-lite-v1 | Amazon | 🔒 Propriétaire | 48,0 % | 2 décembre 2025 | ✅ Mesuré |
| 118 | l3.1-euryale-70b | sao10k | ▫ n.d. | 47,5 % | — | ✅ Mesuré |
| 119 | Writer: Palmyra X5 | Writer | ▫ n.d. | 47,1 % | 21 janvier 2026 | ✅ Mesuré |
| 120 | Sonar Pro | Perplexity | 🔒 Propriétaire | 47,0 % | 7 mars 2025 | ✅ Mesuré |
| 121 | granite-4.0-h-micro | IBM | 🟢 Ouvert | 46,2 % | 20 octobre 2025 | ✅ Mesuré |
| 122 | AI21: Jamba Large 1.7 | AI21 Labs | 🟢 Ouvert | 44,0 % | 8 août 2025 | ✅ Mesuré |
| 123 | ministral-8b-2512 | mistralai | ▫ n.d. | 44,0 % | — | ✅ Mesuré |
| 124 | skyfall-36b-v2 | thedrummer | ▫ n.d. | 42,0 % | — | ✅ Mesuré |
| 125 | l3-lunaris-8b | sao10k | ▫ n.d. | 41,0 % | — | ✅ Mesuré |
| 126 | qwen3-235b-a22b-04-28 | Alibaba Cloud / Qwen Team | ▫ n.d. | 40,4 % | — | ✅ Mesuré |
| 127 | qwen3-235b-a22b-07-25 | Alibaba Cloud / Qwen Team | ▫ n.d. | 39,7 % | — | ✅ Mesuré |
| 128 | unslopnemo-12b | thedrummer | ▫ n.d. | 39,0 % | — | ✅ Mesuré |
| 129 | llama-4-scout-17b-16e-instruct | Meta | ▫ n.d. | 38,7 % | — | ✅ Mesuré |
| 130 | OpenAI: GPT-3.5 Turbo Instruct | OpenAI | 🔒 Propriétaire | 38,0 % | 28 septembre 2023 | ✅ Mesuré |
| 131 | Mistral NeMo | Mistral AI | ▫ n.d. | 37,0 % | 18 juillet 2024 | ✅ Mesuré |
| 132 | rocinante-12b | thedrummer | ▫ n.d. | 36,0 % | — | ✅ Mesuré |
| 133 | ministral-3b-2512 | mistralai | ▫ n.d. | 33,0 % | — | ✅ Mesuré |
| 134 | ui-tars-1.5-7b | ByteDance | ▫ n.d. | 32,3 % | — | ✅ Mesuré |
| 135 | Sonar | Perplexity | 🔒 Propriétaire | 32,0 % | 29 janvier 2025 | ✅ Mesuré |
| 136 | llama-4-maverick-17b-128e-instruct | Meta | ▫ n.d. | 30,3 % | — | ✅ Mesuré |
| 137 | Cohere: Command R7B (12-2024) | Cohere | ▫ n.d. | 27,0 % | 14 décembre 2024 | ✅ Mesuré |
| 138 | MiniMax: MiniMax M2-her | MiniMax | ▫ n.d. | 27,0 % | 23 janvier 2026 | ✅ Mesuré |
| 139 | Meta: Llama 3.2 1B Instruct | Meta | 🟢 Ouvert | 18,9 % | 25 septembre 2024 | ✅ Mesuré |
| 140 | IBM: Granite 4.1 8B | IBM | 🟢 Ouvert | 16,0 % | 30 avril 2026 | ✅ Mesuré |
| 141 | qwen3-next-80b-a3b-thinking-2509 | Alibaba Cloud / Qwen Team | ▫ n.d. | 14,7 % | — | ✅ Mesuré |
| 142 | nova-micro-v1 | Amazon | 🔒 Propriétaire | 14,0 % | — | ✅ Mesuré |
| 143 | inclusionAI: Ling-2.6-flash | InclusionAI | ▫ n.d. | 12,0 % | 21 avril 2026 | ✅ Mesuré |
| 144 | weaver | mancer | ▫ n.d. | 4,0 % | — | ✅ Mesuré |
| 145 | nova-premier-v1 | Amazon | 🔒 Propriétaire | 3,0 % | — | ✅ Mesuré |
| 146 | Arcee AI: Trinity Large Thinking | Arcee AI | 🟢 Ouvert | 0,0 % | 1 avril 2026 | ✅ Mesuré |
| 147 | Command A+ | Cohere | 🟢 Ouvert | 0,0 % | 20 mai 2026 | ✅ Mesuré |
| 148 | Inflection: Inflection 3 Pi | Inflection | 🔒 Propriétaire | 0,0 % | 11 octobre 2024 | ✅ Mesuré |
| 149 | Inflection: Inflection 3 Productivity | Inflection | 🔒 Propriétaire | 0,0 % | 11 octobre 2024 | ✅ Mesuré |
| 150 | Kwaipilot: KAT-Coder-Air V2.5 | kwaipilot | ▫ n.d. | 0,0 % | 10 juillet 2026 | ✅ Mesuré |
| 151 | Kwaipilot: KAT-Coder-Pro V2.5 | kwaipilot | ▫ n.d. | 0,0 % | 10 juillet 2026 | ✅ Mesuré |
| 152 | Meta: Llama Guard 4 12B | Meta | 🟢 Ouvert | 0,0 % | 30 avril 2025 | ✅ Mesuré |
| 153 | Qwen: Qwen3 30B A3B Thinking 2507 | Alibaba Cloud / Qwen Team | 🟢 Ouvert | 0,0 % | 28 août 2025 | ✅ Mesuré |
| 154 | Reka Flash 3 | Reka AI | 🟢 Ouvert | 0,0 % | 12 mars 2025 | ✅ Mesuré |
| 155 | Tencent: Hunyuan A13B Instruct | Tencent | 🟢 Ouvert | 0,0 % | 8 juillet 2025 | ✅ Mesuré |
| 156 | WizardLM-2 8x22B | Microsoft | 🟢 Ouvert | 0,0 % | 16 avril 2024 | ✅ Mesuré |
| 157 | aion-rp-llama-3.1-8b | Aion Labs | ▫ n.d. | 0,0 % | — | ✅ Mesuré |
| 158 | gpt-4o-mini-search-preview-2025-03-11 | OpenAI | 🔒 Propriétaire | 0,0 % | 12 mars 2025 | ✅ Mesuré |
| 159 | gpt-4o-search-preview-2025-03-11 | OpenAI | 🔒 Propriétaire | 0,0 % | 12 mars 2025 | ✅ Mesuré |
| 160 | laguna-m.1 | poolside | ▫ n.d. | 0,0 % | — | ✅ Mesuré |
| 161 | mythomax-l2-13b | gryphe | ▫ n.d. | 0,0 % | — | ✅ Mesuré |
| 162 | nemotron-nano-12b-v2-vl | NVIDIA | 🟢 Ouvert | 0,0 % | 28 octobre 2025 | ✅ Mesuré |
| 163 | olmo-3-32b-think | allenai | ▫ n.d. | 0,0 % | — | ✅ Mesuré |
| 164 | remm-slerp-l2-13b | undi95 | ▫ n.d. | 0,0 % | — | ✅ Mesuré |
Classement établi sur 164 modèles évalués, dont 69 de grands éditeurs. Score médian de l'ensemble : 62,0 %.
Notre analyse
Un score élevé indique qu’un modèle respecte littéralement les contraintes demandées, y compris lorsque leur nombre et leur imbrication augmentent. La correspondance exacte est particulièrement rigoureuse : une réponse globalement correcte peut échouer si son texte ne satisfait pas entièrement le format, l’ordre ou les conditions imposées. Les scores étant au moins partiellement mesurés par un tiers, leur fiabilité ne repose pas uniquement sur des résultats auto-déclarés. Le classement met en évidence une forte dispersion entre la médiane de l’ensemble et DeepSeek R1 Distill Llama 70B, qui atteint la conformité maximale. Ce résultat signale aussi un risque de saturation au sommet, puisque le benchmark ne peut plus départager les modèles obtenant 100 %. Sa portée reste centrée sur le suivi d’instructions en anglais et sur l’Exact Match, plutôt que sur une évaluation générale des capacités. Enfin, le caractère public du jeu implique un risque de contamination par exposition préalable aux tâches, ce qui invite à interpréter les meilleurs scores avec prudence.
Sources des scores : benchable.