Hallucinations (Baseline)
Benchable : Hallucinations (Baseline) est un benchmark créé par Benchable pour évaluer l’humilité épistémique des modèles d’IA. Il vérifie leur capacité à reconnaître l’incertitude face à des concepts, événements ou entités entièrement fictifs, plutôt qu’à produire une réponse inventée.
Benchable : Hallucinations (Baseline) est un benchmark créé par Benchable pour évaluer l’humilité épistémique des modèles d’IA. Il vérifie leur capacité à reconnaître l’incertitude face à des concepts, événements ou entités entièrement fictifs, plutôt qu’à produire une réponse inventée.
L’épreuve repose sur des QCM en anglais dans lesquels « Je ne sais pas » constitue systématiquement la bonne réponse. Elle fournit ainsi un indicateur ciblé de résistance aux hallucinations, utile pour comparer la prudence des modèles lorsqu’une question repose sur de fausses prémisses.
Carte d'identité
| Caractéristique | Valeur |
|---|---|
| Éditeur du benchmark | Benchable |
| Capacités mesurées | Humilite epistemique / resistance aux hallucinations : reconnaitre l'incertitude face a des concepts, evenements ou entites entierement fictifs |
| Modalité | Texte |
| Type de questions | QCM (A/B/C/D) ou 'Je ne sais pas' est toujours la bonne reponse |
| Métrique d'évaluation | Exactitude (100% = aucune hallucination, le modele repond correctement 'Je ne sais pas') |
| Accès | Public |
| Langues | anglais |
| Taille du jeu | 50 questions |
| Ressources | Site / dépôt officiel |
Classement des modèles (147)
| # | Modèle | Éditeur | Licence | Score | Sortie | Fiabilité |
|---|---|---|---|---|---|---|
| 1 | AionLabs: Aion-2.0 | Aion Labs | ▫ n.d. | 100,0 % | 23 février 2026 | ✅ Mesuré |
| 2 | AionLabs: Aion-3.0 | Aion Labs | ▫ n.d. | 100,0 % | 7 juillet 2026 | ✅ Mesuré |
| 3 | Arcee AI: Virtuoso Large | Arcee AI | ▫ n.d. | 100,0 % | 5 mai 2025 | ✅ Mesuré |
| 4 | Claude Haiku 4.5 | Anthropic | 🔒 Propriétaire | 100,0 % | 15 octobre 2025 | ✅ Mesuré |
| 5 | Claude Opus 4 | Anthropic | 🔒 Propriétaire | 100,0 % | 22 mai 2025 | ✅ Mesuré |
| 6 | Claude Opus 4.1 | Anthropic | 🔒 Propriétaire | 100,0 % | 5 août 2025 | ✅ Mesuré |
| 7 | Claude Sonnet 4 | Anthropic | 🔒 Propriétaire | 100,0 % | 22 mai 2025 | ✅ Mesuré |
| 8 | Claude Sonnet 4.5 | Anthropic | 🔒 Propriétaire | 100,0 % | 29 septembre 2025 | ✅ Mesuré |
| 9 | Deep Cogito: Cogito v2.1 671B | Deep Cogito | ▫ n.d. | 100,0 % | 13 novembre 2025 | ✅ Mesuré |
| 10 | DeepSeek V3.1 Terminus | DeepSeek | 🟢 Ouvert | 100,0 % | 22 septembre 2025 | ✅ Mesuré |
| 11 | GPT-4o | OpenAI | 🔒 Propriétaire | 100,0 % | 27 mars 2025 | ✅ Mesuré |
| 12 | GPT-5 mini | OpenAI | 🔒 Propriétaire | 100,0 % | 7 août 2025 | ✅ Mesuré |
| 13 | GPT-5 nano | OpenAI | 🔒 Propriétaire | 100,0 % | 7 août 2025 | ✅ Mesuré |
| 14 | GPT-5.1 | OpenAI | 🔒 Propriétaire | 100,0 % | 13 novembre 2025 | ✅ Mesuré |
| 15 | GPT-5.2 | OpenAI | 🔒 Propriétaire | 100,0 % | 11 décembre 2025 | ✅ Mesuré |
| 16 | Gemini 3.1 Pro Preview | ▫ n.d. | 100,0 % | 19 février 2026 | ✅ Mesuré | |
| 17 | Google: Gemini 3.1 Pro Preview Custom Tools | ▫ n.d. | 100,0 % | 25 février 2026 | ✅ Mesuré | |
| 18 | Kwaipilot: KAT-Coder-Pro V2 | kwaipilot | ▫ n.d. | 100,0 % | 27 mars 2026 | ✅ Mesuré |
| 19 | Magnum v4 72B | anthracite-org | 🟢 Ouvert | 100,0 % | 22 octobre 2024 | ✅ Mesuré |
| 20 | Mistral Large | Mistral AI | ▫ n.d. | 100,0 % | 26 février 2024 | ✅ Mesuré |
| 21 | Mistral Large 2407 | Mistral AI | ▫ n.d. | 100,0 % | 19 novembre 2024 | ✅ Mesuré |
| 22 | Mistral: Mistral Medium 3 | Mistral AI | ▫ n.d. | 100,0 % | 7 mai 2025 | ✅ Mesuré |
| 23 | Mistral: Mistral Medium 3.1 | Mistral AI | ▫ n.d. | 100,0 % | 13 août 2025 | ✅ Mesuré |
| 24 | Mistral: Mixtral 8x22B Instruct | Mistral AI | 🟢 Ouvert | 100,0 % | 17 avril 2024 | ✅ Mesuré |
| 25 | Nous: Hermes 4 405B | Nous Research | 🟢 Ouvert | 100,0 % | 26 août 2025 | ✅ Mesuré |
| 26 | OpenAI: GPT-5.1 Chat | OpenAI | 🔒 Propriétaire | 100,0 % | 13 novembre 2025 | ✅ Mesuré |
| 27 | OpenAI: GPT-5.2 Chat | OpenAI | 🔒 Propriétaire | 100,0 % | 10 décembre 2025 | ✅ Mesuré |
| 28 | OpenAI: GPT-5.4 Image 2 | OpenAI | 🔒 Propriétaire | 100,0 % | 21 avril 2026 | ✅ Mesuré |
| 29 | Qwen: Qwen3 Coder Plus | Alibaba Cloud / Qwen Team | ▫ n.d. | 100,0 % | 23 septembre 2025 | ✅ Mesuré |
| 30 | Qwen: Qwen3.6 Max Preview | Alibaba Cloud / Qwen Team | ▫ n.d. | 100,0 % | 27 avril 2026 | ✅ Mesuré |
| 31 | Sakana: Fugu Ultra | Sakana AI | ▫ n.d. | 100,0 % | 24 juin 2026 | ✅ Mesuré |
| 32 | Seed 1.6 | ByteDance Seed | ▫ n.d. | 100,0 % | 23 décembre 2025 | ✅ Mesuré |
| 33 | StepFun: Step 3.7 Flash | StepFun | 🟢 Ouvert | 100,0 % | 28 mai 2026 | ✅ Mesuré |
| 34 | Writer: Palmyra X5 | Writer | ▫ n.d. | 100,0 % | 21 janvier 2026 | ✅ Mesuré |
| 35 | Z.ai: GLM 5 Turbo | Zhipu AI | ▫ n.d. | 100,0 % | 15 mars 2026 | ✅ Mesuré |
| 36 | deepseek-chat-v3 | DeepSeek | ▫ n.d. | 100,0 % | — | ✅ Mesuré |
| 37 | devstral-2512 | mistralai | ▫ n.d. | 100,0 % | — | ✅ Mesuré |
| 38 | gemini-3.1-flash-image | ▫ n.d. | 100,0 % | — | ✅ Mesuré | |
| 39 | gemini-3.1-flash-image-preview | ▫ n.d. | 100,0 % | — | ✅ Mesuré | |
| 40 | inclusionAI: Ling-2.6-1T | InclusionAI | ▫ n.d. | 100,0 % | 23 avril 2026 | ✅ Mesuré |
| 41 | inclusionAI: Ring-2.6-1T | InclusionAI | ▫ n.d. | 100,0 % | 8 mai 2026 | ✅ Mesuré |
| 42 | kimi-k2.5-0127 | Moonshot AI | ▫ n.d. | 100,0 % | — | ✅ Mesuré |
| 43 | nova-premier-v1 | Amazon | 🔒 Propriétaire | 100,0 % | — | ✅ Mesuré |
| 44 | qwen-plus-2025-01-25 | Alibaba Cloud / Qwen Team | ▫ n.d. | 100,0 % | 8 septembre 2025 | ✅ Mesuré |
| 45 | qwen3-next-80b-a3b-instruct-2509 | Alibaba Cloud / Qwen Team | ▫ n.d. | 100,0 % | — | ✅ Mesuré |
| 46 | qwen3.6-plus-04-02 | Alibaba Cloud / Qwen Team | ▫ n.d. | 100,0 % | 2 avril 2026 | ✅ Mesuré |
| 47 | xAI: Grok Build 0.1 | xAI | ▫ n.d. | 100,0 % | 20 mai 2026 | ✅ Mesuré |
| 48 | AionLabs: Aion-3.0-Mini | Aion Labs | ▫ n.d. | 98,0 % | 7 juillet 2026 | ✅ Mesuré |
| 49 | Claude Opus 4.5 | Anthropic | 🔒 Propriétaire | 98,0 % | 24 novembre 2025 | ✅ Mesuré |
| 50 | Command A+ | Cohere | 🟢 Ouvert | 98,0 % | 20 mai 2026 | ✅ Mesuré |
| 51 | GPT-5 | OpenAI | 🔒 Propriétaire | 98,0 % | 7 août 2025 | ✅ Mesuré |
| 52 | GPT-5.6 Luna | OpenAI | 🔒 Propriétaire | 98,0 % | 9 juillet 2026 | ✅ Mesuré |
| 53 | GPT-5.6 Sol | OpenAI | 🔒 Propriétaire | 98,0 % | 9 juillet 2026 | ✅ Mesuré |
| 54 | Hy3 | Tencent | 🟢 Ouvert | 98,0 % | 6 juillet 2026 | ✅ Mesuré |
| 55 | IBM: Granite 4.1 8B | IBM | 🟢 Ouvert | 98,0 % | 30 avril 2026 | ✅ Mesuré |
| 56 | Nous: Hermes 4 70B | Nous Research | 🟢 Ouvert | 98,0 % | 26 août 2025 | ✅ Mesuré |
| 57 | OpenAI: GPT Chat Latest | OpenAI | 🔒 Propriétaire | 98,0 % | 5 mai 2026 | ✅ Mesuré |
| 58 | OpenAI: GPT-5.6 Luna Pro | OpenAI | 🔒 Propriétaire | 98,0 % | 9 juillet 2026 | ✅ Mesuré |
| 59 | Qwen 3.5 Plus | Qwen | ▫ n.d. | 98,0 % | 16 février 2026 | ✅ Mesuré |
| 60 | Qwen: Qwen3 30B A3B Instruct 2507 | Alibaba Cloud / Qwen Team | 🟢 Ouvert | 98,0 % | 29 juillet 2025 | ✅ Mesuré |
| 61 | deepseek-chat-v3-0324 | DeepSeek | ▫ n.d. | 98,0 % | — | ✅ Mesuré |
| 62 | deepseek-chat-v3.1 | DeepSeek | ▫ n.d. | 98,0 % | — | ✅ Mesuré |
| 63 | gemini-3-pro-image | ▫ n.d. | 98,0 % | — | ✅ Mesuré | |
| 64 | mistral-saba-2502 | mistralai | ▫ n.d. | 98,0 % | — | ✅ Mesuré |
| 65 | o1 | OpenAI | 🔒 Propriétaire | 98,0 % | 17 décembre 2024 | ✅ Mesuré |
| 66 | qwen3-next-80b-a3b-thinking-2509 | Alibaba Cloud / Qwen Team | ▫ n.d. | 98,0 % | — | ✅ Mesuré |
| 67 | Qwen: Qwen3.5-Flash | Alibaba Cloud / Qwen Team | ▫ n.d. | 97,1 % | 25 février 2026 | ✅ Mesuré |
| 68 | Baidu: ERNIE 4.5 VL 424B A47B | Baidu | 🟢 Ouvert | 96,0 % | 30 juin 2025 | ✅ Mesuré |
| 69 | GPT-4.1 nano | OpenAI | 🔒 Propriétaire | 96,0 % | 14 avril 2025 | ✅ Mesuré |
| 70 | GPT-5.6 Terra | OpenAI | 🔒 Propriétaire | 96,0 % | 9 juillet 2026 | ✅ Mesuré |
| 71 | OpenAI: GPT-5.6 Sol Pro | OpenAI | 🔒 Propriétaire | 96,0 % | 9 juillet 2026 | ✅ Mesuré |
| 72 | Qwen: Qwen3.6 Flash | Alibaba Cloud / Qwen Team | ▫ n.d. | 96,0 % | 27 avril 2026 | ✅ Mesuré |
| 73 | gemini-2.5-pro-preview-03-25 | ▫ n.d. | 96,0 % | — | ✅ Mesuré | |
| 74 | hermes-3-llama-3.1-405b | nousresearch | ▫ n.d. | 96,0 % | — | ✅ Mesuré |
| 75 | mistral-small-24b-instruct-2501 | mistralai | ▫ n.d. | 96,0 % | — | ✅ Mesuré |
| 76 | nova-pro-v1 | Amazon | 🔒 Propriétaire | 96,0 % | — | ✅ Mesuré |
| 77 | qwen3-30b-a3b-04-28 | Alibaba Cloud / Qwen Team | ▫ n.d. | 96,0 % | — | ✅ Mesuré |
| 78 | Thinking Machines: Inkling | Thinking Machines | 🟢 Ouvert | 95,8 % | 17 juillet 2026 | ✅ Mesuré |
| 79 | Kimi K2 | Moonshot AI | ▫ n.d. | 94,0 % | 6 novembre 2025 | ✅ Mesuré |
| 80 | OpenAI: GPT-4 Turbo Preview | OpenAI | 🔒 Propriétaire | 94,0 % | 25 janvier 2024 | ✅ Mesuré |
| 81 | Qwen: Qwen3 Coder 30B A3B Instruct | Alibaba Cloud / Qwen Team | 🟢 Ouvert | 94,0 % | 31 juillet 2025 | ✅ Mesuré |
| 82 | nova-micro-v1 | Amazon | 🔒 Propriétaire | 94,0 % | — | ✅ Mesuré |
| 83 | qwen3-coder-next-2025-02-03 | Alibaba Cloud / Qwen Team | 🟢 Ouvert | 94,0 % | 4 février 2026 | ✅ Mesuré |
| 84 | DeepSeek V4 Pro | DeepSeek | ▫ n.d. | 93,1 % | 24 avril 2026 | ✅ Mesuré |
| 85 | Arcee AI: Trinity Large Thinking | Arcee AI | 🟢 Ouvert | 92,0 % | 1 avril 2026 | ✅ Mesuré |
| 86 | DeepSeek V4 Flash | DeepSeek | ▫ n.d. | 92,0 % | 24 avril 2026 | ✅ Mesuré |
| 87 | gpt-5-chat-2025-08-07 | OpenAI | 🔒 Propriétaire | 92,0 % | 7 août 2025 | ✅ Mesuré |
| 88 | Kwaipilot: KAT-Coder-Air V2.5 | kwaipilot | ▫ n.d. | 91,7 % | 10 juillet 2026 | ✅ Mesuré |
| 89 | Kwaipilot: KAT-Coder-Pro V2.5 | kwaipilot | ▫ n.d. | 91,3 % | 10 juillet 2026 | ✅ Mesuré |
| 90 | Inflection: Inflection 3 Pi | Inflection | 🔒 Propriétaire | 90,0 % | 11 octobre 2024 | ✅ Mesuré |
| 91 | OpenAI: GPT-5.6 Terra Pro | OpenAI | 🔒 Propriétaire | 90,0 % | 9 juillet 2026 | ✅ Mesuré |
| 92 | hermes-3-llama-3.1-70b | nousresearch | ▫ n.d. | 90,0 % | — | ✅ Mesuré |
| 93 | qwen3-32b-04-28 | Alibaba Cloud / Qwen Team | ▫ n.d. | 90,0 % | — | ✅ Mesuré |
| 94 | skyfall-36b-v2 | thedrummer | ▫ n.d. | 90,0 % | — | ✅ Mesuré |
| 95 | MiniMax M1 | MiniMax | 🟢 Ouvert | 89,8 % | 17 juin 2025 | ✅ Mesuré |
| 96 | uncensored | venice | ▫ n.d. | 89,8 % | — | ✅ Mesuré |
| 97 | Perceptron: Perceptron Mk1 | Perceptron | ▫ n.d. | 88,4 % | 12 mai 2026 | ✅ Mesuré |
| 98 | GPT-4.1 | OpenAI | 🔒 Propriétaire | 88,0 % | 14 avril 2025 | ✅ Mesuré |
| 99 | Inflection: Inflection 3 Productivity | Inflection | 🔒 Propriétaire | 88,0 % | 11 octobre 2024 | ✅ Mesuré |
| 100 | OpenAI: GPT-5.1-Codex-Max | OpenAI | 🔒 Propriétaire | 88,0 % | 4 décembre 2025 | ✅ Mesuré |
| 101 | Relace: Relace Search | relace | ▫ n.d. | 88,0 % | 8 décembre 2025 | ✅ Mesuré |
| 102 | llama-4-maverick-17b-128e-instruct | Meta | ▫ n.d. | 88,0 % | — | ✅ Mesuré |
| 103 | Claude 3 Haiku | Anthropic | 🔒 Propriétaire | 86,0 % | 13 mars 2024 | ✅ Mesuré |
| 104 | OpenAI: gpt-oss-safeguard-20b | OpenAI | 🟢 Ouvert | 86,0 % | 29 octobre 2025 | ✅ Mesuré |
| 105 | Qwen: Qwen3 30B A3B Thinking 2507 | Alibaba Cloud / Qwen Team | 🟢 Ouvert | 84,0 % | 28 août 2025 | ✅ Mesuré |
| 106 | Tencent: Hy3 preview | Tencent | 🟢 Ouvert | 84,0 % | 22 avril 2026 | ✅ Mesuré |
| 107 | laguna-m.1 | poolside | ▫ n.d. | 84,0 % | — | ✅ Mesuré |
| 108 | ByteDance Seed: Seed 1.6 Flash | ByteDance Seed | ▫ n.d. | 82,0 % | 23 décembre 2025 | ✅ Mesuré |
| 109 | Perplexity: Sonar Pro Search | Perplexity | 🔒 Propriétaire | 80,0 % | 30 octobre 2025 | ✅ Mesuré |
| 110 | nemotron-nano-12b-v2-vl | NVIDIA | 🟢 Ouvert | 80,0 % | 28 octobre 2025 | ✅ Mesuré |
| 111 | Qwen: Qwen3 Coder Flash | Alibaba Cloud / Qwen Team | ▫ n.d. | 78,0 % | 17 septembre 2025 | ✅ Mesuré |
| 112 | command-r-plus-08-2024 | Cohere | ▫ n.d. | 78,0 % | — | ✅ Mesuré |
| 113 | olmo-3-32b-think | allenai | ▫ n.d. | 78,0 % | — | ✅ Mesuré |
| 114 | GPT-4o mini | OpenAI | 🔒 Propriétaire | 76,0 % | 18 juillet 2024 | ✅ Mesuré |
| 115 | unslopnemo-12b | thedrummer | ▫ n.d. | 74,0 % | — | ✅ Mesuré |
| 116 | cydonia-24b-v4.1 | thedrummer | ▫ n.d. | 72,7 % | — | ✅ Mesuré |
| 117 | Mistral: Voxtral Small 24B 2507 | Mistral AI | 🟢 Ouvert | 72,0 % | 30 octobre 2025 | ✅ Mesuré |
| 118 | laguna-xs-2.1 | poolside | ▫ n.d. | 72,0 % | — | ✅ Mesuré |
| 119 | GPT-4.1 mini | OpenAI | 🔒 Propriétaire | 70,0 % | 14 avril 2025 | ✅ Mesuré |
| 120 | Nex AGI: Nex-N2-Pro | Nex AGI | 🟢 Ouvert | 68,0 % | 8 juin 2026 | ✅ Mesuré |
| 121 | llama-4-scout-17b-16e-instruct | Meta | ▫ n.d. | 68,0 % | — | ✅ Mesuré |
| 122 | mistral-small-2603 | mistralai | ▫ n.d. | 66,0 % | — | ✅ Mesuré |
| 123 | ui-tars-1.5-7b | ByteDance | ▫ n.d. | 64,0 % | — | ✅ Mesuré |
| 124 | Mistral NeMo | Mistral AI | ▫ n.d. | 62,0 % | 18 juillet 2024 | ✅ Mesuré |
| 125 | OpenAI: GPT-3.5 Turbo Instruct | OpenAI | 🔒 Propriétaire | 60,0 % | 28 septembre 2023 | ✅ Mesuré |
| 126 | Nex AGI: Nex-N2-Mini | Nex AGI | 🟢 Ouvert | 58,0 % | 24 juin 2026 | ✅ Mesuré |
| 127 | granite-4.0-h-micro | IBM | 🟢 Ouvert | 57,5 % | 20 octobre 2025 | ✅ Mesuré |
| 128 | ByteDance Seed: Seed-2.0-Mini | ByteDance Seed | ▫ n.d. | 56,0 % | 26 février 2026 | ✅ Mesuré |
| 129 | MiniMax: MiniMax-01 | MiniMax | 🟢 Ouvert | 56,0 % | 15 janvier 2025 | ✅ Mesuré |
| 130 | Cohere: Command R7B (12-2024) | Cohere | ▫ n.d. | 54,0 % | 14 décembre 2024 | ✅ Mesuré |
| 131 | AI21: Jamba Large 1.7 | AI21 Labs | 🟢 Ouvert | 52,0 % | 8 août 2025 | ✅ Mesuré |
| 132 | MiniMax: MiniMax M2-her | MiniMax | ▫ n.d. | 52,0 % | 23 janvier 2026 | ✅ Mesuré |
| 133 | Sonar | Perplexity | 🔒 Propriétaire | 52,0 % | 29 janvier 2025 | ✅ Mesuré |
| 134 | Upstage: Solar Pro 3 | Upstage | ▫ n.d. | 52,0 % | 27 janvier 2026 | ✅ Mesuré |
| 135 | l3-lunaris-8b | sao10k | ▫ n.d. | 52,0 % | — | ✅ Mesuré |
| 136 | inclusionAI: Ling-2.6-flash | InclusionAI | ▫ n.d. | 50,0 % | 21 avril 2026 | ✅ Mesuré |
| 137 | nova-2-lite-v1 | Amazon | 🔒 Propriétaire | 46,0 % | 2 décembre 2025 | ✅ Mesuré |
| 138 | Cohere: Command R (08-2024) | Cohere | 🟢 Ouvert | 44,0 % | 30 août 2024 | ✅ Mesuré |
| 139 | gpt-4o-mini-search-preview-2025-03-11 | OpenAI | 🔒 Propriétaire | 28,0 % | 12 mars 2025 | ✅ Mesuré |
| 140 | qwen3-14b-04-28 | Alibaba Cloud / Qwen Team | ▫ n.d. | 28,0 % | — | ✅ Mesuré |
| 141 | reka-edge-2603 | Reka AI | ▫ n.d. | 16,0 % | — | ✅ Mesuré |
| 142 | aion-rp-llama-3.1-8b | Aion Labs | ▫ n.d. | 2,0 % | — | ✅ Mesuré |
| 143 | gemini-3-pro-image-preview | ▫ n.d. | 0,0 % | — | ✅ Mesuré | |
| 144 | ministral-14b-2512 | mistralai | ▫ n.d. | 0,0 % | — | ✅ Mesuré |
| 145 | ministral-3b-2512 | mistralai | ▫ n.d. | 0,0 % | — | ✅ Mesuré |
| 146 | ministral-8b-2512 | mistralai | ▫ n.d. | 0,0 % | — | ✅ Mesuré |
| 147 | mistral-large-2512 | mistralai | ▫ n.d. | 0,0 % | — | ✅ Mesuré |
Classement établi sur 147 modèles évalués, dont 63 de grands éditeurs. Score médian de l'ensemble : 96,0 %.
Notre analyse
Un score élevé signifie que le modèle choisit fréquemment « Je ne sais pas » devant des informations totalement inventées. Un résultat de 100 % correspond à l’absence d’hallucination sur les 50 questions du jeu. La fiabilité bénéficie de scores au moins partiellement mesurés par un tiers, ce qui apporte davantage de rigueur qu’une évaluation reposant uniquement sur des résultats auto-déclarés.
Le score médian de 96 % parmi 235 modèles et le résultat parfait d’AionLabs: Aion-2.0 montrent une forte concentration près du plafond. Cette saturation réduit la capacité du benchmark à départager finement les meilleurs modèles. Son accès public peut aussi favoriser une exposition préalable aux questions, un facteur de contamination à considérer dans l’interprétation. Sa portée reste ciblée : il mesure une réaction à des prémisses entièrement fictives, en anglais, avec une réponse correcte toujours identique. Le classement révèle donc surtout la propension des modèles à exprimer leur incertitude dans ce cadre précis, plutôt qu’une mesure générale de leur exactitude factuelle.
Sources des scores : benchable.