AgMoDB
ModelsAgentsEvalsCompositesVisualizeIndustry
AgMoDB by @mistakeknot
Methodology workbench · live calculation

Composite Benchmark Builder

Combine evidence, define the population, and inspect how every model earns its score.

Loading local draft

01 · Inputs

Source Benchmarks

1/226

02 · Method

Formula

Intelligence Index100%
Population filters

Agent support · any selected

Include or exclude specific models

Selecting any Include entry creates an allowlist. Exclude entries always win.

gpt-oss-20b (low)
gpt-oss-120b (low)
GPT-5.1 Codex mini (high)
gpt-oss-120b (high)
gpt-oss-20b (high)
GPT-5.2 (medium)
GPT-5 nano (high)
Grok-1
GPT-5.2 Codex (xhigh)
o3
GPT-5.2 (Non-reasoning)
GPT-5.2 (xhigh)
GPT-5.1 Codex (high)
GPT-5 mini (high)
Llama 3.3 Instruct 70B
Llama 3.1 Instruct 405B
Llama 3.2 Instruct 90B (Vision)
Llama 3.2 Instruct 11B (Vision)
Llama 4 Maverick
Llama 4 Scout
Gemma 3 1B Instruct
Gemini 3 Pro Preview (low)
Gemma 3 4B Instruct
Gemma 3n E2B Instruct
Gemini 2.5 Flash-Lite Preview (Sep '25) (Non-reasoning)
Gemini 2.5 Flash-Lite Preview (Sep '25) (Reasoning)
Gemini 3 Flash Preview (Non-reasoning)
Gemma 3 270M
Gemini 2.5 Pro
Gemma 3 27B Instruct
Gemma 3 12B Instruct
Gemini 3 Flash Preview (Reasoning)
Gemini 3 Pro Preview (high)
Gemma 3n E4B Instruct
Claude 4.5 Haiku (Reasoning)
Claude 4.5 Sonnet (Non-reasoning)
Claude Opus 4.6 (Adaptive Reasoning, Max Effort)
Claude 4.5 Sonnet (Reasoning)
Claude 4.5 Haiku (Non-reasoning)
Claude Opus 4.6 (Non-reasoning, High Effort)
Claude Opus 4.5 (Non-reasoning)
Claude Opus 4.5 (Reasoning)
Magistral Small 1.2
Magistral Medium 1.2
Mistral Large 3
Mistral Small 3.2
Ministral 3 8B
Ministral 3 3B
Devstral Small 2
Mistral Medium 3.1
Ministral 3 14B
Devstral 2
DeepSeek R1 Distill Llama 70B
DeepSeek V3.2 Speciale
DeepSeek V3.2 (Reasoning)
DeepSeek R1 0528 (May '25)
DeepSeek V3.2 (Non-reasoning)
DeepSeek R1 0528 Qwen3 8B
DeepSeek-OCR
R1 1776
Falcon-H1R-7B
Grok 4.1 Fast (Reasoning)
Grok 3 mini Reasoning (high)
Grok 4
Grok 4.1 Fast (Non-reasoning)
Grok Voice Agent
Grok Code Fast 1
Nova Micro
Nova Premier
Nova 2.0 Omni (Non-reasoning)
Nova 2.0 Omni (medium)
Nova 2.0 Lite (medium)
Nova 2.0 Lite (Non-reasoning)
Nova 2.0 Pro Preview (medium)
Nova 2.0 Omni (low)
Nova 2.0 Pro Preview (low)
Nova 2.0 Pro Preview (Non-reasoning)
Nova 2.0 Lite (low)
Phi-4
Phi-4 Mini Instruct
Phi-4 Multimodal Instruct
LFM2.5-VL-1.6B
LFM2.5-1.2B-Thinking
LFM2.5-1.2B-Instruct
LFM2 2.6B
LFM2 8B A1B
Solar Open 100B (Reasoning)
Solar Pro 2 (Reasoning)
Solar Pro 2 (Non-reasoning)
MiniMax-M2.1
Llama 3.1 Nemotron Instruct 70B
NVIDIA Nemotron Nano 12B v2 VL (Non-reasoning)
NVIDIA Nemotron 3 Nano 30B A3B (Non-reasoning)
NVIDIA Nemotron 3 Nano 30B A3B (Reasoning)
NVIDIA Nemotron Nano 9B V2 (Reasoning)
Llama 3.3 Nemotron Super 49B v1 (Non-reasoning)
NVIDIA Nemotron Nano 12B v2 VL (Reasoning)
Llama Nemotron Super 49B v1.5 (Reasoning)
Llama 3.3 Nemotron Super 49B v1 (Reasoning)
Llama Nemotron Super 49B v1.5 (Non-reasoning)
Llama 3.1 Nemotron Ultra 253B v1 (Reasoning)
NVIDIA Nemotron Nano 9B V2 (Non-reasoning)
Llama 3.1 Nemotron Nano 4B v1.1 (Reasoning)
Kimi K2.5 (Reasoning)
Kimi K2 Thinking
Kimi K2.5 (Non-reasoning)
Kimi K2 0905
Kimi Linear 48B A3B Instruct
Step3 VL 10B
Olmo 3 7B Instruct
Olmo 3.1 32B Think
Olmo 3 7B Think
Olmo 3.1 32B Instruct
Molmo 7B-D
Molmo2-8B
Granite 4.0 H Small
Granite 4.0 H 350M
Granite 4.0 H 1B
Granite 4.0 1B
Granite 4.0 Micro
Granite 4.0 350M
Reka Flash 3
Hermes 4 - Llama-3.1 405B (Reasoning)
DeepHermes 3 - Llama-3.1 8B Preview (Non-reasoning)
Hermes 4 - Llama-3.1 70B (Reasoning)
Hermes 4 - Llama-3.1 70B (Non-reasoning)
Hermes 4 - Llama-3.1 405B (Non-reasoning)
DeepHermes 3 - Mistral 24B Preview (Non-reasoning)
K-EXAONE (Non-reasoning)
EXAONE 4.0 32B (Non-reasoning)
EXAONE 4.0 32B (Reasoning)
Exaone 4.0 1.2B (Non-reasoning)
Exaone 4.0 1.2B (Reasoning)
K-EXAONE (Reasoning)
MiMo-V2-Flash (Reasoning)
MiMo-V2-Flash (Non-reasoning)
ERNIE 4.5 300B A47B
ERNIE 5.0 Thinking Preview
Llama 65B
Cogito v2.1 (Reasoning)
KAT-Coder-Pro V1
INTELLECT-3
Motif-2-12.7B-Reasoning
K2-V2 (low)
K2-V2 (high)
K2 Think V2
K2-V2 (medium)
Mi:dm K 2.5 Pro
Mi:dm K 2.5 Pro Preview
HyperCLOVA X SEED Think (32B)
GLM-4.7 (Non-reasoning)
GLM-4.5-Air
GLM-4.7 (Reasoning)
GLM-4.6V (Reasoning)
GLM-4.6V (Non-reasoning)
GLM-4.7-Flash (Reasoning)
GLM-4.7-Flash (Non-reasoning)
Command A
Apriel-v1.6-15B-Thinker
Jamba 1.7 Mini
Jamba Reasoning 3B
Jamba 1.7 Large
Qwen3 Omni 30B A3B Instruct
Qwen3 Coder 30B A3B Instruct
Qwen3 30B A3B 2507 Instruct
Qwen3 235B A22B 2507 Instruct
Qwen3 VL 30B A3B Instruct
Qwen3 235B A22B 2507 (Reasoning)
Qwen Chat 14B
Qwen3 30B A3B 2507 (Reasoning)
Qwen3 VL 4B (Reasoning)
Qwen3 VL 8B Instruct
Qwen3 VL 4B Instruct
Qwen3 Max Thinking
Qwen3 Next 80B A3B (Reasoning)
Qwen3 Coder 480B A35B Instruct
Qwen3 Max Thinking (Preview)
Qwen3 Next 80B A3B Instruct
Qwen3 1.7B (Reasoning)
Qwen3 VL 235B A22B Instruct
Qwen3 VL 30B A3B (Reasoning)
Qwen3 VL 32B Instruct
Qwen3 4B 2507 (Reasoning)
Qwen3 4B 2507 Instruct
Qwen3 VL 235B A22B (Reasoning)
Qwen3 Coder Next
Qwen3 Omni 30B A3B (Reasoning)
Qwen3 0.6B (Non-reasoning)
Qwen3 0.6B (Reasoning)
Qwen3 1.7B (Non-reasoning)
Qwen3 VL 8B (Reasoning)
Qwen3 VL 32B (Reasoning)
Qwen3 Max
Ling-mini-2.0
Ring-1T
Ling-1T
Ring-flash-2.0
Ling-flash-2.0
Doubao Seed Code
Doubao-Seed-1.8
o1
o1-preview
o1-mini
GPT-4o (Aug '24)
GPT-4o (May '24)
GPT-4 Turbo
GPT-4o (Nov '24)
GPT-4o mini
GPT-3.5 Turbo
GPT-4.1 nano
GPT-5.1 (high)
GPT-5 (minimal)
o4-mini (high)
GPT-4.1
GPT-5 Codex (high)
o3-pro
GPT-5.1 (Non-reasoning)
GPT-5 (high)
GPT-5 nano (medium)
GPT-5 (medium)
GPT-4.1 mini
GPT-4
GPT-5 (low)
GPT-5 mini (medium)
GPT-5 nano (minimal)
o3-mini
GPT-4o mini Realtime (Dec '24)
o3-mini (high)
GPT-5 mini (minimal)
GPT-4.5 (Preview)
GPT-4o (ChatGPT)
GPT-4o Realtime (Dec '24)
GPT-3.5 Turbo (0613)
o1-pro
GPT-5 (ChatGPT)
GPT-4o (March 2025, chatgpt-4o-latest)
Llama 3.1 Instruct 70B
Llama 3.1 Instruct 8B
Llama 3.2 Instruct 3B
Llama 3 Instruct 70B
Llama 3 Instruct 8B
Llama 3.2 Instruct 1B
Llama 2 Chat 7B
Llama 2 Chat 70B
Llama 2 Chat 13B
Gemini 2.0 Pro Experimental (Feb '25)
Gemini 2.0 Flash (experimental)
Gemini 1.5 Pro (Sep '24)
Gemini 2.0 Flash-Lite (Preview)
Gemini 2.0 Flash (Feb '25)
Gemini 1.5 Flash (Sep '24)
Gemini 1.5 Flash-8B
Gemini 1.0 Ultra
Gemma 3n E4B Instruct Preview (May '25)
Gemini 2.5 Flash (Non-reasoning)
Gemini 2.5 Flash Preview (Non-reasoning)
PALM-2
Gemini 2.5 Flash Preview (Sep '25) (Non-reasoning)
Gemini 2.5 Flash (Reasoning)
Gemini 2.5 Flash Preview (Reasoning)
Gemini 2.0 Flash Thinking Experimental (Jan '25)
Gemini 1.5 Pro (May '24)
Gemini 1.5 Flash (May '24)
Gemini 2.5 Flash-Lite (Non-reasoning)
Gemini 1.0 Pro
Gemini 2.5 Pro Preview (May' 25)
Gemini 2.5 Flash Preview (Sep '25) (Reasoning)
Gemini 2.0 Flash-Lite (Feb '25)
Gemini 2.0 Flash Thinking Experimental (Dec '24)
Gemini 2.5 Pro Preview (Mar' 25)
Gemini 2.5 Flash-Lite (Reasoning)
Claude 3.5 Sonnet (Oct '24)
Claude 3.5 Sonnet (June '24)
Claude 3 Opus
Claude 3.5 Haiku
Claude 3 Sonnet
Claude 3 Haiku
Claude Instant
Claude 3.7 Sonnet (Non-reasoning)
Claude 4.1 Opus (Reasoning)
Claude 4.1 Opus (Non-reasoning)
Claude 4 Sonnet (Non-reasoning)
Claude 3.7 Sonnet (Reasoning)
Claude 4 Opus (Non-reasoning)
Claude 4 Sonnet (Reasoning)
Claude 4 Opus (Reasoning)
Claude 2.1
Claude 2.0
Mistral Large 2 (Nov '24)
Mistral Large 2 (Jul '24)
Pixtral Large
Mistral Small 3
Mistral Small (Sep '24)
Mixtral 8x22B Instruct
Mistral Small (Feb '24)
Mistral Large (Feb '24)
Mixtral 8x7B Instruct
Mistral 7B Instruct
Mistral Small 3.1
Mistral Saba
Devstral Medium
Mistral Medium 3
Devstral Small (Jul '25)
Devstral Small (May '25)
Magistral Small 1
Mistral Medium
Magistral Medium 1
DeepSeek R1 Distill Qwen 32B
DeepSeek V3 (Dec '24)
DeepSeek R1 Distill Qwen 14B
DeepSeek-V2.5 (Dec '24)
DeepSeek-Coder-V2
DeepSeek R1 Distill Llama 8B
DeepSeek LLM 67B Chat (V1)
DeepSeek R1 Distill Qwen 1.5B
DeepSeek V3.1 Terminus (Reasoning)
DeepSeek Coder V2 Lite Instruct
DeepSeek R1 (Jan '25)
DeepSeek V3.1 Terminus (Non-reasoning)
DeepSeek V3 0324
DeepSeek V3.2 Exp (Non-reasoning)
DeepSeek-V2.5
DeepSeek V3.1 (Non-reasoning)
DeepSeek V3.1 (Reasoning)
DeepSeek-V2-Chat
DeepSeek V3.2 Exp (Reasoning)
Sonar Pro
Sonar Reasoning Pro
Sonar Reasoning
Sonar
Grok Beta
Grok 3
Grok 4 Fast (Reasoning)
Grok 4 Fast (Non-reasoning)
Grok 3 Reasoning Beta
Grok 2 (Dec '24)
OpenChat 3.5 (1210)
Nova Pro
Nova Lite
Phi-3 Mini Instruct 3.8B
LFM 40B
LFM2 1.2B
Solar Mini
Solar Pro 2 (Preview) (Reasoning)
Solar Pro 2 (Preview) (Non-reasoning)
DBRX Instruct
MiniMax-M2
MiniMax M1 80k
MiniMax M1 40k
Kimi K2
Llama 3.1 Tulu3 405B
OLMo 2 32B
OLMo 2 7B
Olmo 3 32B Think
Granite 3.3 8B (Non-reasoning)
Reka Flash (Sep '24)
Hermes 3 - Llama-3.1 70B
GLM-4.5 (Reasoning)
GLM-4.5V (Reasoning)
GLM-4.6 (Non-reasoning)
GLM-4.6 (Reasoning)
GLM-4.5V (Non-reasoning)
Command-R+ (Apr '24)
Command-R (Mar '24)
Apriel-v1.5-15B-Thinker
Jamba 1.5 Mini
Jamba 1.5 Large
Jamba 1.6 Large
Jamba 1.6 Mini
Arctic Instruct
Qwen2.5 Max
Qwen2.5 Instruct 72B
Qwen2.5 Coder Instruct 32B
Qwen2.5 Turbo
Qwen2 Instruct 72B
Qwen3 8B (Reasoning)
Qwen3 8B (Non-reasoning)
Qwen3 4B (Reasoning)
QwQ 32B-Preview
Qwen2.5 Coder Instruct 7B
QwQ 32B
Qwen2.5 Instruct 32B
Qwen Chat 72B
Qwen1.5 Chat 110B
Qwen3 235B A22B (Reasoning)
Qwen3 32B (Non-reasoning)
Qwen3 30B A3B (Reasoning)
Qwen3 32B (Reasoning)
Qwen3 30B A3B (Non-reasoning)
Qwen3 235B A22B (Non-reasoning)
Qwen3 14B (Reasoning)
Qwen3 4B (Non-reasoning)
Qwen3 14B (Non-reasoning)
Qwen3 Max (Preview)
Seed-OSS-36B-Instruct
MiMo-V2-Flash (Feb 2026)
GLM-5 (Reasoning)
Aurora Alpha
Free Models Router
StepFun: Step 3.5 Flash (free)
Arcee AI: Trinity Large Preview (free)
Upstage: Solar Pro 3 (free)
MiniMax: MiniMax M2-her
Writer: Palmyra X5
LiquidAI: LFM2.5-1.2B-Thinking (free)
LiquidAI: LFM2.5-1.2B-Instruct (free)
OpenAI: GPT Audio
OpenAI: GPT Audio Mini
AllenAI: Molmo2 8B
ByteDance Seed: Seed 1.6 Flash
ByteDance Seed: Seed 1.6
Google: Gemini 3 Flash Preview
Mistral: Mistral Small Creative
OpenAI: GPT-5.2 Chat
OpenAI: GPT-5.2 Pro
Mistral: Devstral 2 2512
Relace: Relace Search
Nex AGI: DeepSeek V3.1 Nex N1
EssentialAI: Rnj 1 Instruct
Body Builder (beta)
OpenAI: GPT-5.1-Codex-Max
Mistral: Ministral 3 14B 2512
Mistral: Ministral 3 8B 2512
Mistral: Ministral 3 3B 2512
Mistral: Mistral Large 3 2512
Arcee AI: Trinity Mini (free)
TNG: R1T Chimera (free)
Google: Nano Banana Pro (Gemini 3 Pro Image Preview)
Google: Gemini 3 Pro Preview
Deep Cogito: Cogito v2.1 671B
OpenAI: GPT-5.1 Chat
Kwaipilot: KAT-Coder-Pro V1
Perplexity: Sonar Pro Search
Mistral: Voxtral Small 24B 2507
OpenAI: gpt-oss-safeguard-20b
LiquidAI: LFM2-2.6B
IBM: Granite 4.0 Micro
OpenAI: GPT-5 Image Mini
Anthropic: Claude Haiku 4.5
Qwen: Qwen3 VL 8B Thinking
OpenAI: GPT-5 Image
OpenAI: o3 Deep Research
OpenAI: o4 Mini Deep Research
NVIDIA: Llama 3.3 Nemotron Super 49B V1.5
Baidu: ERNIE 4.5 21B A3B Thinking
Google: Gemini 2.5 Flash Image (Nano Banana)
Qwen: Qwen3 VL 30B A3B Thinking
OpenAI: GPT-5 Pro
Anthropic: Claude Sonnet 4.5
DeepSeek: DeepSeek V3.2 Exp
TheDrummer: Cydonia 24B V4.1
Relace: Relace Apply 3
Qwen: Qwen3 VL 235B A22B Thinking
Qwen: Qwen3 Coder Plus
Tongyi DeepResearch 30B A3B
Qwen: Qwen3 Coder Flash
OpenGVLab: InternVL3 78B
Qwen: Qwen3 Next 80B A3B Thinking
Meituan: LongCat Flash Chat
Qwen: Qwen Plus 0728
Qwen: Qwen3 30B A3B Thinking 2507
Nous: Hermes 4 70B
Nous: Hermes 4 405B
DeepSeek: DeepSeek V3.1
OpenAI: GPT-4o Audio
Baidu: ERNIE 4.5 21B A3B
Baidu: ERNIE 4.5 VL 28B A3B
AI21: Jamba Large 1.7
OpenAI: GPT-5 Chat
Anthropic: Claude Opus 4.1
Mistral: Codestral 2508
Qwen: Qwen3 30B A3B Instruct 2507
Z.ai: GLM 4.5
Qwen: Qwen3 235B A22B Thinking 2507
Z.ai: GLM 4 32B
Qwen: Qwen3 Coder 480B A35B (free)
ByteDance: UI-TARS 7B
Qwen: Qwen3 235B A22B Instruct 2507
Switchpoint Router
Venice: Uncensored (free)
Google: Gemma 3n 2B (free)
Tencent: Hunyuan A13B Instruct
TNG: DeepSeek R1T2 Chimera (free)
Morph: Morph V3 Large
Morph: Morph V3 Fast
Baidu: ERNIE 4.5 VL 424B A47B
Inception: Mercury
Mistral: Mistral Small 3.2 24B
MiniMax: MiniMax M1
xAI: Grok 3 Mini
Google: Gemini 2.5 Pro Preview 06-05
DeepSeek: R1 0528 (free)
Anthropic: Claude Opus 4
Anthropic: Claude Sonnet 4
Google: Gemma 3n 4B (free)
Google: Gemini 2.5 Pro Preview 05-06
Arcee AI: Spotlight
Arcee AI: Maestro Reasoning
Arcee AI: Virtuoso Large
Arcee AI: Coder Large
Inception: Mercury Coder
Qwen: Qwen3 4B (free)
Meta: Llama Guard 4 12B
Qwen: Qwen3 30B A3B
Qwen: Qwen3 8B
Qwen: Qwen3 14B
Qwen: Qwen3 32B
Qwen: Qwen3 235B A22B
TNG: DeepSeek R1T Chimera (free)
OpenAI: o4 Mini High
EleutherAI: Llemma 7b
AlfredPros: CodeLLaMa 7B Instruct Solidity
xAI: Grok 3 Mini Beta
xAI: Grok 3 Beta
NVIDIA: Llama 3.1 Nemotron Ultra 253B v1
Qwen: Qwen2.5 VL 32B Instruct
DeepSeek: DeepSeek V3 0324
Mistral: Mistral Small 3.1 24B (free)
AllenAI: Olmo 2 32B Instruct
Google: Gemma 3 4B (free)
Google: Gemma 3 12B (free)
OpenAI: GPT-4o-mini Search Preview
OpenAI: GPT-4o Search Preview
Google: Gemma 3 27B (free)
TheDrummer: Skyfall 36B V2
Perplexity: Sonar Deep Research
Llama Guard 3 8B
Google: Gemini 2.0 Flash
Qwen: Qwen VL Plus
AionLabs: Aion-1.0
AionLabs: Aion-1.0-Mini
AionLabs: Aion-RP 1.0 (8B)
Qwen: Qwen VL Max
Qwen: Qwen2.5 VL 72B Instruct
Qwen: Qwen-Plus
Qwen: Qwen-Max
Mistral: Mistral Small 3
MiniMax: MiniMax-01
Sao10K: Llama 3.1 70B Hanami x1
DeepSeek: DeepSeek V3
Sao10K: Llama 3.3 Euryale 70B
Cohere: Command R7B (12-2024)
Meta: Llama 3.3 70B Instruct (free)
OpenAI: GPT-4o (2024-11-20)
Mistral Large 2411
Qwen2.5 Coder 32B Instruct
SorcererLM 8x22B
TheDrummer: UnslopNemo 12B
Magnum v4 72B
Anthropic: Claude 3.5 Sonnet
Qwen: Qwen2.5 7B Instruct
NVIDIA: Llama 3.1 Nemotron 70B Instruct
Inflection: Inflection 3 Pi
Inflection: Inflection 3 Productivity
TheDrummer: Rocinante 12B
Meta: Llama 3.2 3B Instruct (free)
Meta: Llama 3.2 1B Instruct
Meta: Llama 3.2 11B Vision Instruct
Qwen2.5 72B Instruct
NeverSleep: Lumimaid v0.2 8B
Mistral: Pixtral 12B
Cohere: Command R (08-2024)
Cohere: Command R+ (08-2024)
Sao10K: Llama 3.1 Euryale 70B v2.2
Qwen: Qwen2.5-VL 7B Instruct
Nous: Hermes 3 405B Instruct (free)
OpenAI: ChatGPT-4o
Sao10K: Llama 3 8B Lunaris
Meta: Llama 3.1 405B (base)
Meta: Llama 3.1 8B Instruct
Meta: Llama 3.1 405B Instruct
Meta: Llama 3.1 70B Instruct
Mistral: Mistral Nemo
OpenAI: GPT-4o-mini (2024-07-18)
Google: Gemma 2 27B
Google: Gemma 2 9B
Sao10k: Llama 3 Euryale 70B v2.1
NousResearch: Hermes 2 Pro - Llama-3 8B
Mistral: Mistral 7B Instruct v0.3
Meta: LlamaGuard 2 8B
Meta: Llama 3 70B Instruct
Meta: Llama 3 8B Instruct
Mistral: Mixtral 8x22B Instruct
WizardLM-2 8x22B
OpenAI: GPT-4 Turbo Preview
Mistral: Mistral 7B Instruct v0.2
Noromaid 20B
Goliath 120B
Auto Router
OpenAI: GPT-4 Turbo (older v1106)
OpenAI: GPT-3.5 Turbo Instruct
Mistral: Mistral 7B Instruct v0.1
OpenAI: GPT-3.5 Turbo 16k
Mancer: Weaver (alpha)
ReMM SLERP 13B
MythoMax 13B
OpenAI: GPT-4 (older v0314)
OpenAI: GPT-3.5 Turbo
MiniMax-M2.5
GLM-5 (Non-reasoning)
Qwen: Qwen Plus 0728 (thinking)
Qwen: Qwen3 Coder 480B A35B
Qwen: Qwen3 4B
Mistral: Mistral Small 3.1 24B
Meta: Llama 3.3 70B Instruct
Meta: Llama 3.2 3B Instruct
Nous: Hermes 3 405B Instruct
Qwen: Qwen3.5 Plus 2026-02-15
Qwen: Qwen3.5 397B A17B
Qwen3.5 397B A17B (Reasoning)
Claude Sonnet 4.6 (Adaptive Reasoning, Max Effort)
Claude Sonnet 4.6 (Non-reasoning, High Effort)
Amazon: Nova 2 Lite
Amazon: Nova Premier 1.0
Amazon: Nova Lite 1.0
Amazon: Nova Micro 1.0
Amazon: Nova Pro 1.0
Qwen3.5 397B A17B (Non-reasoning)
Gemini 3.1 Pro Preview
Tri-21B-Think
Tri-21B-think Preview
Claude Sonnet 4.6 (Non-reasoning, Low Effort)
Tiny Aya Global
Doubao Seed 2.0 lite (Reasoning)
OpenAI: GPT-5.3-Codex
AionLabs: Aion-2.0
Mercury 2
Google: Gemini 3.1 Pro Preview Custom Tools
Qwen3.5 35B A3B (Reasoning)
Qwen3.5 122B A10B (Reasoning)
Qwen3.5 27B (Reasoning)
Qwen: Qwen3.5-Flash
LiquidAI: LFM2-24B-A2B
ByteDance Seed: Seed-2.0-Mini
Google: Nano Banana 2 (Gemini 3.1 Flash Image Preview)
GPT-5.3 Codex (xhigh)
LFM2 24B A2B
Qwen3.5 122B A10B (Non-reasoning)
Qwen3.5 35B A3B (Non-reasoning)
Qwen3.5 27B (Non-reasoning)
Gemini 3.1 Flash-Lite
Qwen3.5 9B (Reasoning)
Qwen3.5 4B (Reasoning)
Qwen3.5 0.8B (Reasoning)
Qwen3.5 2B (Reasoning)
OpenAI: GPT-5.4 Pro
OpenAI: GPT-5.4
OpenAI: GPT-5.3 Chat
Step 3.5 Flash
GPT-5.4 (xhigh)
GPT-5.4 Pro (xhigh)
Qwen3.5 2B (Non-reasoning)
Qwen3.5 4B (Non-reasoning)
Qwen3.5 9B (Non-reasoning)
LongCat Flash Lite
Nemotron 3 Super 120B A12B (Reasoning)
Qwen3.5 0.8B (Non-reasoning)
Sarvam M (Reasoning)
Grok 4.20 0309 (Non-reasoning)
Grok 4.20 0309 (Reasoning)
Z.ai: GLM 5 Turbo
xAI: Grok 4.20 Multi-Agent Beta
xAI: Grok 4.20 Beta
Hunter Alpha
Healer Alpha
ByteDance Seed: Seed-2.0-Lite
Mistral: Mistral Small 4
OpenAI: GPT-5.4 Nano
OpenAI: GPT-5.4 Mini
MiMo-V2-Pro
Mistral Small 4 (Reasoning)
Mistral Small 4 (Non-reasoning)
MiniMax: MiniMax M2.7
GPT-5.4 (Non-reasoning)
MiniMax-M2.7
Sarvam 30B (high)
Sarvam 105B (high)
Xiaomi: MiMo-V2-Omni
GPT-5.4 mini (xhigh)
GPT-5.4 mini (Non-Reasoning)
GPT-5.4 nano (xhigh)
GPT-5.4 mini (medium)
GPT-5.4 nano (medium)
GPT-5.4 nano (Non-Reasoning)
MiMo-V2-Omni
Nanbeige4.1-3B
Apertus 70B Instruct
Apertus 8B Instruct
GLM-5-Turbo
NVIDIA Nemotron 3 Nano 4B
Reka Edge
Gemini 3 Deep Think
Kwaipilot: KAT-Coder-Pro V2
KAT Coder Pro V2
Nemotron Cascade 2 30B A3B
Google: Lyria 3 Pro Preview
Google: Lyria 3 Clip Preview
Qwen: Qwen3.6 Plus Preview (free)
xAI: Grok 4.20 Multi-Agent
Reka Edge
Z.ai: GLM 5V Turbo
Arcee AI: Trinity Large Thinking
GLM 5V Turbo (Reasoning)
Qwen: Qwen3.6 Plus (free)
Google: Gemma 4 31B
Gemma 4 31B (Reasoning)
Gemma 4 26B A4B (Reasoning)
Qwen3.5 Omni Plus
Gemma 4 E2B (Reasoning)
Gemma 4 E4B (Reasoning)
Gemma 4 E2B (Non-reasoning)
Gemma 4 E4B (Non-reasoning)
Gemma 4 31B (Non-reasoning)
Gemma 4 26B A4B (Non-reasoning)
Solar Pro 3
Nova 2.0 Lite (high)
Muse Spark
MiMo-V2-Omni-0327
Trinity Large Thinking
GLM-5.1 (Non-reasoning)
GLM-5.1 (Reasoning)
Qwen3.6 Plus
Qwen3.5 Omni Flash
Grok 4.20 0309 v2 (Non-reasoning)
Grok 4.20 0309 v2 (Reasoning)
Step 3.5 Flash 2603
JT-MINI
Anthropic: Claude Opus 4.7
Elephant
Anthropic: Claude Opus 4.6 (Fast)
Google: Gemma 4 26B A4B (free)
Meta: Llama Guard 4 12B (free)
Claude Opus 4.7
Claude Mythos Preview
Nanonets OCR-3
Nanonets OCR2+
GLM-OCR
GPT-5.5 (xhigh)
GPT-5.5 (high)
GPT-5.5 (low)
GPT-5.5 (Non-reasoning)
GPT-5.5 (medium)
Claude Opus 4.7 (Non-reasoning, High Effort)
DeepSeek V4 Flash (Reasoning, High Effort)
DeepSeek V4 Pro (Reasoning, High Effort)
DeepSeek V4 Pro Preview
DeepSeek V4 Flash Preview
Kimi K2.6
MiMo-V2.5-Pro
MiMo-V2.5
Qwen3.6 27B (Reasoning)
Qwen3.6 35B A3B (Reasoning)
Qwen3.6 35B A3B (Non-reasoning)
Qwen3.6 27B (Non-reasoning)
Qwen3.6 Max Preview
Ling-2.6-1T
Ling 2.6 Flash
Anthropic Claude Haiku Latest
OpenAI GPT Mini Latest
Google Gemini Pro Latest
MoonshotAI Kimi Latest
Google Gemini Flash Latest
Anthropic Claude Sonnet Latest
OpenAI GPT Latest
Qwen: Qwen3.5 Plus 2026-04-20
Qwen: Qwen3.6 Flash
Qwen: Qwen3.6 Max Preview
OpenAI: GPT-5.5 Pro
Tencent: Hy3 preview (free)
Xiaomi: MiMo-V2.5
OpenAI: GPT-5.4 Image 2
Anthropic: Claude Opus Latest
Pareto Code Router
Baidu: Qianfan-OCR-Fast (free)
EXAONE 4.5 33B (Non-reasoning)
EXAONE 4.5 33B
Hy3-preview (Reasoning)
NVIDIA: Nemotron 3 Nano Omni (free)
Poolside: Laguna XS.2 (free)
Poolside: Laguna M.1 (free)
Granite 4.1 8B
Granite 4.1 30B
Granite 4.1 3B
inclusionAI: Ling-2.6-flash
MiniMax: MiniMax M2.5 (free)
Z.ai: GLM 4.6
Qwen: Qwen3 Next 80B A3B Instruct (free)
MoonshotAI: Kimi K2 0905
OpenAI: gpt-oss-120b (free)
OpenAI: gpt-oss-20b (free)
Z.ai: GLM 4.5 Air (free)
Anthropic: Claude 3.7 Sonnet (thinking)
OpenAI: GPT-4o
GPT-5.4 (low)
DeepSeek V4 Flash (Non-reasoning)
DeepSeek V4 Pro (Non-reasoning)
Kimi K2.6 (Non-reasoning)
MiMo-V2.5-Pro (Non-reasoning)
Owl Alpha
GPT-5.5 Pro (xhigh)
Mistral Medium 3.5
Grok 4.3 (high)
Hy3-preview (Non-reasoning)
NVIDIA Nemotron 3 Nano Omni 30B A3B
OpenAI: GPT Chat Latest
Microsoft: Phi 4 Mini Instruct
Baidu Qianfan: CoBuddy (free)
Google: Gemini 3.1 Flash Lite
inclusionAI: Ling-2.6-1T
Grok 4.3 (Non-reasoning)
inclusionAI: Ring-2.6-1T (free)
MiniCPM-V 4.6 1.3B
ZAYA1-8B
Anthropic: Claude Opus 4.7 (Fast)
Perceptron: Perceptron Mk1
JT-35B-Flash
DeepSeek: DeepSeek V4 Flash (free)
Ring-2.6-1T
Gemini 3.5 Flash (high)
Qwen3.7 Max
Command A+
xAI: Grok Build 0.1
GPT-5.5 Instant (May 2026)
Grok 4.3 (low)
Gemini 3.5 Flash (minimal)
Grok 4.3 (medium)
MiniCPM5-1B (Non-reasoning)
Gemini 3.5 Flash (medium)
MoonshotAI: Kimi K2.6 (free)
Claude Opus 4.8 (Adaptive Reasoning, Max Effort)
Anthropic: Claude Opus 4.8 (Fast)
StepFun: Step 3.7 Flash
MiniMax: MiniMax M3
Step 3.7 Flash
OpenRouter: Fusion
Qwen3.7 Plus
MiniMax-M3
NVIDIA Nemotron 3 Ultra 550B A55B
MiniCPM5-1B (Reasoning)
NVIDIA: Nemotron 3.5 Content Safety (free)
Gemma 4 12B (Reasoning)
LFM2.5-8B-A1B
Nex AGI: Nex-N2-Pro (free)
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)
North Mini Code
Anthropic: Claude Fable Latest
Gemma 4 12B (Non-reasoning)
HyperNova 60B 2605
MoonshotAI: Kimi K2.7 Code
Z.ai: GLM 5.2
Kimi K2.7 Code
GLM-5.2 (max)
Google: Nano Banana 2 (Gemini 3.1 Flash Image)
Google: Nano Banana Pro (Gemini 3 Pro Image)
Grok Build 0.1 0616
Sakana: Fugu Ultra
Nex-N2-Pro
GPT-5.5 Instant (June 2026)
DiffusionGemma 26B A4B
Claude Sonnet 5 (Adaptive Reasoning, Max Effort)
Claude Sonnet 5 (Adaptive Reasoning, High Effort)
Claude Sonnet 5 (Adaptive Reasoning, Medium Effort)
Claude Sonnet 5 (Adaptive Reasoning, Low Effort)
Claude Sonnet 5 (Adaptive Reasoning, Xhigh Effort)
Google: Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image)
Claude Sonnet 5 (Non-reasoning, High Effort)
Poolside: Laguna XS 2.1 (free)
Tencent: Hy3
Nex AGI: Nex-N2-Mini
AionLabs: Aion-3.0-Mini
AionLabs: Aion-3.0
Grok 4.5 (high)
xAI: Grok Latest
GPT-5.6 Terra (medium)
GPT-5.6 Sol (low)
GPT-5.6 Luna (low)
GPT-5.6 Sol (xhigh)
GPT-5.6 Sol (max)
GPT-5.6 Terra (low)
GPT-5.6 Terra (max)
GPT-5.6 Luna (high)
GPT-5.6 Terra (xhigh)
GPT-5.6 Terra (Non-reasoning)
GPT-5.6 Sol (high)
GPT-5.6 Luna (xhigh)
GPT-5.6 Terra (high)
GPT-5.6 Sol (Non-reasoning)
GPT-5.6 Sol (medium)
GPT-5.6 Luna (Non-reasoning)
GPT-5.6 Luna (medium)
GPT-5.6 Luna (max)
OpenAI: GPT-5.6 Luna Pro
OpenAI: GPT-5.6 Terra Pro
OpenAI: GPT-5.6 Sol Pro
JT-4.1 Flash 236B A21B
Muse Spark 1.1 (xhigh)
Gemini 3.6 Flash (high)
Gemini 3.5 Flash-Lite
Claude Opus 5 (Adaptive Reasoning, Max Effort)
Claude Opus 5 (Adaptive Reasoning, High Effort)
Claude Opus 5 (Adaptive Reasoning, Xhigh Effort)
Claude Opus 5 (Adaptive Reasoning, Low Effort)
Claude Opus 5 (Adaptive Reasoning, Medium Effort)
Kimi K3
Motif 3 (Beta)
LongCat 2.0
Inkling (xhigh)
G9v3-3B
GLM-5.2 (Non-reasoning)
Agnes 2.5 Pro Alpha
Claude Opus 5 (Fast)
Ling-3.0-flash (free)
Poolside: Laguna S 2.1
Meituan: LongCat 2.0
Auto Router (Beta)
MoonshotAI: Kimi K3
Kwaipilot: KAT-Coder-Air V2.5
Kwaipilot: KAT-Coder-Pro V2.5
Hy3
Qwen: Qwen3.7 Flash
Claude Opus 5 (batch)
Anthropic: Claude Sonnet 5 (batch)
Anthropic: Claude Fable 5 (batch)
Anthropic: Claude Opus 4.8 (batch)
OpenAI: GPT-5.5 (batch)
Anthropic: Claude Opus 4.6 (batch)
OpenAI: GPT-5.2 (batch)
Anthropic: Claude Opus 4.5 (batch)
OpenAI: GPT-5.1 (batch)
OpenAI: GPT-5 (batch)
OpenAI: GPT-5 Mini (batch)
OpenAI: GPT-5 Nano (batch)
Google: Gemini 3.6 Flash (batch)
Google: Gemini 3.5 Flash Lite (batch)
Google: Gemini 3.5 Flash (batch)
Google: Gemini 3.1 Pro Preview (batch)
Google: Gemini 2.5 Flash Lite (batch)
Google: Gemini 2.5 Flash (batch)
Google: Gemini 2.5 Pro (batch)
Claude Opus 4.7 (Adaptive Reasoning, Max Effort)
Inkling Small
Kimi K3 (low)
DeepSeek V4 Flash 0731
Celeris-1
DeepSeek: DeepSeek V4 Flash 0731
DeepSeek V4 Flash Latest
G9v3-39A5B
Qwen: Qwen3.8 Max
Qwen3.8 Max Preview
Muse Spark 1.2 (xhigh)
Shieldstral 1.0 3B
Apertus v1.5 8B
Apertus v1.5 70B
Leanstral 1.5 119B A6B
Qwen3.8 Max
Ling 3.0 Flash
Thinking Machines: Inkling (batch)
OpenAI: GPT-5.6 Luna (batch)
OpenAI: GPT-5.6 Terra (batch)
OpenAI: GPT-5.6 Sol (batch)
Anthropic: Claude Sonnet 4.6 (batch)
OpenAI: GPT-5 Codex (batch)
OpenAI: o3 Pro (batch)
OpenAI: o3 (batch)
OpenAI: o4 Mini (batch)
OpenAI: GPT-4.1 (batch)
OpenAI: GPT-4.1 Mini (batch)
OpenAI: GPT-4.1 Nano (batch)
OpenAI: o1-pro (batch)
OpenAI: o3 Mini High (batch)
OpenAI: o3 Mini (batch)
OpenAI: o1 (batch)
OpenAI: GPT-4o-mini (batch)
OpenAI: GPT-4 Turbo (batch)
Ling 3.0 Tiny
inclusionAI: Ling 3.0 Tiny (free)
Muse Glimmer (high)
Upstage: Solar Pro 4
Meta: Muse Glimmer 30B
Sakana: Sakana Namazu
Nemotron 3.5 Lightning
NVIDIA: Nemotron 3.5 Lightning
ByteDance Seed: Seed-2.0-Code
LiquidAI: LFM2.5-2.6B (free)
Grok 4.6 (high)
DeepSeek: DeepSeek V4 Pro 0813
ByteDance Seed: Seed 2.1 Turbo
Solar Pro 4
Qwen: Qwen3.8 2.4T A95B
Solar Open2 250B
K-EXAONE 2.0 0803
Motif 3
A.X-K2
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 5 Fallback)
DeepSeek V4 Pro 0813 (Reasoning, Max Effort)
Constraints & coverage

Live methodology

1 Source Benchmark

Snapshot 8/13/2026, 5:52:10 AM
Intelligence Index × 10–100 Composite Score

Result · updates immediately

Leaderboard

629 ranked · 366 excluded

Selected claude-opus-5, rank 1, Composite Score 100.0.

629 ranked models and 366 excluded models. Select a model name to inspect its score evidence.
ChangeContextDecision
1—Anthropic100.0100%0%1052.61Mpareto
2—xhigh reasoningAnthropic99.8100%0%1051.5—dominated
3—Anthropic99.7100%0%2066.81Mdominated
4—high reasoningAnthropic99.5100%0%1050.1—dominated
5—max reasoningOpenAI99.3100%0%11.363.91.1Mdominated
5—high reasoningSpaceXAI99.3100%0%362.8500Kpareto
7—Kimi99.0100%0%638.61.0Mdominated
8—xhigh reasoningOpenAI98.9100%0%11.357.9—dominated
9—medium reasoningAnthropic98.7100%0%1048.9—dominated
10—Alibaba98.6100%0%347.6—dominated
11—Anthropic98.3100%0%10—1Mdominated
11—high reasoningOpenAI98.3100%0%11.357.3—dominated
13—xhigh reasoningMeta98.1100%0%2—1.0Mpareto
14—max reasoningOpenAI97.9100%0%4.5120.51.1Mdominated
15—xhigh reasoningOpenAI97.8100%0%11.3—1.1Mdominated
16—high reasoningSpaceXAI97.6100%0%351500Kdominated
17—medium reasoningOpenAI97.5100%0%11.352.3—dominated
18—Anthropic97.3100%0%480.91Mdominated
19—Anthropic97.1100%0%10——dominated
20—high reasoningOpenAI97.0100%0%11.3——dominated
21—OpenAI96.8100%0%4.890.5400Kdominated
22—max reasoningAlibaba96.7100%0%349.31Mdominated
23—xhigh reasoningMeta96.5100%0%2—1.0Mdominated
24—xhigh reasoningOpenAI96.3100%0%5.6—1.1Mdominated
25—DeepSeek96.2100%0%0.569.3—pareto
26—xhigh reasoningOpenAI96.0100%0%4.596.4—dominated
27—max reasoningZ AI95.9100%0%2.2119.3512Kdominated
28—low reasoningAnthropic95.7100%0%1048.7—dominated
29—max reasoningOpenAI95.5100%0%0.5155.61.1Mpareto
30—high reasoningGoogle95.4100%0%3.4—1.0Mdominated
31—Alibaba95.1100%0%2.933.7262.1Kdominated
31—DeepSeek95.1100%0%0.2115.91.0Mpareto
33—high reasoningGoogle94.9100%0%3215.91.0Mdominated
34—medium reasoningOpenAI94.7100%0%11.3——dominated
35—low reasoningOpenAI94.6100%0%11.349.7—dominated
36—xhigh reasoningOpenAI94.3100%0%0.5146.7—dominated
36—high reasoningOpenAI94.3100%0%4.588.8—dominated
38—MiniMax94.1100%0%0.546204.8Kdominated
39—Xiaomi93.9100%0%0.8—1.0Mdominated
40—OpenAI93.8100%0%1.7174.4400Kdominated
41—Anthropic93.6100%0%6——dominated
42—low reasoningKimi93.5100%0%636.31.0Mdominated
43—Google93.3100%0%4.51241.0Mdominated
44—Motif Technologies93.2100%0%0——pareto
45—high reasoningOpenAI93.0100%0%0.5145.5—dominated
46—medium reasoningOpenAI92.8100%0%4.587—dominated
46—Z AI92.8100%0%1.5—202.8Kdominated
48—medium reasoningGoogle92.4100%0%3.4233.2—dominated
48—max reasoningAlibaba92.4100%0%3.8—1Mdominated
50—xhigh reasoningOpenAI92.2100%0%4.8115.3—dominated
51—MiniMax92.0100%0%0.596.8524.3Kdominated
52—DeepSeek91.8100%0%0.559.91.0Mdominated
52—Motif Technologies91.8100%0%0——dominated
54—Kimi91.6100%0%1.7—256Kdominated
55—Alibaba91.4100%0%1.350.9262.1Kdominated
56—Anthropic91.2100%0%10——dominated
57—low reasoningOpenAI91.1100%0%11.3——dominated
58—Meta90.9100%0%0——dominated
59—OpenAI90.8100%0%0.5163.2400Kdominated
60—non-reasoning reasoningAnthropic90.6100%0%10——dominated
61—KwaiKAT90.4100%0%0.5111.5256Kdominated
62—high reasoningDeepSeek90.3100%0%0.555.61.0Mdominated
63—Xiaomi90.1100%0%0.8—262.1Kdominated
64—xhigh reasoningOpenAI90.0100%0%4.8—400Kdominated
65—Kimi89.8100%0%1.740.5—dominated
66—Xiaomi89.6100%0%0.548.91.0Mdominated
66—Z AI89.6100%0%1.9—202.8Kdominated
68—non-reasoning reasoningAnthropic89.3100%0%457.7—dominated
69—xhigh reasoningThinking Machines89.2100%0%1.875.31.0Mdominated
70—Tencent89.0100%0%0.265.3—dominated
71—DeepSeek88.8100%0%0.2—1.0Mdominated
71—Nex AGI88.8100%0%1140.7—dominated
73—Anthropic88.4100%0%10——dominated
73—non-reasoning reasoningOpenAI88.4100%0%11.353.8—dominated
73—Tencent88.4100%0%086.4262.1Kdominated
76—Upstage88.1100%0%0.566.3—dominated
77—Xiaomi87.9100%0%0—1.0Mdominated
78—low reasoningOpenAI87.7100%0%4.575.9—dominated
79—xhigh reasoningOpenAI87.5100%0%4.8—400Kdominated
79—Thinking Machines87.5100%0%0.5131.5524.3Kdominated
81—max reasoningAlibaba87.3100%0%2.9——dominated
82—Z AI87.1100%0%2.1—202.8Kdominated
83—xhigh reasoningOpenAI86.9100%0%1.7—400Kdominated
84—SpaceXAI86.8100%0%1.3——dominated
85—high reasoningGoogle86.5100%0%4.5——dominated
85—Z AI86.5100%0%1.6—204.8Kdominated
87—Alibaba86.3100%0%1.1—1Mdominated
88—low reasoningOpenAI86.1100%0%5.6——dominated
89—China Mobile86.0100%0%0——dominated
90—Sapiens AI85.7100%0%0.6131.1—dominated
90—xhigh reasoningOpenAI85.7100%0%0.5—400Kdominated
92—Alibaba85.5100%0%0.754.91Mdominated
93—Z AI85.4100%0%0——dominated
94—high reasoningDeepSeek85.2100%0%0.2—1.0Mdominated
95—medium reasoningOpenAI84.9100%0%4.8——dominated
95—medium reasoningOpenAI84.9100%0%0.5140—dominated
95—MiniMax84.9100%0%0.5——dominated
98—Anthropic84.6100%0%10—1Mdominated
99—Google84.4100%0%1.1——dominated
100—NVIDIA84.2100%0%1.1119.71Mdominated
101—SpaceXAI84.0100%0%1.6——dominated
101—Xiaomi84.0100%0%0.278.7—dominated
103—high reasoningSpaceXAI83.8100%0%1.6—1Mdominated
104—InclusionAI83.6100%0%0.1384.5131.1Kdominated
105—Alibaba83.4100%0%1.455.5262.1Kdominated
106—high reasoningOpenAI83.3100%0%3.4—400Kdominated
107—Anthropic82.9100%0%6——dominated
107—Google82.9100%0%0.9326.91.0Mdominated
107—SpaceXAI82.9100%0%3—2Mdominated
107—Upstage82.9100%0%0——dominated
111—Xiaomi82.5100%0%0——dominated
112—Anthropic82.3100%0%649.21Mdominated
113—high reasoningOpenAI82.2100%0%3.4—400Kdominated
114—medium reasoningSpaceXAI82.0100%0%1.6121.1—dominated
115—Anthropic81.8100%0%6—1Mdominated
116—ByteDance Seed81.5100%0%0——dominated
116—non-reasoning reasoningZ AI81.5100%0%2.1——dominated
116—low reasoningSpaceXAI81.5100%0%1.6107.9—dominated
119—Anthropic81.1100%0%3038.1200Kdominated
119—Kimi81.1100%0%1.2—262.1Kdominated
121—Xiaomi80.9100%0%0——dominated
122—minimal reasoningGoogle80.7100%0%3.4217.4—dominated
122—non-reasoning reasoningOpenAI80.7100%0%11.3——dominated
124—non-reasoning reasoningAnthropic80.3100%0%10—200Kdominated
124—high reasoningOpenAI80.3100%0%3.4—400Kdominated
126—non-reasoning reasoningKimi80.1100%0%1.7——dominated
127—Z AI79.9100%0%0——dominated
127—high reasoningOpenAI79.9100%0%3.4—400Kdominated
129—Anthropic79.5100%0%647.2—dominated
129—high reasoningMeta79.5100%0%0.696.1—dominated
131—SK Telecom79.2100%0%0——dominated
131—Google79.2100%0%1.1179.31.0Mdominated
133—non-reasoning reasoningZ AI79.0100%0%2.2119—dominated
134—non-reasoning reasoningOpenAI78.7100%0%4.583.9—dominated
134—medium reasoningOpenAI78.7100%0%3.4——dominated
134—Alibaba78.7100%0%0.8—262.1Kdominated
137—Anthropic78.1100%0%30——dominated
137—Z AI78.1100%0%1—202.8Kdominated
137—KwaiKAT78.1100%0%0.5108.1—dominated
137—MiniMax78.1100%0%0.5—196.6Kdominated
141—Tencent77.7100%0%0.1—262.1Kdominated
142—OpenAI77.4100%0%11.3——dominated
142—LongCat77.4100%0%1.337.9—dominated
142—Alibaba77.4100%0%1.479.3—dominated
145—SpaceXAI77.1100%0%6—256Kdominated
146—Xiaomi76.9100%0%0——dominated
147—low reasoningGoogle76.7100%0%4.5——dominated
147—low reasoningOpenAI76.7100%0%0.5140.2—dominated
149—Kimi76.4100%0%1.1—262.1Kdominated
150—OpenAI76.3100%0%35—200Kdominated
151—non-reasoning reasoningZ AI76.1100%0%1.6——dominated
152—Anthropic75.9100%0%3037.9200Kdominated
152—Anthropic75.9100%0%649.41Mdominated
154—DeepSeek75.6100%0%0.3——dominated
154—Alibaba75.6100%0%1.1134.2262.1Kdominated
156—non-reasoning reasoningAlibaba75.3100%0%1.478.2—dominated
157—Alibaba75.2100%0%1.6—262.1Kdominated
158—MiniMax74.9100%0%0.5—196.6Kdominated
158—Alibaba74.9100%0%0.8156.3262.1Kdominated
160—non-reasoning reasoningDeepSeek74.5100%0%0.5571.0Mdominated
160—low reasoningOpenAI74.5100%0%3.4——dominated
160—Xiaomi74.5100%0%0.2——dominated
163—Anthropic74.2100%0%30——dominated
164—AI9Stars74.0100%0%0——dominated
164—medium reasoningOpenAI74.0100%0%0.7——dominated
166—high reasoningOpenAI73.4100%0%0.7—400Kdominated
166—SpaceXAI73.4100%0%0——dominated
166—Alibaba73.4100%0%1.553.3—dominated
166—non-reasoning reasoningAlibaba73.4100%0%1.456—dominated
166—InclusionAI73.4100%0%0.9124.8262.1Kdominated
171—Anthropic72.8100%0%2109.9200Kdominated
171—DeepSeek72.8100%0%1.9——dominated
171—OpenAI72.8100%0%3.5118.1200Kdominated
174—LG AI Research72.5100%0%0——dominated
175—StepFun72.3100%0%0.4389.8—dominated
176—medium reasoningOpenAI72.1100%0%0.5——dominated
177—medium reasoningOpenAI72.0100%0%1.7——dominated
178—Mistral AI71.8100%0%3121.9262.1Kdominated
179—non-reasoning reasoningKimi71.7100%0%1.2——dominated
180—non-reasoning reasoningAlibaba71.5100%0%0.8——dominated
181—Anthropic71.2100%0%2117.9—dominated
181—non-reasoning reasoningAnthropic71.2100%0%6——dominated
181—Alibaba71.2100%0%0.7—262.1Kdominated
184—Anthropic70.9100%0%6——dominated
185—Google70.7100%0%035.5—dominated
186—non-reasoning reasoningDeepSeek70.5100%0%0.2—1.0Mdominated
186—Z AI70.5100%0%1——dominated
188—OpenAI70.2100%0%11.3116.1—dominated
189—China Mobile70.1100%0%0——dominated
190—KwaiKAT69.8100%0%0——dominated
190—MiniMax69.8100%0%0.5—196.6Kdominated
192—non-reasoning reasoningAnthropic69.6100%0%30——dominated
193—non-reasoning reasoningXiaomi69.4100%0%0.549.5—dominated
194—non-reasoning reasoningOpenAI69.3100%0%5.6——dominated
195—non-reasoning reasoningAlibaba69.1100%0%1.1151.3—dominated
196—non-reasoning reasoningGoogle68.9100%0%1.1——dominated
196—SpaceXAI68.9100%0%0.3——dominated
198—Anthropic68.6100%0%0——dominated
199—non-reasoning reasoningZ AI68.5100%0%1——dominated
200—non-reasoning reasoningOpenAI68.3100%0%0.5132.4—dominated
201—non-reasoning reasoningTencent68.1100%0%0.1——dominated
201—InclusionAI68.1100%0%0.9—262.1Kdominated
203—ByteDance Seed67.7100%0%0——dominated
203—non-reasoning reasoningOpenAI67.7100%0%4.8——dominated
203—StepFun67.7100%0%0.2——dominated
206—Z AI67.4100%0%0.648.9131.1Kdominated
207—Google67.1100%0%0.2—262.1Kdominated
207—high reasoningOpenAI67.1100%0%1.9—200Kdominated
209—non-reasoning reasoningAnthropic66.7100%0%30——dominated
209—non-reasoning reasoningAnthropic66.7100%0%6——dominated
209—StepFun66.7100%0%0.2—256Kdominated
212—DeepSeek66.3100%0%0.3——dominated
212—Google66.3100%0%3.4—1.0Mdominated
214—high reasoningOpenAI66.1100%0%0.7—400Kdominated
215—NVIDIA65.9100%0%0.4140262.1Kdominated
216—Google65.8100%0%0.6—1.0Mdominated
217—Alibaba65.6100%0%2.4——dominated
218—non-reasoning reasoningDeepSeek65.4100%0%0.3—163.8Kdominated
218—non-reasoning reasoningXiaomi65.4100%0%0—262.1Kdominated
220—non-reasoning reasoningSpaceXAI65.1100%0%1.6104.4—dominated
221—non-reasoning reasoningAlibaba65.0100%0%0.8175.9—dominated
222—InclusionAI64.7100%0%0272.4—dominated
222—max reasoningAlibaba64.7100%0%2.4—262.1Kdominated
224—non-reasoning reasoningAlibaba64.5100%0%0.7163.5—dominated
225—Google64.3100%0%0——dominated
226—non-reasoning reasoningAnthropic64.1100%0%296.3—dominated
226—high reasoningOpenAI64.1100%0%0.3151.9131.1Kdominated
228—Kimi63.9100%0%1.1—131.1Kdominated
229—non-reasoning reasoningAnthropic63.6100%0%6—200Kdominated
229—OpenAI63.6100%0%26.3—200Kdominated
231—NVIDIA63.4100%0%0.1270.81Mdominated
232—Google63.1100%0%0——dominated
232—non-reasoning reasoningZ AI63.1100%0%1—202.8Kdominated
234—Z AI62.9100%0%0.2—202.8Kdominated
235—high reasoningSpaceXAI62.7100%0%0.4——dominated
235—non-reasoning reasoningSpaceXAI62.7100%0%3——dominated
237—Cohere62.4100%0%0191.3131.1Kdominated
238—Google62.3100%0%3.4——dominated
239—DeepSeek62.1100%0%0—163.8Kdominated
240—LG AI Research61.9100%0%0——dominated
241—Baidu61.7100%0%0——dominated
241—non-reasoning reasoningGoogle61.7100%0%0.255.9—dominated
243—Google61.4100%0%0.282.8—dominated
243—non-reasoning reasoningSpaceXAI61.4100%0%1.6——dominated
245—medium reasoningAmazon61.1100%0%3.4125.2—dominated
246—SpaceXAI61.0100%0%0—256Kdominated
247—Inception60.8100%0%0.4939.8128Kdominated
248—Alibaba60.7100%0%0.275.9256Kdominated
249—non-reasoning reasoningDeepSeek60.4100%0%0.5—163.8Kdominated
249—non-reasoning reasoningDeepSeek60.4100%0%0.3——dominated
251—ServiceNow60.2100%0%0——dominated
252—non-reasoning reasoningDeepSeek60.0100%0%0.8——dominated
253—medium reasoningAmazon59.8100%0%0.9——dominated
253—Alibaba59.8100%0%0.6124.8262.1Kdominated
255—DeepSeek59.6100%0%0.9——dominated
256—Alibaba59.4100%0%2.6——dominated
257—ServiceNow59.2100%0%0——dominated
257—high reasoningAmazon59.2100%0%0.9197.4—dominated
259—non-reasoning reasoningOpenAI58.9100%0%3.4——dominated
260—non-reasoning reasoningAlibaba58.8100%0%0.278.9—dominated
261—LG AI Research58.6100%0%0——dominated
262—DeepSeek58.3100%0%2.1—64Kdominated
262—non-reasoning reasoningGoogle58.3100%0%0.259.6—dominated
262—Alibaba58.3100%0%0.123.6—dominated
265—Google58.0100%0%0.9——dominated
266—Cohere57.8100%0%029.2256Kdominated
267—high reasoningOpenAI57.6100%0%0.1—400Kdominated
268—Alibaba57.5100%0%2.6——dominated
269—low reasoningAmazon57.3100%0%3.4131.8—dominated
270—Z AI57.0100%0%0——dominated
270—Kimi57.0100%0%1—131.1Kdominated
270—Mistral AI57.0100%0%0.3145.5—dominated
273—OpenAI56.7100%0%3.5—1.0Mdominated
274—Alibaba56.5100%0%2.4——dominated
275—Mistral AI56.1100%0%049.3—dominated
275—medium reasoningOpenAI56.1100%0%0.1——dominated
275—medium reasoningAmazon56.1100%0%0.9216.7—dominated
275—OpenAI56.1100%0%1.9—200Kdominated
275—Alibaba56.1100%0%0.3230.6—dominated
280—non-reasoning reasoningGoogle55.5100%0%0—1.0Mdominated
280—OpenAI55.5100%0%262.5—200Kdominated
282—China Mobile55.3100%0%0——dominated
283—DeepSeek55.0100%0%2.4——dominated
283—SpaceXAI55.0100%0%8—131.1Kdominated
285—ByteDance Seed54.8100%0%0.3——dominated
286—Alibaba54.5100%0%1.2——dominated
286—Arcee AI54.5100%0%0.4215.5—dominated
288—Multiverse Computing54.3100%0%0——dominated
289—Alibaba54.1100%0%3——dominated
290—Alibaba54.0100%0%2.6——dominated
291—Mistral AI53.7100%0%2.8106.6—dominated
291—low reasoningAmazon53.7100%0%0.9207.9—dominated
291—Perplexity53.7100%0%3.5—128Kdominated
294—MiniMax53.3100%0%1——dominated
295—non-reasoning reasoningOpenAI53.1100%0%0.5——dominated
295—NVIDIA53.1100%0%0——dominated
297—Mistral AI52.8100%0%0118.5—dominated
297—Google52.8100%0%0——dominated
299—MBZUAI Institute of Foundation Models52.5100%0%0——dominated
299—LongCat52.5100%0%0——dominated
301—minimal reasoningOpenAI52.2100%0%3.4——dominated
302—Naver52.0100%0%0——dominated
302—OpenAI52.0100%0%28.9——dominated
304—non-reasoning reasoningSpaceXAI51.8100%0%0—2Mdominated
305—Z AI51.4100%0%0.5——dominated
305—non-reasoning reasoningLG AI Research51.4100%0%0——dominated
305—Alibaba51.4100%0%1.9201.3—dominated
308—non-reasoning reasoningOpenAI51.1100%0%1.7——dominated
309—Z AI50.9100%0%0.4—131.1Kdominated
309—low reasoningAmazon50.9100%0%0.9——dominated
311—non-reasoning reasoningSpaceXAI50.6100%0%0.3—2Mdominated
311—Korea Telecom50.6100%0%0——dominated
313—DeepSeek50.3100%0%0.5—163.8Kdominated
314—InclusionAI50.2100%0%0——dominated
315—AI9Stars50.0100%0%0——dominated
316—non-reasoning reasoningAlibaba49.8100%0%0.125.6—dominated
317—Mistral AI49.7100%0%0.856—dominated
318—Prime Intellect49.4100%0%0—131.1Kdominated
318—high reasoningOpenAI49.4100%0%1.9—200Kdominated
320—non-reasoning reasoningZ AI49.2100%0%0.2——dominated
321—OpenAI49.0100%0%0——dominated
322—DeepSeek48.6100%0%0.5——dominated
322—Google48.6100%0%0.2——dominated
322—high reasoningOpenAI48.6100%0%0.1118.4131.1Kdominated
322—SpaceXAI48.6100%0%0——dominated
322—Upstage48.6100%0%0——dominated
327—NVIDIA48.0100%0%0.1321.2262.1Kdominated
327—Alibaba48.0100%0%0.171262.1Kdominated
329—low reasoningOpenAI47.7100%0%0.3158.4—dominated
329—Mistral AI47.7100%0%0.2——dominated
331—OpenAI47.5100%0%0.7—1.0Mdominated
332—Mistral AI47.3100%0%0.8—131.1Kdominated
333—Alibaba47.1100%0%0.8——dominated
334—Meta46.6100%0%0.4124.91.0Mdominated
334—Meta46.6100%0%0.290.3131.1Kdominated
334—Meta46.6100%0%090.3128Kdominated
334—MiniMax46.6100%0%0——dominated
334—NVIDIA46.6100%0%0.1269.2—dominated
334—Upstage46.6100%0%0.3122.9128Kdominated
340—low reasoningOpenAI45.9100%0%0.1151.5—dominated
340—non-reasoning reasoningAmazon45.9100%0%3.4120.6—dominated
340—Alibaba45.9100%0%1.2—262.1Kdominated
343—minimal reasoningOpenAI45.5100%0%0.7——dominated
344—DeepSeek45.1100%0%0.5——dominated
344—non-reasoning reasoningGoogle45.1100%0%0.9—1.0Mdominated
344—high reasoningMBZUAI Institute of Foundation Models45.1100%0%0——dominated
344—InclusionAI45.1100%0%0.2—262.1Kdominated
348—Zyphra44.7100%0%0——dominated
349—OpenAI44.6100%0%0——dominated
350—Alibaba44.4100%0%0.9187.6262.1Kdominated
351—OpenAI44.1100%0%0——dominated
351—Alibaba44.1100%0%0.9—160Kdominated
351—Trillion Labs44.1100%0%0——dominated
354—Google43.7100%0%0——dominated
354—Alibaba43.7100%0%2.6——dominated
356—NVIDIA43.3100%0%1.236.6131.1Kdominated
356—Alibaba43.3100%0%0.8——dominated
356—Alibaba43.3100%0%0.7—32.8Kdominated
359—Google43.0100%0%0——dominated
360—non-reasoning reasoningGoogle42.8100%0%0.278.6—dominated
361—non-reasoning reasoningGoogle42.7100%0%0.2—1.0Mdominated
362—Motif Technologies42.5100%0%0——dominated
363—InclusionAI42.3100%0%0——dominated
363—Amazon42.3100%0%561—dominated
365—medium reasoningMistral AI41.8100%0%0——dominated
365—Meta41.8100%0%0.431.2131.1Kdominated
365—Mistral AI41.8100%0%0.8—131.1Kdominated
365—Upstage41.8100%0%0——dominated
369—Celeris41.2100%0%0.32,146.7—dominated
369—medium reasoningMistral AI41.2100%0%0—131.1Kdominated
369—medium reasoningMBZUAI Institute of Foundation Models41.2100%0%0——dominated
369—NVIDIA41.2100%0%0.475.9—dominated
373—OpenAI40.6100%0%0——dominated
373—non-reasoning reasoningMistral AI40.6100%0%0.3138.9—dominated
373—Trillion Labs40.6100%0%0——dominated
376—Anthropic40.0100%0%0—200Kdominated
376—Google40.0100%0%0——dominated
376—Google40.0100%0%062.3—dominated
376—NVIDIA40.0100%0%0——dominated
380—OpenBMB39.5100%0%0——dominated
380—Alibaba39.5100%0%0——dominated
380—high reasoningSarvam39.5100%0%0.1——dominated
383—Anthropic38.9100%0%30——dominated
383—Mistral AI38.9100%0%0——dominated
383—Google38.9100%0%0——dominated
383—Meta38.9100%0%0169.216.4Kdominated
383—non-reasoning reasoningAmazon38.9100%0%0.9219.2—dominated
388—non-reasoning reasoningOpenBMB38.4100%0%0——dominated
389—non-reasoning reasoningGoogle38.1100%0%0——dominated
389—Perplexity38.1100%0%0——dominated
391—Mistral AI37.9100%0%0.8138.2—dominated
392—Google37.7100%0%0.2——dominated
392—Alibaba37.7100%0%2.6——dominated
394—Mistral AI37.4100%0%0.297.8—dominated
395—OpenAI37.3100%0%4.4—128Kdominated
396—DeepSeek36.9100%0%0—32.8Kdominated
396—Nanbeige36.9100%0%0——dominated
396—Alibaba36.9100%0%1.2—131.1Kdominated
399—AI21 Labs36.5100%0%3.557.4256Kdominated
399—non-reasoning reasoningZ AI36.5100%0%0.4—131.1Kdominated
401—non-reasoning reasoningAlibaba36.3100%0%1.2——dominated
402—Mistral AI36.1100%0%0.2——dominated
403—Google35.9100%0%0——dominated
403—Mistral AI35.9100%0%0——dominated
405—LG AI Research35.6100%0%0——dominated
405—Alibaba35.6100%0%0.7——dominated
407—non-reasoning reasoningAmazon35.3100%0%0.9——dominated
407—Alibaba35.3100%0%1.3——dominated
409—DeepSeek35.0100%0%0——dominated
409—Meta35.0100%0%0.3121.5327.7Kdominated
411—max reasoningAlibaba34.7100%0%0——dominated
412—Google34.3100%0%0——dominated
412—Nous Research34.3100%0%0.289.8—dominated
412—Alibaba34.3100%0%0.4—131.1Kdominated
412—non-reasoning reasoningUpstage34.3100%0%0——dominated
416—Anthropic33.8100%0%6——dominated
416—DeepSeek33.8100%0%0.8—131.1Kdominated
416—Mistral AI33.8100%0%3—65.5Kdominated
419—DeepSeek33.2100%0%0——dominated
419—TII UAE33.2100%0%0——dominated
419—Meta33.2100%0%052.1131.1Kdominated
419—Meta33.2100%0%052.1131.1Kdominated
423—OpenAI32.7100%0%0.2—1.0Mdominated
423—InclusionAI32.7100%0%0.2——dominated
425—Google32.4100%0%0——dominated
425—Alibaba32.4100%0%0.4108.4—dominated
427—OpenAI32.0100%0%4.4—128Kdominated
427—Alibaba32.0100%0%0.5——dominated
427—Perplexity32.0100%0%1—127.1Kdominated
430—Meta31.6100%0%0.781.2—dominated
430—StepFun31.6100%0%0——dominated
432—Alibaba31.4100%0%0.8——dominated
433—Mistral AI31.1100%0%0—131.1Kdominated
433—Alibaba31.1100%0%0——dominated
433—Perplexity31.1100%0%6—200Kdominated
436—Z AI30.5100%0%0.9——dominated
436—Mistral AI30.5100%0%0.2101.7—dominated
436—Mistral AI30.5100%0%0——dominated
436—OpenAI30.5100%0%0.8103.616.4Kdominated
440—Baidu29.9100%0%0.5—123Kdominated
440—NVIDIA29.9100%0%0.953—dominated
440—Meta29.9100%0%0.6468.2Kdominated
440—Alibaba29.9100%0%0.4——dominated
444—Nous Research29.3100%0%1.529.3—dominated
444—NVIDIA29.3100%0%0.385.2—dominated
444—Upstage29.3100%0%0——dominated
447—non-reasoning reasoningGoogle28.7100%0%058.5—dominated
447—IBM28.7100%0%0——dominated
447—Meta28.7100%0%077.4131.1Kdominated
447—NVIDIA28.7100%0%0.1165.7—dominated
451—Google28.2100%0%0—1.0Mdominated
451—non-reasoning reasoningNous Research28.2100%0%1.532.2—dominated
451—NVIDIA28.2100%0%0——dominated
454—non-reasoning reasoningNVIDIA27.8100%0%0.4103.9—dominated
454—non-reasoning reasoningAlibaba27.8100%0%1.2——dominated
456—Google27.2100%0%0——dominated
456—OpenAI27.2100%0%7.5—128Kdominated
456—low reasoningMBZUAI Institute of Foundation Models27.2100%0%0——dominated
456—Kimi27.2100%0%0——dominated
456—NVIDIA27.2100%0%0——dominated
461—Meta26.6100%0%0——dominated
461—non-reasoning reasoningNVIDIA26.6100%0%0——dominated
461—Alibaba26.6100%0%0.7——dominated
464—Alibaba26.2100%0%0——dominated
464—Alibaba26.2100%0%0.3—131.1Kdominated
466—Anthropic25.7100%0%6——dominated
466—OpenAI25.7100%0%0——dominated
466—Liquid AI25.7100%0%0344.4—dominated
466—Allen Institute for AI25.7100%0%0——dominated
470—Mistral AI25.2100%0%0—131.1Kdominated
470—InclusionAI25.2100%0%0.2——dominated
472—Allen Institute for AI25.0100%0%0—65.5Kdominated
473—Google24.7100%0%0——dominated
473—minimal reasoningOpenAI24.7100%0%0.1——dominated
473—SpaceXAI24.7100%0%0——dominated
476—OpenAI24.3100%0%15—128Kdominated
476—Alibaba24.3100%0%0——dominated
478—non-reasoning reasoningUpstage24.0100%0%0——dominated
479—Cohere23.8100%0%4.458.9256Kdominated
479—Amazon23.8100%0%1.4——dominated
481—Google23.3100%0%0——dominated
481—Meta23.3100%0%0——dominated
481—NVIDIA23.3100%0%1.280.6—dominated
481—Alibaba23.3100%0%0——dominated
485—SpaceXAI22.9100%0%0——dominated
486—non-reasoning reasoningNVIDIA22.6100%0%0.1192256Kdominated
486—non-reasoning reasoningNVIDIA22.6100%0%0.1161.8128Kdominated
486—Alibaba22.6100%0%0——dominated
489—Mistral AI22.3100%0%0.1200.6—dominated
490—Mistral AI22.1100%0%3—131.1Kdominated
491—Alibaba21.9100%0%0——dominated
491—Alibaba21.9100%0%0——dominated
493—non-reasoning reasoningZ AI21.5100%0%0.9—65.5Kdominated
493—OpenAI21.5100%0%37.5—8.2Kdominated
493—non-reasoning reasoningAlibaba21.5100%0%0.6——dominated
496—non-reasoning reasoningGoogle20.9100%0%0.2—1.0Mdominated
496—OpenAI20.9100%0%0.3—128Kdominated
496—non-reasoning reasoningNous Research20.9100%0%0.298.2—dominated
496—Mistral AI20.9100%0%0.2——dominated
496—Amazon20.9100%0%0.1——dominated
501—non-reasoning reasoningAlibaba20.4100%0%0.4——dominated
502—DeepSeek20.1100%0%0——dominated
502—Meta20.1100%0%0.6——dominated
502—non-reasoning reasoningAlibaba20.1100%0%0——dominated
505—DeepSeek19.4100%0%0——dominated
505—Google19.4100%0%0——dominated
505—IBM19.4100%0%0.1126.6131.1Kdominated
505—Meta19.4100%0%082.28.2Kdominated
505—high reasoningSarvam19.4100%0%0——dominated
510—Meta18.9100%0%0.1139.860Kdominated
511—DeepSeek18.6100%0%0——dominated
511—non-reasoning reasoningGoogle18.6100%0%0——dominated
511—Mistral AI18.6100%0%0.3—32.8Kdominated
511—Allen Institute for AI18.6100%0%0—65.5Kdominated
515—Google18.1100%0%0——dominated
515—Allen Institute for AI18.1100%0%0.2—65.5Kdominated
517—Meta17.5100%0%0——dominated
517—Alibaba17.5100%0%0.1—131.1Kdominated
517—Perplexity17.5100%0%0——dominated
517—Reka AI17.5100%0%0.4——dominated
517—Upstage17.5100%0%0.2——dominated
522—SpaceXAI17.0100%0%0——dominated
523—non-reasoning reasoningLG AI Research16.7100%0%0——dominated
523—Microsoft16.7100%0%045.2—dominated
523—Alibaba16.7100%0%0——dominated
526—Google16.4100%0%0——dominated
527—non-reasoning reasoningAlibaba16.2100%0%0——dominated
528—Google16.0100%0%0——dominated
528—Alibaba16.0100%0%0——dominated
530—non-reasoning reasoningNous Research15.7100%0%0—32.8Kdominated
530—AI21 Labs15.7100%0%0——dominated
532—IBM15.4100%0%0.17.7—dominated
533—Nous Research15.0100%0%0.7—65.5Kdominated
533—AI21 Labs15.0100%0%3.5——dominated
533—non-reasoning reasoningAlibaba15.0100%0%0.3——dominated
533—Alibaba15.0100%0%0.4105.1—dominated
537—DeepSeek14.5100%0%0——dominated
537—AI21 Labs14.5100%0%0——dominated
537—Allen Institute for AI14.5100%0%0——dominated
540—Google14.0100%0%0——dominated
540—Liquid AI14.0100%0%0——dominated
540—Microsoft14.0100%0%0.244.216.4Kdominated
543—Anthropic13.5100%0%0——dominated
543—IBM13.5100%0%0——dominated
543—Amazon13.5100%0%0.1286.7—dominated
546—Google13.1100%0%0——dominated
546—Mistral AI13.1100%0%0.3——dominated
546—Microsoft13.1100%0%0——dominated
549—Google12.6100%0%0——dominated
549—non-reasoning reasoningNVIDIA12.6100%0%0.3177.6128Kdominated
549—Microsoft12.6100%0%017.7—dominated
552—Mistral AI12.2100%0%6—128Kdominated
552—Alibaba12.2100%0%0—32.8Kdominated
554—Mistral AI11.9100%0%0——dominated
555—Meta11.7100%0%0.1——dominated
555—Meta11.7100%0%0——dominated
557—AI21 Labs11.4100%0%0——dominated
557—OpenBMB11.4100%0%0——dominated
559—Alibaba11.0100%0%0——dominated
559—Alibaba11.0100%0%0——dominated
559—Reka AI11.0100%0%0.494.465.5Kdominated
562—Allen Institute for AI10.7100%0%0—65.5Kdominated
563—Anthropic10.4100%0%0——dominated
563—Anthropic10.4100%0%0.5—200Kdominated
563—Allen Institute for AI10.4100%0%0——dominated
566—InclusionAI10.0100%0%0——dominated
566—Allen Institute for AI10.0100%0%0——dominated
568—Anthropic9.6100%0%0——dominated
568—DeepSeek9.6100%0%0——dominated
568—DeepSeek9.6100%0%0——dominated
571—OpenAI9.1100%0%0.8——dominated
571—medium reasoningMistral AI9.1100%0%3——dominated
571—Mistral AI9.1100%0%0.3——dominated
574—Meta8.8100%0%1.2——dominated
575—Snowflake8.4100%0%0——dominated
575—Liquid AI8.4100%0%0——dominated
575—Meta8.4100%0%0.318.9—dominated
575—Alibaba8.4100%0%0——dominated
579—non-reasoning reasoningAlibaba8.0100%0%0——dominated
580—Google7.8100%0%0——dominated
581—DeepSeek7.6100%0%0——dominated
581—Google7.6100%0%0——dominated
583—Cohere6.8100%0%6——dominated
583—Databricks6.8100%0%0——dominated
583—DeepSeek6.8100%0%0——dominated
583—Meta6.8100%0%0——dominated
583—Meta6.8100%0%0——dominated
583—OpenChat6.8100%0%0——dominated
583—Sarvam6.8100%0%0——dominated
590—LG AI Research6.2100%0%0——dominated
591—non-reasoning reasoningLG AI Research6.0100%0%0——dominated
591—Allen Institute for AI6.0100%0%0.1—65.5Kdominated
593—AI21 Labs5.4100%0%0.3——dominated
593—AI21 Labs5.4100%0%0——dominated
593—Liquid AI5.4100%0%0——dominated
593—Liquid AI5.4100%0%0——dominated
593—Liquid AI5.4100%0%0——dominated
598—IBM4.9100%0%0——dominated
598—Alibaba4.9100%0%0——dominated
600—AI21 Labs4.6100%0%0——dominated
601—Swiss AI Initiative4.2100%0%1.3——dominated
601—Google4.2100%0%0——dominated
601—IBM4.2100%0%0——dominated
601—Mistral AI4.2100%0%0.5—32.8Kdominated
605—non-reasoning reasoningNous Research3.8100%0%0——dominated
606—Anthropic3.3100%0%0——dominated
606—Cohere3.3100%0%0.8——dominated
606—Meta3.3100%0%0——dominated
606—Mistral AI3.3100%0%0.3—32.8Kdominated
606—Alibaba3.3100%0%0——dominated
611—IBM2.8100%0%0——dominated
611—Allen Institute for AI2.8100%0%0——dominated
613—non-reasoning reasoningIBM2.5100%0%0.1——dominated
613—Liquid AI2.5100%0%0—32.8Kdominated
615—non-reasoning reasoningAlibaba2.2100%0%0——dominated
616—Swiss AI Initiative1.0100%0%0.1——dominated
616—Google1.0100%0%0——dominated
616—Google1.0100%0%0——dominated
616—Google1.0100%0%0——dominated
616—Google1.0100%0%0.1——dominated
616—IBM1.0100%0%0——dominated
616—IBM1.0100%0%0——dominated
616—Liquid AI1.0100%0%0——dominated
616—Liquid AI1.0100%0%0356.4—dominated
616—Meta1.0100%0%0——dominated
616—Meta1.0100%0%0.1——dominated
616—non-reasoning reasoningAlibaba1.0100%0%0——dominated
616—Alibaba1.0100%0%0——dominated
616—Cohere1.0100%0%0129.7—dominated
366 excludedby population, evidence, or constraints
  • DeepSeek-OCR
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Grok Voice Agent
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Cogito v2.1 (Reasoning)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Mi:dm K 2.5 Pro Preview
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Doubao-Seed-1.8
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • GPT-4o mini Realtime (Dec '24)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • GPT-4o Realtime (Dec '24)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • GPT-3.5 Turbo (0613)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Aurora Alpha
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Free Models Router
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • StepFun: Step 3.5 Flash (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Arcee AI: Trinity Large Preview (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Upstage: Solar Pro 3 (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • MiniMax: MiniMax M2-her
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Writer: Palmyra X5
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • LiquidAI: LFM2.5-1.2B-Thinking (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • LiquidAI: LFM2.5-1.2B-Instruct (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: GPT Audio
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: GPT Audio Mini
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • AllenAI: Molmo2 8B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • ByteDance Seed: Seed 1.6 Flash
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • ByteDance Seed: Seed 1.6
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Mistral: Mistral Small Creative
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: GPT-5.2 Chat
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: GPT-5.2 Pro
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Mistral: Devstral 2 2512
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Relace: Relace Search
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Nex AGI: DeepSeek V3.1 Nex N1
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • EssentialAI: Rnj 1 Instruct
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Body Builder (beta)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: GPT-5.1-Codex-Max
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Mistral: Ministral 3 14B 2512
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Mistral: Ministral 3 8B 2512
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Mistral: Ministral 3 3B 2512
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Mistral: Mistral Large 3 2512
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Arcee AI: Trinity Mini (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • TNG: R1T Chimera (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Google: Nano Banana Pro (Gemini 3 Pro Image Preview)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Google: Gemini 3 Pro Preview
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Deep Cogito: Cogito v2.1 671B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: GPT-5.1 Chat
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Kwaipilot: KAT-Coder-Pro V1
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Perplexity: Sonar Pro Search
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Mistral: Voxtral Small 24B 2507
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: gpt-oss-safeguard-20b
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • LiquidAI: LFM2-2.6B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • IBM: Granite 4.0 Micro
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: GPT-5 Image Mini
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Qwen: Qwen3 VL 8B Thinking
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: GPT-5 Image
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: o3 Deep Research
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: o4 Mini Deep Research
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • NVIDIA: Llama 3.3 Nemotron Super 49B V1.5
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Baidu: ERNIE 4.5 21B A3B Thinking
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Google: Gemini 2.5 Flash Image (Nano Banana)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Qwen: Qwen3 VL 30B A3B Thinking
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: GPT-5 Pro
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • DeepSeek: DeepSeek V3.2 Exp
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • TheDrummer: Cydonia 24B V4.1
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Relace: Relace Apply 3
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Qwen: Qwen3 VL 235B A22B Thinking
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Qwen: Qwen3 Coder Plus
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Tongyi DeepResearch 30B A3B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Qwen: Qwen3 Coder Flash
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenGVLab: InternVL3 78B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Qwen: Qwen3 Next 80B A3B Thinking
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Meituan: LongCat Flash Chat
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Qwen: Qwen Plus 0728
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Qwen: Qwen3 30B A3B Thinking 2507
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Nous: Hermes 4 70B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Nous: Hermes 4 405B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • DeepSeek: DeepSeek V3.1
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: GPT-4o Audio
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Baidu: ERNIE 4.5 21B A3B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Baidu: ERNIE 4.5 VL 28B A3B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: GPT-5 Chat
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Mistral: Codestral 2508
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Qwen: Qwen3 235B A22B Thinking 2507
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Z.ai: GLM 4 32B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Qwen: Qwen3 Coder 480B A35B (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • ByteDance: UI-TARS 7B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Qwen: Qwen3 235B A22B Instruct 2507
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Switchpoint Router
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Venice: Uncensored (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Google: Gemma 3n 2B (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Tencent: Hunyuan A13B Instruct
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • TNG: DeepSeek R1T2 Chimera (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Morph: Morph V3 Large
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Morph: Morph V3 Fast
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Baidu: ERNIE 4.5 VL 424B A47B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Inception: Mercury
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Mistral: Mistral Small 3.2 24B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • MiniMax: MiniMax M1
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • xAI: Grok 3 Mini
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Google: Gemini 2.5 Pro Preview 06-05
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • DeepSeek: R1 0528 (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Google: Gemma 3n 4B (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Google: Gemini 2.5 Pro Preview 05-06
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Arcee AI: Spotlight
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Arcee AI: Maestro Reasoning
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Arcee AI: Virtuoso Large
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Arcee AI: Coder Large
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Inception: Mercury Coder
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Qwen: Qwen3 4B (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Meta: Llama Guard 4 12B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Qwen: Qwen3 30B A3B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Qwen: Qwen3 8B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Qwen: Qwen3 14B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Qwen: Qwen3 32B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Qwen: Qwen3 235B A22B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • TNG: DeepSeek R1T Chimera (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: o4 Mini High
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • EleutherAI: Llemma 7b
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • AlfredPros: CodeLLaMa 7B Instruct Solidity
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • xAI: Grok 3 Mini Beta
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • xAI: Grok 3 Beta
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • NVIDIA: Llama 3.1 Nemotron Ultra 253B v1
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Qwen: Qwen2.5 VL 32B Instruct
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • DeepSeek: DeepSeek V3 0324
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Mistral: Mistral Small 3.1 24B (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • AllenAI: Olmo 2 32B Instruct
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Google: Gemma 3 4B (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Google: Gemma 3 12B (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: GPT-4o-mini Search Preview
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: GPT-4o Search Preview
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Google: Gemma 3 27B (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • TheDrummer: Skyfall 36B V2
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Perplexity: Sonar Deep Research
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Llama Guard 3 8B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Google: Gemini 2.0 Flash
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Qwen: Qwen VL Plus
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • AionLabs: Aion-1.0
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • AionLabs: Aion-1.0-Mini
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • AionLabs: Aion-RP 1.0 (8B)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Qwen: Qwen VL Max
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Qwen: Qwen2.5 VL 72B Instruct
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Qwen: Qwen-Plus
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Qwen: Qwen-Max
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Mistral: Mistral Small 3
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • MiniMax: MiniMax-01
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Sao10K: Llama 3.1 70B Hanami x1
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Sao10K: Llama 3.3 Euryale 70B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Cohere: Command R7B (12-2024)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: GPT-4o (2024-11-20)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Mistral Large 2411
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Qwen2.5 Coder 32B Instruct
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • SorcererLM 8x22B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • TheDrummer: UnslopNemo 12B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Magnum v4 72B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Anthropic: Claude 3.5 Sonnet
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Qwen: Qwen2.5 7B Instruct
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Inflection: Inflection 3 Pi
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Inflection: Inflection 3 Productivity
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • TheDrummer: Rocinante 12B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Qwen2.5 72B Instruct
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • NeverSleep: Lumimaid v0.2 8B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Mistral: Pixtral 12B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Cohere: Command R (08-2024)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Cohere: Command R+ (08-2024)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Sao10K: Llama 3.1 Euryale 70B v2.2
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Qwen: Qwen2.5-VL 7B Instruct
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Nous: Hermes 3 405B Instruct (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: ChatGPT-4o
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Sao10K: Llama 3 8B Lunaris
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Meta: Llama 3.1 405B (base)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Meta: Llama 3.1 405B Instruct
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Mistral: Mistral Nemo
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: GPT-4o-mini (2024-07-18)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Google: Gemma 2 27B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Google: Gemma 2 9B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Sao10k: Llama 3 Euryale 70B v2.1
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • NousResearch: Hermes 2 Pro - Llama-3 8B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Mistral: Mistral 7B Instruct v0.3
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Meta: LlamaGuard 2 8B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • WizardLM-2 8x22B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: GPT-4 Turbo Preview
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Mistral: Mistral 7B Instruct v0.2
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Noromaid 20B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Goliath 120B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Auto Router
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: GPT-4 Turbo (older v1106)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: GPT-3.5 Turbo Instruct
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Mistral: Mistral 7B Instruct v0.1
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: GPT-3.5 Turbo 16k
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Mancer: Weaver (alpha)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • ReMM SLERP 13B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • MythoMax 13B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: GPT-4 (older v0314)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Qwen: Qwen Plus 0728 (thinking)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Qwen: Qwen3 Coder 480B A35B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Qwen: Qwen3 4B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Mistral: Mistral Small 3.1 24B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Nous: Hermes 3 405B Instruct
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Qwen: Qwen3.5 Plus 2026-02-15
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Amazon: Nova 2 Lite
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Amazon: Nova Premier 1.0
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Amazon: Nova Lite 1.0
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Amazon: Nova Micro 1.0
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Amazon: Nova Pro 1.0
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • AionLabs: Aion-2.0
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Google: Gemini 3.1 Pro Preview Custom Tools
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Qwen: Qwen3.5-Flash
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • LiquidAI: LFM2-24B-A2B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • ByteDance Seed: Seed-2.0-Mini
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Google: Nano Banana 2 (Gemini 3.1 Flash Image Preview)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: GPT-5.4 Pro
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: GPT-5.4
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: GPT-5.3 Chat
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • GPT-5.4 Pro (xhigh)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • xAI: Grok 4.20 Multi-Agent Beta
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • xAI: Grok 4.20 Beta
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Hunter Alpha
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Healer Alpha
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • ByteDance Seed: Seed-2.0-Lite
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Mistral: Mistral Small 4
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Reka Edge
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Gemini 3 Deep Think
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Google: Lyria 3 Pro Preview
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Google: Lyria 3 Clip Preview
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Qwen: Qwen3.6 Plus Preview (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • xAI: Grok 4.20 Multi-Agent
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Reka Edge
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Arcee AI: Trinity Large Thinking
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Qwen: Qwen3.6 Plus (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Google: Gemma 4 31B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Anthropic: Claude Opus 4.7
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Elephant
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Anthropic: Claude Opus 4.6 (Fast)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Google: Gemma 4 26B A4B (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Meta: Llama Guard 4 12B (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Claude Opus 4.7
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Claude Mythos Preview
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Nanonets OCR-3
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Nanonets OCR2+
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • GLM-OCR
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Anthropic Claude Haiku Latest
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI GPT Mini Latest
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Google Gemini Pro Latest
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • MoonshotAI Kimi Latest
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Google Gemini Flash Latest
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Anthropic Claude Sonnet Latest
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI GPT Latest
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Qwen: Qwen3.5 Plus 2026-04-20
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Qwen: Qwen3.6 Flash
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: GPT-5.5 Pro
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: GPT-5.4 Image 2
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Anthropic: Claude Opus Latest
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Pareto Code Router
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Baidu: Qianfan-OCR-Fast (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • EXAONE 4.5 33B (Non-reasoning)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • NVIDIA: Nemotron 3 Nano Omni (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Poolside: Laguna XS.2 (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Poolside: Laguna M.1 (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • inclusionAI: Ling-2.6-flash
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • MiniMax: MiniMax M2.5 (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Z.ai: GLM 4.6
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Qwen: Qwen3 Next 80B A3B Instruct (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • MoonshotAI: Kimi K2 0905
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: gpt-oss-120b (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: gpt-oss-20b (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Z.ai: GLM 4.5 Air (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Anthropic: Claude 3.7 Sonnet (thinking)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: GPT-4o
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Owl Alpha
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • GPT-5.5 Pro (xhigh)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: GPT Chat Latest
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Microsoft: Phi 4 Mini Instruct
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Baidu Qianfan: CoBuddy (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Google: Gemini 3.1 Flash Lite
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • inclusionAI: Ling-2.6-1T
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • inclusionAI: Ring-2.6-1T (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Anthropic: Claude Opus 4.7 (Fast)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Perceptron: Perceptron Mk1
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • DeepSeek: DeepSeek V4 Flash (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • xAI: Grok Build 0.1
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • MoonshotAI: Kimi K2.6 (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Anthropic: Claude Opus 4.8 (Fast)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • StepFun: Step 3.7 Flash
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • MiniMax: MiniMax M3
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenRouter: Fusion
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • NVIDIA: Nemotron 3.5 Content Safety (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Nex AGI: Nex-N2-Pro (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Anthropic: Claude Fable Latest
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • MoonshotAI: Kimi K2.7 Code
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Z.ai: GLM 5.2
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Google: Nano Banana 2 (Gemini 3.1 Flash Image)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Google: Nano Banana Pro (Gemini 3 Pro Image)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Sakana: Fugu Ultra
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Claude Sonnet 5 (Adaptive Reasoning, High Effort)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Claude Sonnet 5 (Adaptive Reasoning, Medium Effort)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Claude Sonnet 5 (Adaptive Reasoning, Low Effort)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Claude Sonnet 5 (Adaptive Reasoning, Xhigh Effort)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Google: Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Poolside: Laguna XS 2.1 (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Tencent: Hy3
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Nex AGI: Nex-N2-Mini
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • AionLabs: Aion-3.0-Mini
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • AionLabs: Aion-3.0
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • xAI: Grok Latest
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: GPT-5.6 Luna Pro
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: GPT-5.6 Terra Pro
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: GPT-5.6 Sol Pro
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Claude Opus 5 (Fast)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Ling-3.0-flash (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Poolside: Laguna S 2.1
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Meituan: LongCat 2.0
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Auto Router (Beta)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • MoonshotAI: Kimi K3
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Kwaipilot: KAT-Coder-Air V2.5
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Kwaipilot: KAT-Coder-Pro V2.5
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Qwen: Qwen3.7 Flash
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Claude Opus 5 (batch)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Anthropic: Claude Sonnet 5 (batch)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Anthropic: Claude Fable 5 (batch)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Anthropic: Claude Opus 4.8 (batch)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: GPT-5.5 (batch)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Anthropic: Claude Opus 4.6 (batch)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: GPT-5.2 (batch)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Anthropic: Claude Opus 4.5 (batch)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: GPT-5.1 (batch)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: GPT-5 (batch)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: GPT-5 Mini (batch)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: GPT-5 Nano (batch)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Google: Gemini 3.6 Flash (batch)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Google: Gemini 3.5 Flash Lite (batch)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Google: Gemini 3.5 Flash (batch)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Google: Gemini 3.1 Pro Preview (batch)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Google: Gemini 2.5 Flash Lite (batch)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Google: Gemini 2.5 Flash (batch)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Google: Gemini 2.5 Pro (batch)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • DeepSeek: DeepSeek V4 Flash 0731
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • DeepSeek V4 Flash Latest
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Qwen: Qwen3.8 Max
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Shieldstral 1.0 3B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Apertus v1.5 8B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Apertus v1.5 70B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Leanstral 1.5 119B A6B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Thinking Machines: Inkling (batch)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: GPT-5.6 Luna (batch)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: GPT-5.6 Terra (batch)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: GPT-5.6 Sol (batch)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Anthropic: Claude Sonnet 4.6 (batch)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: GPT-5 Codex (batch)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: o3 Pro (batch)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: o3 (batch)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: o4 Mini (batch)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: GPT-4.1 (batch)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: GPT-4.1 Mini (batch)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: GPT-4.1 Nano (batch)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: o1-pro (batch)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: o3 Mini High (batch)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: o3 Mini (batch)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: o1 (batch)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: GPT-4o-mini (batch)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • OpenAI: GPT-4 Turbo (batch)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • inclusionAI: Ling 3.0 Tiny (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Upstage: Solar Pro 4
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Meta: Muse Glimmer 30B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Sakana: Sakana Namazu
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • NVIDIA: Nemotron 3.5 Lightning
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • ByteDance Seed: Seed-2.0-Code
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • LiquidAI: LFM2.5-2.6B (free)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • DeepSeek: DeepSeek V4 Pro 0813
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • ByteDance Seed: Seed 2.1 Turbo
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Qwen: Qwen3.8 2.4T A95B
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.
  • Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 5 Fallback)
    • Too few Source Benchmarks contribute to this score.
    • Too few contributions are directly observed.

Why this score?

Why Claude Opus 5 (Adaptive Reasoning, Max Effort) scored 100.0

#1Rank
100.0%Coverage
1Observed
0Predicted

Intelligence Index

Observed
+100.0score points
Raw
63.10
Normalized
100.00
Configured weight
100.0%
Effective weight
100.0%
artificial_analysis
Source date 2026-08-13 · recorded 2026-08-13
Methodology & snapshot

Formula

  1. Intelligence Index × 100.0%higher · percentile
Missing data
reweight available
Evidence
observed only
Population
Population: all models
Coverage gate
0% · 1 observed · 1 contributing
Snapshot
Captured 2026-08-13 · 995 models · 84460 receipts
Integrity warnings1
  • Intelligence Index has moderate contamination riskinfo

    Weighting is set by Artificial Analysis and encodes subjective design choices even when documented; strong performance on one high-value workload can be hidden by weaker results in other categories.

    Remove it, lower its weight, or document why its contamination risk is acceptable.

Sensitivity

Perturb every configured weight by ±20% to test whether small judgment changes reorder the result.