Skip to main content
AIDiveForge AIDiveForge

Performance Data

Sourced benchmark and pricing figures for a subset of models we track — not the full directory. A dash means we do not have a published number. Sources include Artificial Analysis, Vellum, and vendor documentation. This table is not on a daily refresh clock; treat each listing's last-updated time as the freshness signal.

Tool MMLU HumanEval GPQA Context Input $/1M Output $/1M
ChatGPT 128,000 $2.50 $10.00
Claude 91.1% 95.4% 200,000 $3.00 $15.00
Gemini 91.8% 91.9% 1,048,576 $1.25 $10.00
Llama 3 8,192 $2.65 $3.50
Mistral 32,000 $0.15 $0.20
Grok 131,072 $2.00 $10.00
Qwen2.5 72B 32,768 $0.36 $0.40
DBRX Instruct 80.7% 87.5% 32,768 $1.20 $1.20
Mistral Large 2 86.2% 89.8% 48.6% 262,144 $0.50 $1.50
o1 92.3% 94.5% 96.5% 200,000 $15.00 $60.00
Command R7B 128,000 $0.04 $0.15
Grok Code Fast 1 256,000 $1.00 $2.00
Llama 3.2 90B Vision Instruct 128,000 $2.04 $2.04
DeepSeek V3 131,072 $0.28 $0.42
Claude Sonnet 4.5 1,000,000 $3.00 $15.00
Gemini 2.5 Flash 1,048,576 $0.30 $2.50
Llama 4 Scout 131,072 $0.11 $0.34