Skip to main content
AIDiveForge AIDiveForge

Performance Data

Live benchmark scores and pricing for every LLM and AI tool tracked by AIDiveForge. Data is sourced from Artificial Analysis, Vellum, and community submissions — updated daily.

Tool MMLU HumanEval GPQA Context Input $/1M Output $/1M
ChatGPT 128,000 $2.50 $10.00
Claude 91.1% 95.4% 200,000 $3.00 $15.00
Gemini 91.8% 91.9% 1,048,576 $1.25 $10.00
Llama 3 8,192 $2.65 $3.50
Mistral 32,000 $0.15 $0.20
Grok 131,072 $2.00 $10.00
Qwen2.5 72B 32,768 $0.12 $0.39
DBRX Instruct 80.7% 87.5% 32,768 $1.20 $1.20
Mistral Large 2 86.2% 89.8% 48.6% 262,144 $0.50 $1.50
o1 92.3% 94.5% 96.5% 200,000 $15.00 $60.00
Command R7B 128,000 $0.04 $0.15
Grok Code Fast 1 256,000 $0.20 $1.50
Llama 3.2 90B Vision Instruct 128,000 $2.04 $2.04
DeepSeek V3 131,072 $0.28 $0.42
Claude Sonnet 4.5 200,000 $3.00 $15.00
Gemini 2.5 Flash 1,048,576 $0.30 $2.50
Llama 4 Scout 131,072 $0.11 $0.34