Performance Data
Live benchmark scores and pricing for every LLM and AI tool tracked by AIDiveForge. Data is sourced from Artificial Analysis, Vellum, and community submissions — updated daily.
| Tool | MMLU | HumanEval | GPQA | Context | Input $/1M | Output $/1M |
|---|---|---|---|---|---|---|
| ChatGPT | — | — | — | 128,000 | $2.50 | $10.00 |
| Claude | 91.1% | — | 95.4% | 200,000 | $3.00 | $15.00 |
| Gemini | 91.8% | — | 91.9% | 1,048,576 | $1.25 | $10.00 |
| Llama 3 | — | — | — | 8,192 | $2.65 | $3.50 |
| Mistral | — | — | — | 32,000 | $0.15 | $0.20 |
| Grok | — | — | — | 131,072 | $2.00 | $10.00 |
| Qwen2.5 72B | — | — | — | 32,768 | $0.12 | $0.39 |
| DBRX Instruct | 80.7% | 87.5% | — | 32,768 | $1.20 | $1.20 |
| Mistral Large 2 | 86.2% | 89.8% | 48.6% | 262,144 | $0.50 | $1.50 |
| o1 | 92.3% | 94.5% | 96.5% | 200,000 | $15.00 | $60.00 |
| Command R7B | — | — | — | 128,000 | $0.04 | $0.15 |
| Grok Code Fast 1 | — | — | — | 256,000 | $0.20 | $1.50 |
| Llama 3.2 90B Vision Instruct | — | — | — | 128,000 | $2.04 | $2.04 |
| DeepSeek V3 | — | — | — | 131,072 | $0.28 | $0.42 |
| Claude Sonnet 4.5 | — | — | — | 200,000 | $3.00 | $15.00 |
| Gemini 2.5 Flash | — | — | — | 1,048,576 | $0.30 | $2.50 |
| Llama 4 Scout | — | — | — | 131,072 | $0.11 | $0.34 |