Live rankings of 31 large language models by quality, cost, and speed. Independent benchmark data from Artificial Analysis. Find the best LLM for your workload.
Live LLM leaderboard with 31 language models ranked by independent benchmark scores, pricing, speed, and context window. Filter by task type or price tier to find the best model for your workload.
| # | Model | Tier | Quality | Price (In/Out) | Speed | Context | Value Score |
|---|---|---|---|---|---|---|---|
| 1 | Claude Opus 5 Anthropic | Frontier | 61 | $5.00 / $25.00 | 52 tok/s | 1.0M | 2 |
| 2 | Claude Fable 5 Anthropic | Frontier | 60 | $10.00 / $50.00 | 71 tok/s | 1.0M | 1 |
| 3 | GPT-5.6 Sol OpenAI | Frontier | 59 | $5.00 / $30.00 | 85 tok/s | 1.1M | 2 |
| 4 | Kimi K3 Moonshot AI | Frontier | 57 | $3.00 / $15.00 | 62 tok/s | 1.0M | 3 |
| 5 | Claude Opus 4.8 Anthropic | Frontier | 56 | $5.00 / $25.00 | 30 tok/s | 1.0M | 2 |
| 6 | Gemini 3.7 Flash Google | Mid-Range | 56 | $0.75 / $3.75 | 340 tok/s | 1.0M | 12 |
| 7 | GPT-5.6 Terra OpenAI | Premium | 55 | $2.00 / $12.00 | 75 tok/s | 1.1M | 4 |
| 8 | GPT-5.5 OpenAI | Premium | 55 | $5.00 / $30.00 | 80 tok/s | 1.1M | 2 |
| 9 | Claude Opus 4.7 Anthropic | Premium | 54 | $5.00 / $25.00 | 48 tok/s | 1.0M | 2 |
| 10 | Grok 4.5 xAI | Premium | 54 | $2.00 / $6.00 | 70 tok/s | 500K | 7 |
| 11 | Claude Sonnet 5 Anthropic | Premium | 53 | $2.00 / $10.00 | 78 tok/s | 1.0M | 4 |
| 12 | GPT-5.4 OpenAI | Mid-Range | 51 | $2.50 / $15.00 | 116 tok/s | 1.1M | 3 |
| 13 | GPT-5.6 Luna OpenAI | Budget | 51 | $0.20 / $1.20 | 150 tok/s | 1.1M | 36 |
| 14 | GLM-5.2 Z AI | Mid-Range | 51 | $0.50 / $2.00 | 170 tok/s | 128K | 20 |
| 15 | Gemini 3.6 Flash Google | Mid-Range | 50 | $1.50 / $7.50 | 251 tok/s | 1.0M | 6 |
| 16 | Gemini 3.5 Flash Google | Budget | 50 | $1.50 / $9.00 | 178 tok/s | 1.0M | 5 |
| 17 | Gemini 3.1 Pro Google | Frontier | 47 | $2.50 / $15.00 | 113 tok/s | 1.0M | 3 |
| 18 | Claude Sonnet 4.6 Anthropic | Mid-Range | 47 | $3.00 / $15.00 | 46 tok/s | 1.0M | 3 |
| 19 | Gemini 3 Flash Google | Budget | 46 | $0.07 / $0.30 | 160 tok/s | 1.0M | 123 |
| 20 | DeepSeek V4 Pro DeepSeek | Mid-Range | 44 | $0.43 / $0.87 | 71 tok/s | 128K | 34 |
| 21 | KAT-Coder-Pro V2 KwaiPilot | Mid-Range | 44 | $0.30 / $1.20 | 100 tok/s | 256K | 29 |
| 22 | MiniMax M3 MiniMax | Mid-Range | 44 | $0.30 / $1.20 | 93 tok/s | 205K | 29 |
| 23 | Grok 4 xAI | Premium | 43 | $3.00 / $15.00 | 66 tok/s | 2.0M | 2 |
| 24 | Gemini 3 Pro Google | Premium | 40 | $2.00 / $12.00 | 90 tok/s | 1.0M | 3 |
| 25 | DeepSeek V4 Flash DeepSeek | Budget | 40 | $0.14 / $0.28 | 118 tok/s | 128K | 95 |
| 26 | MiniMax M2.7 MiniMax | Mid-Range | 38 | $0.30 / $1.20 | 51 tok/s | 205K | 25 |
| 27 | GPT-4o Mini OpenAI | Budget | 38 | $0.15 / $0.60 | 180 tok/s | 128K | 51 |
| 28 | Claude Haiku 4.5 Anthropic | Budget | 37 | $1.00 / $5.00 | 95 tok/s | 200K | 6 |
| 29 | DeepSeek R1 (Free) DeepSeek | Free | 27 | Free / Free | 40 tok/s | 64K | 270 |
| 30 | Qwen 2.5 VL 72B (Free) Alibaba | Free | 15 | Free / Free | 50 tok/s | 128K | 150 |
| 31 | Llama 3.3 70B (Free) Meta | Free | 14 | Free / Free | 80 tok/s | 128K | 140 |
Most LLM leaderboards show a single metric. Quality score, or ELO rating, or some proprietary number. That tells you which model is smartest but nothing about whether you can afford to use it.
This leaderboard combines four dimensions: quality (from the Artificial Analysis Intelligence Index), price per million tokens (from OpenRouter and official APIs), generation speed (tokens per second), and context window size. The Value Score combines quality and price into a single number so you can find models that punch above their cost.
We exclude deprecated models, extremely limited-availability models, and models that don't offer meaningful advantages over alternatives in their price range. The goal is a useful ranking, not an exhaustive list.
Claude Opus 5 now leads the Intelligence Index at 61, edging out Fable 5 (60) and GPT-5.6 Sol (59). At $5/$25, Opus 5 costs half of Fable 5 and matches Sol's input pricing with cheaper output ($25 vs $30). The frontier tier is three-deep and closer than ever.
The biggest shakeup is GPT-5.6 Luna. On July 30, OpenAI cut Luna's price by 80%: from $1/$6 to $0.20/$1.20 per million tokens. Luna scores 51 on the Intelligence Index at a blended cost under $1/M. It now sits on the Pareto frontier, offering the highest intelligence score at its price point by a wide margin.
The mid-tier remains strong. Grok 4.5 (score 54) and Kimi K3 (score 57) handle 90% of tasks that premium models can, at 20-50% of the price. GPT-5.6 Terra also got a 20% price cut to $2/$12. If you're running agents or batch processing, these are where the economics make sense.
For deeper analysis of each model, see our Best LLM in 2026 guide, coding model rankings, GPT-5.6 breakdown, or cost comparison. Use the Model Selector for personalized recommendations.
Free AI optimization and data conversion tools.