LLM Leaderboard: Compare Top AI Models
A manually curated, sourced comparison of current large language models. Set your own priorities below to get a personalized ranking, or browse the full sortable table with a computed value score, context window, and multimodal support. Refreshed periodically rather than pulled from a live feed, so every number carries the date and source it came from.
Build your ranking
Full comparison table
| # | Model | Arena Elo | SWE-bench | Price in/out per 1M | Value* | Context | Multimodal |
|---|
*Value = SWE-bench Verified score divided by output price per 1M tokens, higher means more coding capability per dollar. Only calculated when both figures are published. Context window tiers and multimodal support are approximate and vary by source and configuration, confirm exact limits on the provider's own documentation. Arena Elo, SWE-bench Verified, and pricing sourced from Arena (arena.ai), public SWE-bench trackers (Vals AI, BenchLM, llm-stats), and provider pricing pages. Last verified August 12, 2026.
// outline
How to Read This Leaderboard
Every AI lab claims their model is best, and there's usually a benchmark screenshot to back it up. The honest picture is messier: top models on the main leaderboards now sit within a small margin of each other, and different benchmarks disagree about who's actually first, because they're measuring different things.
This page tracks several different kinds of scores side by side on purpose. Arena Elo measures what real users prefer in blind, head to head comparisons. SWE-bench Verified measures whether a model can fix a real bug in a real codebase. Value score is a metric we calculate ourselves, coding capability divided by output price, to surface which models punch above their price tag. A model can lead on one and trail on another. Use the ranking builder above to weight these by what actually matters for your situation, or browse the full table and sort any column yourself.
What Each Benchmark Actually Measures
- Arena Elo (human preference). Real users send a prompt, get answers from two anonymous models, and vote for the better one. Millions of votes get converted into a chess style Elo rating. It captures overall usefulness and tone, not any single narrow skill.
- SWE-bench Verified (coding ability). Models are given real, unresolved GitHub issues from popular Python projects and must submit a code patch that fixes the bug without breaking existing tests. It's a much closer approximation of real software engineering work than simpler code completion tests.
- Price per million tokens. What the provider charges for input and output tokens through their API. This is separate from subscription pricing for the chat app, and it's what actually determines your bill if you're building on top of a model. Our AI Model Cost Calculator turns these rates into an estimated monthly cost for your own usage.
- Value score. A number we calculate, not one any lab publishes: SWE-bench score divided by output price. It answers a question raw benchmarks can't, which model gives you the most coding capability per dollar spent, and it often tells a very different story than the Arena or SWE-bench rankings alone.
- Context window and multimodal support. How much text a model can consider at once, shown as a tier rather than an exact figure since providers report this differently, and what input types beyond text it accepts, like images, audio, or video.
Why Scores Move Around So Much
Three things drive most of the movement you'll see between one leaderboard snapshot and the next:
- New releases reset the field. Major labs now ship meaningful updates every few weeks, and each one can reshuffle the top of the table.
- Benchmark contamination. When test questions, or very similar ones, existed in a model's training data, it can appear to "solve" a problem by recalling a pattern rather than reasoning fresh. This is a known issue on SWE-bench Verified specifically, which is part of why the harder SWE-bench Pro exists.
- Different scaffolding, different scores. Labs often report benchmark results using their own tuned agent setup, with custom tool access and retry budgets. Independent trackers run a standardized harness instead. Neither is wrong, but the two numbers are not directly comparable, which is exactly why scores for the same model can differ by 10 or more points depending on the source.
What Happened to Claude Fable 5 and Mythos 5
If you've seen references to Claude Fable 5 or Mythos 5 disappearing and reappearing, here's the short version. Anthropic released both models on June 9, 2026. On June 12, access was suspended to comply with US Department of Commerce export controls. The Department lifted those controls on June 30, and Anthropic restored access on July 1, 2026, with Mythos 5 remaining limited to approved partners under Anthropic's Project Glasswing rather than general availability.
It's a useful reminder alongside the benchmark numbers above: a model's availability, not just its score, can change on short notice for reasons that have nothing to do with its capability.
3 Common Leaderboard Mistakes
- Treating a 5 point gap as meaningful. At the top of most leaderboards, small gaps are within normal statistical noise. A model ranked third isn't necessarily worse for your task than the model ranked first.
- Comparing scores from different sources as if they're the same test. A SWE-bench Pro score and a SWE-bench Verified score for two different models tell you almost nothing when compared directly, even though they share a name.
- Picking a model and never checking again. Given how often this field moves, a model you chose six months ago may no longer be the strongest option in its category, and a model you dismissed then may now lead.
Once you've narrowed to a couple of candidates from the table above, our AI Tool Recommendation Quiz can help match the right one to your actual use case, budget, and skill level, and the Prompt Quality Checker helps make sure you're getting a fair test out of whichever model you pick.
Frequently Asked Questions
What is the LMArena or Chatbot Arena leaderboard?
Arena, formerly called LMArena and originally LMSYS Chatbot Arena, is a public platform where real users compare two anonymous AI models side by side and vote for the better answer. Those votes feed a Bradley-Terry rating system, similar to Elo in chess, to produce a human preference ranking.
What is SWE-bench and why do scores vary so much between sources?
SWE-bench tests models on real software bugs pulled from actual GitHub projects, asking them to submit a working code fix. Scores vary between sources because labs run the benchmark on their own tuned agent scaffolding, while independent trackers use a standardized harness, and the two are not directly comparable.
Which AI model is currently ranked first?
Rankings shift by benchmark and change often. As of this page's last update, Claude models led both the Arena human preference leaderboard and the SWE-bench Verified coding leaderboard, with GPT and Gemini close behind in a tight cluster. Check the table above for the current snapshot.
Why do benchmark scores change so often?
New model versions ship every few weeks from major labs, and each release resets part of the ranking. Arena scores also shift gradually as more votes are cast, and benchmark providers periodically update their methodology, which can move scores without any model actually changing.
Are AI benchmarks reliable?
They are useful directional signals, not precise measurements. Top models on Arena often sit within a statistical margin of each other, and coding benchmarks can be affected by contamination, where test questions overlapped with a model's training data. Treat rank order as a rough guide, not a guarantee.
What is benchmark contamination?
Contamination happens when benchmark questions, or very similar ones, appeared in a model's training data before the test was run. The model can then appear to solve a problem by recalling a pattern it has seen before, rather than reasoning through it fresh, which inflates the score.
What is the difference between SWE-bench Verified and SWE-bench Pro?
SWE-bench Verified is the original, now widely saturated benchmark where top models cluster near the high 80s percent. SWE-bench Pro is a harder, newer version designed to reduce contamination, and models typically score 15 to 35 points lower on Pro than they do on Verified.
Should I choose a model based on benchmark rank alone?
No. Benchmarks are a compass, not a map. They narrow the field to a few reasonable candidates, but the only way to know which model actually fits your specific task is to test 2 or 3 top-ranked options on your own real work.
What is the best value AI model right now?
Based on coding score divided by output price, the model with the highest value score on this page is usually a lower-cost option rather than a flagship, since flagship models charge a premium that outpaces their capability edge. Check the "Best value" quick answer card above the table for the current pick, since this shifts whenever pricing or scores update.
What happened to Claude Fable 5 and Mythos 5?
Anthropic released both models on June 9, 2026. On June 12, access was suspended to comply with US Department of Commerce export controls. The Department lifted those controls on June 30, and Anthropic restored access on July 1, 2026, with Mythos 5 remaining limited to approved partners.
How often should I check an LLM leaderboard?
Checking monthly is enough for most people, since that roughly matches how often major labs ship meaningful updates. If you're actively deciding between models for a new project, it's worth a fresh check right before you commit, since a new release can shift the picture quickly.
Summary: A Shortlist, Not a Verdict
Use this table to narrow a field of over 200 tools down to 2 or 3 realistic candidates for your task, then test those candidates directly rather than trusting a single number to decide for you. That combination, a sourced leaderboard plus your own quick test, beats chasing whichever model currently holds the top spot.
Khalid Hussain
Founder of Review Publically. Holds a Master's in Computer Science with professional training in Google Advanced Data Analytics and ML. Cross-checks benchmark trackers and provider documentation before publishing scores on this page, and revisits it periodically as new models ship.
// related reads