Type “llm leaderboard” into Google in September 2026 and you will not get one answer. You will get five. LMArena says one model is on top. Artificial Analysis says another. OpenRouter’s usage rankings say something else entirely, and if you stumble onto an old bookmark for Hugging Face’s Open LLM Leaderboard, you will land on a page that has not updated in over a year. None of these sites are lying. They are measuring different things and calling the result the same word: “best.”
That confusion has real cost. Engineering teams choosing between GPT-6 Astra, Claude Opus 5.5, Claude Sonnet 5.5, Gemini 3.8 Flash, and DeepSeek V4.1 Flash often start their evaluation by screenshotting a leaderboard and pasting it into a Slack channel, without checking whether that leaderboard measures human preference, curated benchmark performance, or raw API traffic. Those are three different questions with three different correct answers. This piece breaks down how LMArena, Artificial Analysis, OpenRouter Rankings, and the smaller trackers (SWFTE, BenchLM, Vellum) actually work, where they agree, where they diverge by dozens of points on the same model, and which one you should actually be checking before you sign an API contract. For the full current rundown of frontier model releases, see our 2026 AI models hub.
Why “LLM Leaderboard” Search Results Don’t Agree With Each Other
The short version: LMArena measures what humans prefer in a blind side-by-side vote. Artificial Analysis measures how a model scores on a fixed basket of graded tests. OpenRouter measures what developers actually pay to run in production. Hugging Face’s original Open LLM Leaderboard measured open-weight models against a static benchmark suite, and Hugging Face itself confirms that leaderboard was archived in June 2024, with a secondary report from Label Your Data pegging final retirement at March 2025 after it had scored more than 13,000 models. None of that stopped it from still showing up in Google’s top 10 for “llm leaderboard” well into 2026, which is part of why so many people land on stale data without realizing it.
Each platform also updates on a different clock. Artificial Analysis pushes its Intelligence Index in numbered releases (it sat at v4.3.2 as of late September 2026, per the company’s own changelog). LMArena’s Bradley-Terry ratings shift with every batch of new votes, sometimes hourly. OpenRouter’s usage rankings move with whatever developers happen to be routing traffic through that week, which means a cheap, fast model can rank above a smarter one simply because it’s the default fallback in someone’s popular open-source framework. None of these are wrong. They’re just not interchangeable, and treating them as interchangeable is the single most common mistake in how engineering teams pick a model.
LMArena: What 6 Million Human Votes Actually Measure
LMArena began as Chatbot Arena, a research project inside UC Berkeley’s Sky Computing Lab tied to the LMSYS group and Databricks/Anyscale co-founder Ion Stoica. The pitch was disarmingly simple: show a user two anonymized model responses to the same prompt, let them vote for the better one, and aggregate those votes into a ranking. It has since become a company. A funding writeup from AgentMarketCap, dated April 2026, puts LMArena’s raise history at roughly $100 million in May 2025 at close to a $600 million valuation, followed by a $150 million Series A on January 6, 2026, at a $1.7 billion post-money valuation, led by Felicis and UC Investments with participation from firms including Kleiner Perkins and Lightspeed. That is a research side-project that turned into a unicorn in under two years.
The mechanics matter for anyone reading the rankings. LMArena no longer runs a pure online Elo system; per the platform’s own methodology description, it now uses a Bradley-Terry statistical model to estimate the probability that one model beats another across pairwise comparisons. A secondary source cited in AI Wiki’s LMArena entry reports the platform has collected more than 6 million user votes across more than 400 models, though that specific figure comes from a third-party writeup rather than an official LMArena statistics page, so treat it as a reported estimate rather than an audited count. Voting itself remains free for anyone who wants to participate.
The limitation is baked into the method. Human raters are not a random sample of enterprise use cases, they tend to reward longer and more confident-sounding answers even when those answers are less accurate, and models with less voting history carry wider uncertainty bands. If your use case is a customer support bot judged on tone and helpfulness, LMArena’s signal is close to what you care about. If your use case is a coding agent that has to compile on the first try, it is measuring something adjacent to what you need, not the thing itself.
How Bradley-Terry and Arena Elo Scoring Actually Work
Most people reading an LMArena leaderboard assume the numbers behind each model work like a chess rating: win a match, gain points; lose one, drop them. That was true in the platform’s early days, when it ran a straightforward online Elo update after every vote. LMArena’s current methodology description states it has since moved to a Bradley-Terry model, a statistical approach that estimates, across the full set of pairwise comparisons collected, the underlying probability that one specific model beats another in a head-to-head matchup. The practical difference matters: Bradley-Terry fits all the accumulated vote data at once rather than updating incrementally after each match, which produces more stable ratings but also means a model’s score can shift when the platform reprocesses its whole dataset, not just when that specific model wins or loses a new vote.
This is also why a model with very few recorded votes can show a misleadingly high or low rating. A newly added model with only a few hundred comparisons carries a wide confidence interval around its score, even if the point estimate looks impressive on the leaderboard’s front page. LMArena typically flags this with a vote-count column or an uncertainty range, and it is worth checking that column before treating any single-digit rank difference between two models as meaningful. A two-point Elo gap between models with thousands of votes each is a real signal. The same two-point gap between a well-established model and one added last week is closer to noise.
Artificial Analysis: The Curated-Benchmark Approach
Artificial Analysis takes the opposite approach: instead of crowdsourced preference, it runs models through a fixed, weighted basket of graded evaluations and produces a single Intelligence Index score from 0 to 100. As of late September 2026, the company’s own site lists Intelligence Index v4.3.2 as combining ten evaluations: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity’s Last Exam, GDP.pdf, CritPt, AA-Omniscience, and AA-LCR v1.1. The index is described on the company’s evaluations page as a weighted average of these production benchmark scores.
On that index, as tracked in late September 2026, Claude Opus 5.5 (adaptive reasoning, max effort) led the pack at a score of roughly 57.6 to 58, followed by Claude Sonnet 5.5 at 56 points, just two behind its more expensive sibling. GPT-6 Astra posted an Intelligence Index of 52.7 on its “max” configuration, alongside a Coding Index of 76.9 and an Agentic Index of 51.0, according to OpenRouter’s own model page for Astra pricing and benchmarks. Gemini 3.8 Flash and DeepSeek V4.1 Flash landed further down most tracker snapshots, in the high-30s to low-40s range depending on which aggregator’s copy of the index you’re reading, since not every tracker updates on the same release cadence as Artificial Analysis itself.
The business model differs from LMArena’s too. Public model-comparison pages and index results are viewable without payment, but Artificial Analysis also runs a commercial side, revenue described by independent tracker Tom Rochette as coming from “enterprise benchmarking-insights subscriptions” and private custom benchmarking work, with a published terms PDF for its paid data platform but no public price list. No funding round has been publicly disclosed for the company, which is a meaningful contrast with LMArena’s venture-backed unicorn status.
GPT-6 Astra’s Self-Reported Scores vs. Independent Verification
GPT-6 Astra is a useful case study in why “the lab’s own benchmark page” and “an independent tracker” are not the same source, even when they’re describing the same model. OpenAI’s own announcement, published at launch on September 3, 2026, credits Astra with saturating several of its hardest internal tests: a 99.9% score on ARC-AGI-3, a 100% score on ExploitBench, and a 98% score on FrontierMath Tier 4, framing the model as having already helped solve open problems in mathematics. Those are OpenAI’s numbers, on OpenAI’s chosen tests, reported by OpenAI.
Independent trackers tell a more mixed story once Astra is placed against a broader, standardized test set rather than the tests OpenAI chose to headline. BenchLM’s model page for GPT-6 Astra lists a Terminal-Bench 4.0 score of 57.90%, explicitly flagged as 12.7 points behind the best verified score on that benchmark at the time, which BenchLM attributes to Claude Sonnet 5.5’s 70.60%. Artificial Analysis’s own tracker, via OpenRouter’s pricing page, separately lists Astra’s Coding Index at 76.9 and its broader Intelligence Index at 52.7 for the “max” configuration, well below Claude Opus 5.5’s 57.6 to 58 on the same index. None of this makes OpenAI’s claims false. ARC-AGI-3, ExploitBench, and FrontierMath Tier 4 are real, difficult tests, and saturating them is a genuine achievement. But a model can lead on one lab’s chosen benchmark set and trail on a broader independent index at the same time, and a reader who only sees the launch-day press release will miss that second half of the picture entirely.
Benchmark Contamination and the Saturation Problem
Every fixed benchmark has a shelf life, and 2026’s leaderboard landscape is shaped by two prior rounds of that shelf life expiring. The first was Hugging Face’s original Open LLM Leaderboard, archived in June 2024 largely because open models had gotten good enough to cluster near the top of its fixed test suite, making the rankings less useful for distinguishing genuinely stronger models from ones that happened to be tuned toward the test distribution. The second, playing out in real time through 2026, is visible in how quickly frontier labs are saturating individual components of even the newer indices: GPT-6 Astra’s reported 99.9% on ARC-AGI-3 and 100% on ExploitBench leave almost no room for a next-generation model to show measurable improvement on those specific tests, which is exactly the pattern that pushed Hugging Face to retire its leaderboard in the first place.
Artificial Analysis’s response to this problem has been to version its index rather than let any single test dominate indefinitely. The jump from whatever preceded it to Intelligence Index v4.3.2 involved swapping in newer, harder evaluations like Humanity’s Last Exam and CritPt specifically because older components had become less discriminating as models improved. That versioning is also why comparing a score from Intelligence Index v4.1.1 directly against a score from v4.3.2, something this article flagged earlier with DeepSeek V4 Pro’s 44.3-versus-36 discrepancy, produces numbers that look contradictory but are actually measuring against different underlying test batteries. The fix isn’t complicated, but it requires discipline: always cite the index version alongside the score, the same way you’d cite a software version number when reporting a bug.
OpenRouter Rankings: What Developers Actually Pay to Run
OpenRouter’s rankings measure neither human preference nor test performance. They measure production traffic, meaning real API calls routed through OpenRouter’s infrastructure by paying developers. A model climbs this leaderboard by being chosen, repeatedly, for real work, which folds in price, latency, context window, and plain habit alongside raw capability. It is the closest thing to a real-world adoption signal that any of these platforms publish, and it is also the easiest to misread, since a model can top the usage chart because it is the default fallback in a popular open-source SDK rather than because it is the strongest available option.
The one hard number available from OpenRouter’s public rankings page as of late September 2026 is telling on its own: DeepSeek V4.1 Flash is listed as having processed roughly 22 trillion tokens through the platform, with usage growing 24% in the period tracked. That single figure captures something neither LMArena nor Artificial Analysis can: DeepSeek’s open-weight Flash model, which shipped September 10, 2026, with native visual understanding and an asymmetric causal encoder-decoder architecture, is one of the most-used models in production on OpenRouter’s network even though it does not top either the human-preference or the curated-benchmark leaderboards. Cheap, fast, and open-weight beats “smartest” for a huge share of real workloads.
Hugging Face Open LLM Leaderboard: Why It Still Shows Up in Search Results
Hugging Face’s original Open LLM Leaderboard is not a current source, and Hugging Face says so directly in its own documentation: the leaderboard was archived in June 2024 and replaced by a newer evaluation approach. A secondary report from Label Your Data adds that the leaderboard’s final retirement came in March 2025, after it had evaluated more than 13,000 open-weight models, and that Hugging Face pulled the plug because its fixed benchmark suite was becoming saturated as models improved past what the tests could meaningfully distinguish. The archived results remain viewable for historical reference, but nobody should be citing a 2024-era Open LLM Leaderboard rank as evidence about GPT-6 Astra, Claude Opus 5.5, or any other model that shipped in 2026.
The reason it keeps surfacing is SEO inertia: the page accumulated years of backlinks and citations before it stopped updating, so it still ranks. That is a useful lesson on its own. A leaderboard’s Google ranking tells you nothing about whether its data is current, and checking the “last updated” date before trusting any benchmark page should be step one, not an afterthought.
The Smaller Trackers: SWFTE, BenchLM, Vellum, and Docsbot
Below the big three sit a cluster of smaller aggregators that scrape or license data from the larger platforms and repackage it into their own tables. SWFTE’s LM Leaderboard, updated as of September 24, 2026, ranks 56 language models by a blend of LMArena Elo, quality score, inference speed, price, and context window, giving each model a “value” score that tries to combine capability and cost into one number. BenchLM publishes per-benchmark breakdowns, listing GPT-6 Astra’s Terminal-Bench 4.0 score at 57.90% against a “best verified” comparison row, a format aimed at people who want the raw benchmark number rather than a composite index. Docsbot.ai runs a comparison tool pulling from Artificial Analysis’s own index alongside pricing and context-window data, useful mainly as a fast side-by-side for two specific models rather than a full leaderboard.
Vellum occupies a different niche. It is primarily a developer evaluation and workflow-testing platform, not a public crowdsourced arena, and its leaderboard reflects benchmark and evaluation tracking oriented at teams building products rather than casual comparison shoppers. Treating Vellum as methodologically equivalent to LMArena would be a mistake: one measures public preference at scale, the other packages benchmark data for engineering decision-making.
Leaderboard Comparison Table: Signal, Cost, and Status
| Platform | Primary Signal | Methodology | Cost to View | Funding/Status | 2026 Update Cadence |
|---|---|---|---|---|---|
| LMArena | Human preference | Blind pairwise voting, Bradley-Terry model | Free (public vote UI) | $150M Series A, Jan 2026, $1.7B valuation | Continuous, live with votes |
| Artificial Analysis | Curated benchmark score | Weighted average of 10 graded evals (v4.3.2) | Free public pages; paid enterprise tier | No disclosed funding; subscription revenue | Numbered index releases |
| OpenRouter Rankings | Real API usage | Tokens routed / requests through the platform | Free to view | Backed by prior venture rounds (platform, not the rankings page itself) | Rolling, usage-window based |
| Hugging Face Open LLM Leaderboard (v1) | Static benchmark suite | Fixed open-source test battery | Free | Archived June 2024; retired ~March 2025 | Not updated (historical only) |
| SWFTE LM Leaderboard | Blended composite | Arena Elo + speed + price + context + value score | Free | Independent tracker/aggregator | Weekly-ish snapshots |
| BenchLM | Per-benchmark scores | Individual test results vs. “best verified” row | Free | Independent tracker/aggregator | Per-model release pages |
| Docsbot.ai Models | Side-by-side comparison | Pulls Artificial Analysis index + pricing data | Free | Independent tracker | Ongoing |
| Vellum AI Leaderboard | Developer evaluation | Platform-run benchmark/eval tracking | Free tier; paid platform | Developer tooling company | Ongoing |
Where the Same Model Gets Different Scores: A Live Example
The clearest way to see how much methodology matters is to watch one model move across trackers. Claude Sonnet 5.5, which Anthropic released September 28, 2026, scored 56 on Artificial Analysis’s headline Intelligence Index configuration (“adaptive reasoning, max effort, default fallback”), putting it in second place overall behind Claude Opus 5.5’s 57.6 to 58. But the same underlying model, run at a different effort setting labeled “medium with fallback,” shows up in Artificial Analysis’s own leaderboard table at a score of 41, a 15-point swing from a single configuration change, not a different model at all. Anyone screenshotting a leaderboard row without checking which configuration it reflects can walk away with a badly wrong impression of a model’s actual ceiling.
Pricing tells a similar story of a model changing shape depending on where you measure it. Anthropic’s own release messaging, echoed by Japanese tech outlet Gigazine’s September 29, 2026 coverage, describes Sonnet 5.5 as 30% faster and 30% cheaper than its predecessor while outperforming GPT-6 Sol on the same Artificial Analysis index. Meanwhile Anthropic priced the model at roughly $2 per million input tokens and $10 per million output tokens according to model-tracking site Opper.ai’s September 2026 release notes, a price point that matches what GPT-6.1 Sol charges on the API side, sits below Claude Opus 5.5’s $4/$20 tier, but stays above Gemini 3.8 Flash’s considerably cheaper $0.75/$3.75 rate. For a broader three-way look at how Sonnet 5.5 stacks up against other mid-tier models, see this Sonnet 5.5 vs. GPT-6 Luna vs. Gemini 3.8 breakdown. None of these numbers contradict each other. They just come from different vantage points on the same launch.
Frontier Model Specs Compared Across the Trackers
| Model | Maker | Release Date | Context Window | Input/Output Price (per 1M tokens) | AA Intelligence Index | Open Weight? |
|---|---|---|---|---|---|---|
| GPT-6 Astra | OpenAI | Sep 3, 2026 | 1,050,000 tokens (128K max output) | $10 / $50 | ~52.7 (max config) | No |
| GPT-6.1 Sol | OpenAI | Sep 2026 | Not disclosed in tracker data | $2 / $10 | ~51.8 | No |
| Claude Opus 5.5 | Anthropic | Sep 22, 2026 | 1M / 128K max output | $4 / $20 (cache read $0.20) | 57.6–58 (#1 tracked) | No |
| Claude Sonnet 5.5 | Anthropic | Sep 28, 2026 | 1M context | $2 / $10 | 56 (#2 tracked) | No |
| Gemini 3.8 Flash | Sep 2, 2026 | 1,048,576 / 65,536 max output | $0.75 / $3.75 (cached $0.075) | ~40.9 | No | |
| DeepSeek V4.1 Flash | DeepSeek | Sep 10, 2026 | Not fully disclosed in tracker data | ~$0.525 blended (tracker estimate) | ~39.5 | Yes (open-weights) |
| DeepSeek V4 Pro (0813) | DeepSeek | Aug 13, 2026 | Not fully disclosed in tracker data | Varies by reasoning-effort tier | ~36–44 (source-dependent) | Yes (open-weights) |
Two rows in that table carry explicit caveats worth repeating in plain language: DeepSeek V4.1 Flash’s exact context window figure was not consistently published across the trackers checked for this piece, and DeepSeek V4 Pro’s Intelligence Index score varies depending on whether you’re reading an Artificial Analysis v4.1.1 snapshot (44.3, per Docsbot’s comparison tool) or a more recent v4.3.2-era snapshot (closer to 36 on the “max effort” configuration). When a benchmark number moves that much between index versions of the same tracker, the honest move is to cite the version number, not just the score.
Real-World Examples: How Teams Actually Use These Leaderboards
Customer support triage. A team building a support chatbot cares primarily about tone and perceived helpfulness across thousands of short, informal exchanges. LMArena’s blind-vote signal is close to a direct proxy for what a support customer would report as “the bot understood me,” which makes it a reasonable first filter even though it says little about factual accuracy.
Coding agents and CI pipelines. A team shipping an autonomous coding agent cares about Terminal-Bench and SWE-style scores far more than human preference, because the agent’s output either compiles and passes tests or it doesn’t. Artificial Analysis’s Coding Index (76.9 for GPT-6 Astra’s max configuration) and BenchLM’s raw Terminal-Bench 4.0 numbers are the more relevant reference points here, not the Arena Elo column.
High-volume, cost-sensitive batch jobs. A team summarizing millions of documents a month is optimizing for price-per-token and throughput above all else. Gemini 3.8 Flash’s $0.75/$3.75 pricing and DeepSeek V4.1 Flash’s open-weight, self-hostable option both undercut the frontier reasoning models by an order of magnitude, and OpenRouter’s usage rankings, where DeepSeek V4.1 Flash’s 22 trillion tokens processed reflects exactly this kind of workload, are the leaderboard that actually matches this use case.
Regulated or safety-sensitive deployments. A healthcare or financial services team evaluating a model for production needs benchmark evidence with a paper trail, not a crowd vote. Artificial Analysis’s numbered, versioned index (v4.3.2, with each of its ten component evaluations named) gives compliance teams something they can cite in documentation. LMArena’s live, constantly shifting Elo rating is much harder to defend in an audit six months later.
Open-source and self-hosted deployments. Teams that need to run models on their own infrastructure, for data residency or cost reasons, are filtering by license before they filter by score. DeepSeek V4.1 Flash and DeepSeek V4 Pro’s open-weight status puts them in consideration where GPT-6 Astra, Claude Opus 5.5, and Gemini 3.8 Flash simply are not options, regardless of benchmark rank.
Vendor procurement and RFP evaluation. A procurement team writing a request-for-proposal document for an enterprise AI deployment needs to justify its shortlist to a budget committee that may not know what LMArena or Artificial Analysis are. In practice, this means citing a versioned, dated score, like Artificial Analysis’s Intelligence Index v4.3.2, rather than a live Elo number that could read differently by the time the committee reviews the document. The archived Hugging Face leaderboard is the clearest example of what not to cite in this context, since a procurement document built on stale 2024 data would misrepresent every model released since.
IDE and developer-tool integrations. Teams building the model picker inside an IDE plugin or coding assistant, in the mold of what GitHub Copilot, Cursor, and similar tools offer, tend to default users toward whichever model tops the usage rankings rather than the raw benchmark leaderboard, on the theory that revealed developer preference under real deadline pressure is a better predictor of day-to-day satisfaction than a lab-reported score. This is part of why DeepSeek V4.1 Flash’s strong OpenRouter usage numbers matter more to a tool vendor’s default settings than its middling Intelligence Index rank.
Pricing Snapshot: What Each Model Costs to Run
| Model | Input Price (per 1M tokens) | Output Price (per 1M tokens) | Cached Input Discount | Notes |
|---|---|---|---|---|
| GPT-6 Astra | $10.00 | $50.00 | $1.00 cached | Highest-cost frontier tier in this comparison |
| GPT-6.1 Sol | $2.00 | $10.00 | Not disclosed | Mid-tier OpenAI reasoning model |
| Claude Opus 5.5 | $4.00 | $20.00 | $0.20 (60% cheaper vs. Opus 5) | 40% cheaper than Claude Opus 5 at launch |
| Claude Sonnet 5.5 | $2.00 | $10.00 | Not disclosed | 30% faster, 30% cheaper than Sonnet 5 |
| Gemini 3.8 Flash | $0.75 | $3.75 | $0.075 cached | Lowest-cost closed-weight model in this set |
| DeepSeek V4.1 Flash | ~$0.525 blended | tracker estimate | N/A | Open-weight; can be self-hosted at compute cost only |
The spread is stark: running GPT-6 Astra’s output tokens costs roughly 13 times what Gemini 3.8 Flash charges for the same volume, and DeepSeek V4.1 Flash’s open-weight license removes the per-token fee from the equation entirely for teams with the infrastructure to self-host. None of the leaderboards discussed above bake cost into their headline number except SWFTE’s blended “value” score and, indirectly, OpenRouter’s usage rankings, where price is one of several factors pulling traffic toward cheaper models.
Migration Guide: Moving From a Single-Leaderboard Habit to a Multi-Source Workflow
Most teams currently pick a model by checking one leaderboard, usually whichever one ranked highest in their last Google search. Moving to a defensible, multi-source evaluation workflow takes roughly five steps.
- Step 1: Separate your use case into a signal type. Decide upfront whether you care more about perceived quality (LMArena), graded task performance (Artificial Analysis), or real-world cost-efficiency at scale (OpenRouter). Write this down before you look at any rankings, so the rankings don’t retroactively justify a model you already liked.
- Step 2: Check the “last updated” date on every source. A leaderboard that hasn’t moved in months, like the archived Hugging Face Open LLM Leaderboard, should be excluded from any 2026 decision regardless of how well it ranks in search.
- Step 3: Note the exact configuration, not just the model name. As the Claude Sonnet 5.5 example above shows, “max effort” and “medium with fallback” configurations of the same model can differ by 15 points on the same index. Record which configuration a cited score reflects.
- Step 4: Cross-reference at least two independent-methodology sources. A model that ranks well on both a human-preference platform (LMArena) and a curated-benchmark platform (Artificial Analysis) is a safer bet than one that only leads on a single metric.
- Step 5: Run your own held-out evaluation set. None of these public leaderboards test your actual prompts, your actual data, or your actual failure modes. Treat every external leaderboard as a shortlist generator, then validate the top two or three candidates against 50-100 real examples from your own workload before committing to a production contract.
Pros and Cons of Each Leaderboard Type
LMArena — Pros: Free, large sample size, captures real human preference across diverse informal prompts, hard to game with a single well-crafted benchmark question. Cons: Not a random sample of enterprise workloads, rewards fluency and confidence over correctness, ratings for newly added models carry high uncertainty, and the underlying company is now a $1.7 billion venture-backed business with commercial incentives that didn’t exist when it was a Berkeley research project.
Artificial Analysis — Pros: Transparent, versioned methodology with ten named component evaluations, comparable across models on a fixed 0-100 scale, useful for compliance documentation. Cons: A fixed test battery can be gamed or become saturated over time the same way Hugging Face’s original leaderboard did, scores shift meaningfully between index versions, and the composite score can obscure which specific capability (coding vs. reasoning vs. agentic tasks) is driving a model’s rank.
OpenRouter Rankings — Pros: Reflects genuine production behavior at scale, hard to fake since it’s tied to real paid API traffic, naturally weights in price and reliability alongside raw capability. Cons: Popularity and quality are not the same thing, a model can rank high because it’s a framework’s default rather than because it’s the best fit, and the rankings say nothing about a model’s ceiling on specialized tasks.
Hugging Face Open LLM Leaderboard (legacy) — Pros: None, in a 2026 context, beyond historical reference for models released before mid-2024. Cons: Archived and not updated, still ranks in search results due to accumulated backlinks, and citing it for a current model comparison is a factual error waiting to happen.
The Verdict: Which Leaderboard Should You Actually Trust?
There isn’t one. That’s the honest verdict, and it’s a more useful answer than picking a favorite. The data backs a specific, three-part rule: use Artificial Analysis when you need a defensible, versioned score for a technical decision, especially anything involving coding or agentic tasks, since its ten-evaluation Intelligence Index (v4.3.2 as of late September 2026) is the closest thing in this space to a documented methodology. Use LMArena as a sanity check on perceived quality for consumer-facing or conversational products, keeping in mind its rapid rise to a $1.7 billion valuation means it now has real commercial incentives layered on top of its research roots. And use OpenRouter’s usage rankings as your reality check on what the rest of the industry is actually paying to run, since a model’s real-world adoption, like DeepSeek V4.1 Flash’s 22 trillion tokens processed with 24% growth, sometimes tells you more than any benchmark score.
On the models themselves, as of this writing: Claude Opus 5.5 leads Artificial Analysis’s Intelligence Index at roughly 58 points, with Claude Sonnet 5.5 close behind at 56 and priced at less than half of Opus 5.5’s rate, making Sonnet 5.5 the stronger default choice for teams that don’t need the absolute top score. GPT-6 Astra remains the strongest coding-specific option in this set with a 76.9 Coding Index score, at a price that is roughly five times Sonnet 5.5’s. Gemini 3.8 Flash and DeepSeek V4.1 Flash both trail on raw intelligence scores but win decisively on cost, with DeepSeek’s open-weight license adding a self-hosting option none of the closed models offer. No single leaderboard captures all four of those tradeoffs in one number, which is exactly why checking only one of them before making a production decision is the mistake this whole comparison is trying to help you avoid.
Frequently Asked Questions
Is the Hugging Face Open LLM Leaderboard still updated in 2026?
No. Hugging Face’s own documentation confirms the original leaderboard was archived in June 2024, and a secondary report says it was formally retired around March 2025 after scoring more than 13,000 models. It still appears in search results because of accumulated backlinks, not because its data is current.
Why does the same model get different scores on different leaderboards?
Because the platforms measure different things by design. LMArena measures which response humans prefer in a blind vote. Artificial Analysis measures performance on a fixed set of graded tests. OpenRouter measures how much real API traffic a model actually gets. A model can lead one and trail another without any of the leaderboards being wrong.
Is LMArena free to use?
The public voting interface is free. The company behind it, however, raised a $150 million Series A in January 2026 at a $1.7 billion valuation, so its data and enterprise services may involve paid tiers beyond the free public vote.
What is the Artificial Analysis Intelligence Index actually made of?
As of version 4.3.2 in late September 2026, it’s a weighted average of ten evaluations: AA-Briefcase v1.1, GDPval-AA v2.1, AutomationBench-AA, Terminal-Bench 4.0, SciCode, Humanity’s Last Exam, GDP.pdf, CritPt, AA-Omniscience, and AA-LCR v1.1, scaled to a 0-100 index.
Which leaderboard should I trust for a coding assistant?
Artificial Analysis’s Coding Index or a raw Terminal-Bench 4.0 score from a tracker like BenchLM will tell you more than a general Arena Elo rating, since coding tasks have objectively verifiable pass/fail outcomes that benefit from graded benchmarks over subjective human preference.
Do any of these leaderboards account for pricing?
Only indirectly. SWFTE’s LM Leaderboard computes a blended “value” score combining quality, speed, price, and context window. OpenRouter’s usage rankings indirectly favor cheaper models because price influences what developers actually choose to run in production. LMArena and Artificial Analysis’s headline scores do not factor in cost at all.
Has anyone proven a leaderboard was manipulated by an AI lab?
No verified, sourced case of deliberate manipulation of these specific platforms by a named AI lab was found as of late September 2026. Structural vulnerabilities, like benchmark contamination, prompt selection, or coordinated voting, are documented as theoretical risks in how these systems work, but that is different from proof of an actual manipulation event.
Should I build my own leaderboard instead of trusting public ones?
For any production decision that matters, yes, at least as a final step. Public leaderboards are useful for generating a shortlist of two or three candidate models. Testing those candidates against 50-100 examples pulled from your actual workload is the only way to know how a model performs on your specific data, not a generic benchmark’s.
Why did GPT-6 Astra’s self-reported scores look so much stronger than its independent tracker scores?
Because they’re measuring against different test sets. OpenAI’s launch-day figures (99.9% on ARC-AGI-3, 100% on ExploitBench, 98% on FrontierMath Tier 4) come from tests the company chose to highlight. Independent trackers like BenchLM and Artificial Analysis run Astra against broader, standardized batteries like Terminal-Bench 4.0, where it posted 57.90%, trailing the best verified score on that specific test. Both sets of numbers can be accurate at once.
What does it mean when a benchmark is “saturated”?
It means models are scoring so close to the maximum possible result that the test can no longer meaningfully distinguish a stronger model from a weaker one. This is part of why Hugging Face retired its original Open LLM Leaderboard in 2024-2025, and it’s the same pressure pushing Artificial Analysis to keep versioning its Intelligence Index with newer, harder evaluations like Humanity’s Last Exam and CritPt.