The best open-weight model is now three points off the top of the index
Artificial Analysis has Claude Opus 5 leading at 63 and Kimi K3 at 60. The gap that used to be a tier is now a rounding error on a nine-benchmark average.
Artificial Analysis publishes one number per model, averaged from nine evaluations, and its table currently opens with Claude Opus 5 in adaptive reasoning at maximum effort on 63, out of 161 models it has assessed. Claude Fable 5 follows at 62, then Claude Opus 5 at high effort and GPT-5.6 Sol at maximum effort, both on 61. The first model on that list whose weights anyone can download is Kimi K3 at maximum effort, on 60.
| Model | Index | |
|---|---|---|
| Proprietary | ||
| Claude Opus 5 (adaptive reasoning, max) | 63 | |
| Claude Fable 5 (max, Opus 4.8 fallback) | 62 | |
| Claude Opus 5 (adaptive reasoning, high) | 61 | |
| GPT-5.6 Sol (max) | 61 | |
| Open weights | ||
| Kimi K3 (max) | 60 | |
| Qwen3.8 2.4T A95B | 58 | |
| DeepSeek V4 Pro 0813 (max) | 53 | |
| Qwen3.8 27B | 52 | |
| 4550556065 | ||
Three points, top to top. One point between the best open-weight entry and GPT-5.6 Sol. Alibaba's largest Qwen3.8 configuration sits at 58 and DeepSeek V4 Pro 0813 at 53 — a spread among open-weight models that is now wider than the distance from the best of them to the proprietary leader.
Open weights is the right term for that column, and open source mostly is not. What these licences release is a file of weights that can be downloaded, self-hosted and fine-tuned; the training data, the pipeline that produced it and most of the decisions in between stay inside the labs. Qwen3.8 27B is Apache 2.0, which is unusually permissive, and still says nothing about how it was made.
The second chart is where the more interesting result is. Artificial Analysis plots the same index against the parameters a model actually activates at inference, and Qwen3.8 27B lands on its Pareto line: 52 on the index, from a 27-billion-parameter model released four days ago. Sitting just to the left, at a similar score with fewer active parameters, is DeepSeek's frontier entry, which scores 53 on the table. Two labs are now competing on intelligence per active parameter, not only on the absolute number.

That axis is not a price. Active parameters are one term in a cost model that also includes memory bandwidth, batching, quantisation, the hardware underneath and whatever margin a provider adds, and Artificial Analysis publishes cost per task as its own separate measurement for exactly that reason. A model that activates fewer parameters is more parameter-efficient on this chart; whether it is cheaper to serve is a different question with a different answer for every deployment.
None of this makes the index a verdict. It is an aggregate of nine benchmarks, reasoning effort is part of each entry rather than a footnote to it, and a three-point difference on a composite score is not a three-point difference in whatever a particular team is trying to do. Scores also move: Qwen3.8 27B was evaluated within days of release, and the table is different most weeks.
What the numbers do change is the shape of a decision. When the best downloadable model was an entire tier below the best API, capability decided it. At three points, a buyer is trading a small amount of measured intelligence against self-hosting, fine-tuning, data that never leaves, and not being one pricing page away from a bill they do not control. Closed models still lead. There is less room left underneath them than there was.
Sources
This article was written from these pages. Read them.
- primaryArtificial Analysis — model comparison and Intelligence Indexartificialanalysis.ai
- primaryArtificial Analysis — Qwen3.8 27B, intelligence against active parametersartificialanalysis.ai
Written and edited by a person at Epoch, from the primary documents listed above, and checked before publication. Any diagram here is ours.