EPOCH
Benchmarks

The best open-weight model is now three points off the top of the index

Artificial Analysis has Claude Opus 5 leading at 63 and Kimi K3 at 60. The gap that used to be a tier is now a rounding error on a nine-benchmark average.

Artificial Analysis

Artificial Analysis publishes one number per model, averaged from nine evaluations, and its table currently opens with Claude Opus 5 in adaptive reasoning at maximum effort on 63, out of 161 models it has assessed. Claude Fable 5 follows at 62, then Claude Opus 5 at high effort and GPT-5.6 Sol at maximum effort, both on 61. The first model on that list whose weights anyone can download is Kimi K3 at maximum effort, on 60.

The axis is a window on a hundred-point index rather than the whole of it. A dot is read from its position, so narrowing the window claims nothing the numbers do not — three points still measure three points against the ticks. One index is not a verdict: nine evaluations averaged into a single number, across 161 models. Reasoning effort is part of each entry, so these are maximum-effort configurations on both sides of the line, and open weights means downloadable weights rather than an open training stack. Scores read from Artificial Analysis on 18 August 2026, and they move; Qwen3.8 27B was released four days earlier.
ModelIndex
Proprietary
Claude Opus 5 (adaptive reasoning, max)63
Claude Fable 5 (max, Opus 4.8 fallback)62
Claude Opus 5 (adaptive reasoning, high)61
GPT-5.6 Sol (max)61
Open weights
Kimi K3 (max)60
Qwen3.8 2.4T A95B58
DeepSeek V4 Pro 0813 (max)53
Qwen3.8 27B52
4550556065

Three points, top to top. One point between the best open-weight entry and GPT-5.6 Sol. Alibaba's largest Qwen3.8 configuration sits at 58 and DeepSeek V4 Pro 0813 at 53 — a spread among open-weight models that is now wider than the distance from the best of them to the proprietary leader.

Open weights is the right term for that column, and open source mostly is not. What these licences release is a file of weights that can be downloaded, self-hosted and fine-tuned; the training data, the pipeline that produced it and most of the decisions in between stay inside the labs. Qwen3.8 27B is Apache 2.0, which is unusually permissive, and still says nothing about how it was made.

The second chart is where the more interesting result is. Artificial Analysis plots the same index against the parameters a model actually activates at inference, and Qwen3.8 27B lands on its Pareto line: 52 on the index, from a 27-billion-parameter model released four days ago. Sitting just to the left, at a similar score with fewer active parameters, is DeepSeek's frontier entry, which scores 53 on the table. Two labs are now competing on intelligence per active parameter, not only on the absolute number.

Artificial Analysis' scatter chart of Intelligence Index against active parameters at inference time, on a log scale, for open-weight models. A dotted Pareto line runs from small models at the lower left up through Qwen3.8 27B at about 52 and DeepSeek V4 Pro 0813 nearby with fewer active parameters, then GLM-5.2 and Kimi K3 at around 60 near 100 billion active parameters. Many other open-weight models sit below the line.
The same index against the parameters a model activates to answer. Qwen3.8 27B lands on the frontier line; DeepSeek's entry is left of it at a similar score.Artificial Analysis

That axis is not a price. Active parameters are one term in a cost model that also includes memory bandwidth, batching, quantisation, the hardware underneath and whatever margin a provider adds, and Artificial Analysis publishes cost per task as its own separate measurement for exactly that reason. A model that activates fewer parameters is more parameter-efficient on this chart; whether it is cheaper to serve is a different question with a different answer for every deployment.

None of this makes the index a verdict. It is an aggregate of nine benchmarks, reasoning effort is part of each entry rather than a footnote to it, and a three-point difference on a composite score is not a three-point difference in whatever a particular team is trying to do. Scores also move: Qwen3.8 27B was evaluated within days of release, and the table is different most weeks.

What the numbers do change is the shape of a decision. When the best downloadable model was an entire tier below the best API, capability decided it. At three points, a buyer is trading a small amount of measured intelligence against self-hosting, fine-tuning, data that never leaves, and not being one pricing page away from a bill they do not control. Closed models still lead. There is less room left underneath them than there was.

Sources

This article was written from these pages. Read them.

  1. primaryArtificial Analysis — model comparison and Intelligence Indexartificialanalysis.ai
  2. primaryArtificial Analysis — Qwen3.8 27B, intelligence against active parametersartificialanalysis.ai

Written and edited by a person at Epoch, from the primary documents listed above, and checked before publication. Any diagram here is ours.