MODEL INDEX

Which model is
best for your work?

The right model depends on the job. Start with writing, coding, agents, research, or creative work — then see the evidence behind the shortlist.

7public leaderboard sources
0–100consensus score scale
30models in the full ranking

THE CONSENSUS INDEX

Top models, in context.

Last checked Aug 21, 2026 at 20:00 UTC

Scores reflect relative performance across public model evaluations. Coverage and evidence confidence are shown alongside the score so a ranking never has to tell the whole story by itself.

AI Center Model Consensus leaderboard
RankModel
012026-06-09
88.0%
¥67.21¥336.0389.4 Strong evidence
02
Claude Opus 5Anthropic
2026-07-24
84.5%
¥33.60¥168.0286.2 Strong evidence
032026-07-09
84.5%
¥33.60¥201.6283.3 Strong evidence
04
Kimi K3Moonshot AI
2026-07-16
84.5%
¥20.00¥100.0079.9 Strong evidence
052026-08-18
60.0%
¥9.41¥29.5776.9 Moderate evidence
062026-08-13
81.0%
¥5.04¥25.2075.7 Strong evidence
072026-08-12
84.5%
¥13.44¥40.3274.8 Strong evidence
08
GPT-5.5OpenAI
2026-04-23
80.0%
¥33.60¥201.6273.9 Moderate evidence
092026-08-05
73.0%
¥8.40¥28.5672.5 Moderate evidence
102026-07-19
81.0%
¥12.00¥36.0071.7 Strong evidence
112026-05-28
88.0%
¥33.60¥168.0270.6 Strong evidence
122026-07-09
81.0%
¥13.44¥80.6569.3 Moderate evidence
132026-04-16
72.0%
¥33.60¥168.0264.1 Moderate evidence
142026-07-09
76.5%
¥8.40¥28.5661.2 Moderate evidence
152026-08-13
73.0%
¥9.00¥27.0058.7 Moderate evidence
16
GPT-5.4OpenAI
2026-03-05
69.5%
¥16.80¥100.8157.1 Moderate evidence
172026-07-08
88.0%
¥13.44¥40.3256.4 Strong evidence
182026-06-30
84.5%
¥13.44¥67.2155.7 Strong evidence
192026-07-09
81.0%
¥1.34¥8.0651.9 Moderate evidence
202026-07-21
81.0%
¥5.04¥25.2049.9 Moderate evidence
212026-08-14
51.0%
¥3.00¥12.0047.9 Limited evidence
222026-05-19
81.0%
¥10.08¥60.4947.6 Strong evidence
232026-06-16
84.5%
¥8.00¥28.0046.8 Strong evidence
242026-02-05
73.0%
¥33.60¥168.0246.1 Moderate evidence
252026-07-31
76.5%
¥3.00¥9.0042.9 Moderate evidence
262026-02-19
88.0%
¥13.44¥80.6541.1 Strong evidence
272026-02-17
76.5%
¥20.16¥100.8137.4 Moderate evidence
282026-05-19
51.0%
¥12.00¥36.0034.3 Limited evidence
292026-02-05
42.0%
¥11.76¥94.0932.2 Limited evidence
302026-04-24
73.0%
¥2.92¥5.8531.8 Moderate evidence

Benchmark results link directly to the original public evaluation boards. Provider pricing is checked against official websites and converted to CNY where necessary.

START WITH THE WORK

Which model is best
for your work?

There is no universal winner. Pick the job first, then compare the models that make sense for that job.

Editorial starting points based on the current AI Center snapshot — not a replacement for the full evidence table.

READ THE SIGNAL

Three numbers worth understanding.

METHOD v8.6
02

Evaluation coverage

Coverage shows how much evidence a model has actually earned. Missing evaluations are not treated as zeroes, and missing source weight is never redistributed.

More coverage = more confidence in the placement
03

Evidence confidence

High, medium, and low confidence describe how complete and stable the evidence is. They do not describe the model’s quality.

High Medium Low

A TRANSPARENT METHOD

Rankings are useful.
Context makes them honest.

AI Center’s model index is designed for decisions, not hot takes. A model can lead with less evidence, and a narrow score gap does not automatically mean a meaningful capability gap.

Read the methodology
01Normalize model identities
02Compare only shared evaluations
03Estimate latent capability
04Anchor the score and show uncertainty