Benchmarks

What each model is actually good at.

Independent test scores for Claude, GPT, Gemini and the open models, side by side. No marketing numbers: every score links to its source.

Published test scores from independent labs. Pick a model to see how it ranks on each skill; tap a score for the source.

Published test scores, side by side. Each column is one skill; the brighter the cell, the higher that model ranks among the others on that test. Hover a score for details, click it for the source. "—" means the test hasn't been run on that model (text-only models, for example, can't take the image test).

Rank the models on one skill

Smarts vs price

Smartest first (Artificial Analysis Intelligence Index), with the blended API price per 1M tokens. A high score with a low price is the best value.

Up = smarter overall (Artificial Analysis Intelligence Index). Left = cheaper (blended API price per 1M tokens, log scale). The top-left corner is the best value.