Every step, so you can check it or disagree with it.
1. Capture
Every board is fetched daily at 04:30 UTC from its own page, data file or API and parsed without a model. If a parser breaks, the model reads that board’s page instead and those rows carry an Unverified badge.
2. One model, one entry
Variants collapse to the model: effort levels (high, xhigh, max), thinking modes, dated snapshots and agent harnesses. The best variant on a board stands for the model there. The exact name each board published is kept next to every entry.
3. Which boards count
Language-model quality boards (chat, reasoning, coding, agents and tools, vision) that published in the last 180 days. Image, video, speech, embeddings, speed, price and usage get their own category leaders.
4. Points
On each counted board: #1 earns 10 points, #2 earns 9, down to 1 point for #10. Podium score = points earned / (10 × boards counted) × 100. Ties break on #1 finishes, then the median rank across boards that list the model.
A board is one team’s method on one day. Arena does not run evaluations and does not adjust anyone’s numbers; it shows where the boards agree. Vendors: OpenAI, Anthropic, Google and others appear by the name each board uses, linked to their own pages.
Scores as published by each board on the capture date. Model names and logos belong to their owners; logos via logo.dev.