PeopleBench Rate your models

How the boards are made

People rate at least two models they used for the same kind of work. We turn every pair into a head-to-head result and fit a ranking to those results.

What you're asked

  1. What you used AI for in the last 30 days. You can pick several.
  2. For each of those, which models you used. At least two.
  3. One to five stars for each. Optional: how often you used it, where (the official app, the API, a coding tool) and on which plan.
  4. One short question about a product you picked.

You don't see the boards while you rate. Models are grouped by family (Claude, GPT, Gemini and so on), and you open a family to see its versions. Families are listed by how many people rated them for that kind of work in the last 90 days. Until a board has 30 raters, we use a fixed starting order based on how widely each family is used. Family logos appear only in the form, never on the boards, so the results don't look like an ad for anyone.

Quality and limits are kept apart

The stars are about the quality of the results. They are the only thing the ranking uses. Separately, we ask whether you hit usage limits (never, sometimes, often, or constantly), where you used the model (its own app, the API, an IDE or agent tool such as Cursor or Claude Code, or on your own hardware for open-weight models) and on which exact plan (for example Claude Pro or Max 20x, or Cursor Pro).

The boards show, next to each model, the share who hit the limits often or constantly, once at least 10 people have answered. Model pages break this down by plan, so you can see if a model is great but unusable on a cheaper plan. The boards can be filtered by plan type: personal, business, API, or local. A board gets a "Local" toggle as soon as anyone has rated a model they ran locally.

From stars to a ranking

Some people give everything five stars. Others never go above three. So we don't average the stars. If you gave model A four stars and model B two, that's one win for A over B. Equal stars are a draw. Someone who rates three models adds three pairs.

You can rate 2 to 6 models per kind of work, and each pair from your submission counts 1/(k−1), where k is how many models you rated. That way every model you rated gets the same total weight from you whether you rated two models or six, and nobody gets extra influence by rating everything.

We fit a Bradley–Terry model to all the pairs in each board. It's the model behind Elo ratings and Chatbot Arena. From it we show one number per model, the preference score: how often that model would beat an average model on the same board, according to the people who used both. 50% is average. 64% means it wins about two times out of three against a typical model on that board. Scores from different boards can't be compared with each other, because each board has its own average.

The bar next to each score is a 95% bootstrap range: we redo the fit 200 times on random resamples of the submissions and show where the middle 95% of results land. When two models' ranges overlap, we can't tell which is better, so they share a rank range such as "2–4": the best rank is one more than the number of models that are clearly ahead, the worst is the number of models that aren't clearly behind.

The dots

Each row shows one dot per person. For every submission that counts (weight above 0) and rated at least two models on that board, we compare the model's stars with the average stars that person gave the other models they rated there. Higher is a filled dot (rated it higher), equal is a ring (the same), lower is a grey dot. "79 of 103 people rated it higher" means 79 of the 103 people who rated it next to something else gave it more stars than their average for the rest. The dots count people, not weight. Only submissions that count at least half (weight 0.5 or more) are counted as a person, so a repeat submission or one that failed a check never shows up as a whole extra dot. The same rule applies to the "hit limits" shares and the per-plan tables. The ranking itself uses the full weights described below.

Above 100 people, each dot stands for several (the row says how many), and the three groups keep their proportions.

The handwritten notes on a board are generated from the same numbers: "clear favourite" when the top model's range doesn't overlap the second's, "too close to call" over the longest run of models whose rank ranges overlap, "climbing fast" or "slipping" for a change of 10 points or more, and "new this month" for a model released in the last 30 days.

Next to the preference score you see the average stars and the number of raters. The stars are there for reference; the ranking comes from the head-to-head pairs.

The fit includes a weak prior: every model starts with two imaginary games against an average model, one won and one lost. With lots of ratings this makes no visible difference. With only a handful it pulls the score toward 50% and makes the range wider, which is honest about how little we know.

We only rank models that are connected to each other through people who rated both, directly or via other models. A model whose raters never overlap with the rest of the board is listed as "Not comparable yet" instead of getting a score on a scale it isn't really on.

A model gets a rank once 30 weighted raters have rated it next to something else on that board. Below that, it's listed in grey as "not enough data yet". Between 30 and 60 raters it is marked Preliminary. Models released in the last 30 days are marked New.

Time window

Each board uses the last 90 days. The small line in the trend column shows the preference score at the end of each of the last 12 weeks, each over its own 90-day window. The arrow compares with the 90 days before the current window, in percentage points, and only appears when the model had enough raters in both. That's how a model that quietly got worse shows up. The boards are recalculated every hour.

Weighting

Every submission gets a weight between 0 and 1. Nobody is told their weight, and no submission is rejected, because telling people what went wrong teaches cheaters what to fix. Things that lower the weight:

Made-up models. The model lists include a made-up model, and the form says so. It works as an attention check: someone who has really used the models they pick won't pick it, while bots and people clicking at random sometimes do. A submission that includes it is kept but gets weight 0, and nothing tells the person. We change the made-up models from time to time and don't publish which they are.

People who used a model daily count slightly more for that model's pairs than people who tried it a few times.

One per week, without accounts

We don't store IP addresses. We store a keyed hash of the address, with a random key made fresh each week. When the week ends we delete the key and the hashes, so they can't be matched or reversed afterwards. Several people behind the same office or mobile network can all rate. After the first, their ratings just count for less that week. More on the privacy page.

Known biases