PeopleBench Rate your models

Open data

The ratings behind the boards, one row per rating, so you can check our numbers or compute your own. Updated once a week.

Files

No real ratings yet. The files appear here after the first week with real submissions. (Test data never goes into them.)

Licence

CC BY 4.0. Use it for anything, including commercially, as long as you credit “PeopleBench” and link to this site.

Columns

submissionA random id that groups the ratings one person gave in one go. It is new in every export, and the rows come in random order, so rows can't be linked across exports.
monthCalendar month of the submission, e.g. 2026-09. No finer time.
disciplineThe board: coding, writing, research, everyday, images, video.
modelModel id, as in the URLs on this site.
stars1 to 5.
surfaceWhere it was used: app, api, coding_tool (IDE or agent), local (run on your own hardware; open-weight models only), other. Empty when not given.
toolFor IDE/agent use: which tool, e.g. cursor, claude-code, copilot. A tool used by fewer than 5 people on a board in a month becomes other.
planThe exact plan as owner:plan, e.g. claude:max-20x (the model's own app) or cursor:pro (a tool). …:other means "other / not sure"; empty means not given. A plan used by fewer than 5 people on a board in a month is replaced by its plan type (personal, business, api, unknown).
audiencepersonal, business, api, local, or unknown, derived from the plan (or from the surface for API and local use).
limitsDid they hit usage limits: never, sometimes, often, constantly, or empty.
frequencyHow often they used it: few, weekly, daily, or empty.

Only ratings that count are included (weight above 0). The weights themselves are left out on purpose: publishing them would show exactly which checks a submission failed, which would help anyone trying to game the boards. So the file lets you recompute the boards closely, but not to the decimal. Also left out: network hashes, where visitors came from, countries, anything finer than the month, and the made-up attention-check models.