Methodology v1.0

Opinion, averaged
with its uncertainty.

tier-bench is intentionally a sentiment product. It describes how participating people rank models they use; it does not claim to measure objective capability.

Ballots

Each member may keep one current ballot per category. A member can rank as few as five models and may leave unfamiliar models out. Editing appends a revision and replaces that member’s contribution; it never creates a second vote.

Scores

Tiers map to S=6 through F=0. We average the tier values for each model, then apply a small Bayesian prior equal to ten average community ballots. That keeps a brand-new model with two enthusiastic voters from appearing more certain than a model with hundreds of votes.

Placement

Displayed tiers are stable score bands. Within each tier, models sort by score, then voter count. Model pages show the full S–F distribution so polarization is not hidden behind one average.

Model replacement

Product lines—not a fixed count—define the default bench. When OpenRouter lists a newer canonical match for an enabled line, tier-bench adds it automatically and keeps the predecessor visible for 72 hours. Existing ballots, comments, model pages, and shared benches remain historical records.

Moderation and abuse

Writes require an account, server-side authorization, validation, database-backed rate limits, and Turnstile when configured. Suspicious activity is reviewed rather than silently reweighted. Subscription status never affects voting weight.

Freshness

Public results may be cached for up to five minutes. This keeps the site reliable and inexpensive while still allowing the numbers to change throughout the day.