How Bench's AI tool ratings works.
A deterministic, multi-source rating system for AI developer tools — built by WWT to keep our engineering and client-advisory practices grounded in evidence.
Scoring methodology
Every tool is scored by two pillars. The community signal engine continuously turns raw public signals — reviews, issue trackers, forums — into per-dimension numbers. The ACES editorial evaluation is a structured, evidence-graded expert assessment against a versioned rubric, published only after human review. The radar blends both.
Pillar 1 — Community signal engine
Three layers: dimensions (the radar axes), per-source normalization (how raw signals become 0–100 numbers), and the weighted blend (how those numbers combine into the composite headline).
The eight dimensions
Per-source weights & confidence
Each source contributes to specific dimensions via a versioned weight matrix, editable from the admin console without code changes. Review-site sources cover most dimensions equally except Community; GitHub issues drive Reliability; community forums drive the Community dimension; the ACES editorial pillar is weighted by its evidence grade. A dimension renders as Insufficient data when confidence is below 0.30 — the dash is the honest answer until the pipeline accumulates enough signal.
Pillar 2 — ACES editorial evaluations
ACES scores six dimensions on a 1–20 scale against a comparative rubric (the framework WWT's predecessor radar developed as ACES v2). Evaluations are drafted by an automated researcher with live web evidence, pass fail-closed validation gates, and are published only after a human editor reviews them in the admin console. The rubric itself is versioned data in the repository — any change to scoring doctrine is a reviewed pull request, never a silent tweak.
The six ACES dimensions
The comparative five-band scale
All six dimensions share one scale, anchored to the cohort rather than to abstract quality — a score says where a tool stands relative to its peers today. The top two bands are evidence-gated so enthusiasm can't outrun proof.
Rating, tiers, and the three signals
There is deliberately no single blended “adjusted score.” Each evaluation carries three independent signals:
- Rating (0–100) — pure capability: the six dimensions averaged and scaled (× 5). Tiers: Leading ≥ 76, Proven ≥ 64, Emerging ≥ 50, Watch below.
- Signal level — confidence in the assessment, from the kind of evidence behind it: Validated, Assessed, Tracked, or Detected.
- Evidence grade (A–D) — sourcing rigor, scored from recency (35%), depth (35%), and hands-on testing (30%).
Score caps
Fourteen structural caps limit dimension scores while a disqualifying condition holds — the most restrictive cap wins, and it lifts when the condition resolves. A few examples:
Safeguards
- Desk-research ceiling — without hands-on or independent third-party evidence, no dimension exceeds 16.
- Conflict-of-interest exclusion — tools published by Anthropic are excluded from the autonomous evaluation queue, because the researcher runs on Anthropic models. They can only be evaluated by a human.
- Anti-clustering QA— each batch's score distribution is checked statistically; a batch that clusters suspiciously in the middle band is rejected rather than published.
- Human in the loop — every evaluation, and every change to a published one, goes through editor review with a full audit history.
How the pillars combine
A published ACES evaluation feeds the community engine as one weighted source: each 1–20 dimension score maps onto the radar's axes (× 5 to the 0–100 scale) and blends with the community signals. Community and Reliability come from community sources alone — the editorial pillar doesn't vote there.