How we score things (and why "contact sales" costs you points)

The rubric behind every ranking on this site, published so you can argue with it. Ops burden 40%, real cost 30%, fit 20%, tooling 10%.

The bench2 min readTooling

Ranked bar chart of the standard Canyon Review scoring weights, operational burden first

Every ranking here ends in numbers, so it is fair to ask where the numbers come from. This page is the answer, pinned for future arguments. When a specific review weighs things differently, it says so up top; this is the default rubric.

Operational burden, 40%. The biggest weight because it is the biggest cost, and the one pricing pages never show. Who gets paged, how often, and what do they need to know? A tool that runs itself scores high. A tool that runs itself until it doesn't, and then requires archaeology, scores like the archaeology. We weight burden at the two-year mark, not the demo.

Cost at a stated workload, 30%. Not list price. We define a modest, concrete workload (a million messages a day, three repos of CI, 200GB of logs) and price it, with an as-of date, from the public page. If pricing is "contact sales", we score the tier at the worst quote we can document, because that is negotiating advice as much as scoring: the number is designed to be invisible, and invisible numbers round up.

Fit for the actual job, 20%. Not capability, fit. Kafka is enormously capable; as a background-job queue for a five-person team it fits like a forklift in a kitchen. Points here go to tools whose defaults match the job we ranked them for, and we say what that job is in the first paragraph.

Tooling and docs, 10%. The smallest weight, and honestly it functions as a tiebreaker. Good docs cannot rescue a tool that pages you nightly; the reverse trade is one everyone accepts.

Two standing modifiers. The boring bonus: a tool with no surprises in two years of community postmortems gets a bump, because "nothing happened" is the best benchmark result there is. And the 3 a.m. test: if the recovery procedure requires understanding the tool's internals, the burden score takes the hit, no matter how elegant those internals are.

What we do not score: vendor roadmaps (futures are free to promise), logos on the customer page, funding announcements, and whether the tool is fashionable this year. Fashion is how you end up with the forklift.

The chart shows the weights; it is illustrative in the most literal sense. Argue with the rubric by mail. The best objections change it, and changes get printed, because a ranking site that quietly moves its own goalposts is a marketing site with extra steps.