16comparisons on this site publish scores out of 10. A score is worth nothing if you can't see where it came from, so this page explains exactly that — including what we don't do.
Every comparison scores both tools on criteria chosen for that category — ease of use and automation depth for project management, model choice and agent quality for AI coding tools. The criteria differ; the scale doesn't.
Four things happen before a comparison is published.
Accounts are set up the way a reader would set them up, on the plan a reader would actually start on — usually the free tier or the entry paid tier, not a vendor-provided demo account.
Within a comparison, both tools get the same jobs: import the same data, build the same board, send the same campaign. Differences in the result are the comparison.
Including the per-seat minimums, the annual-versus-monthly gap, the feature gates and the overage rates. Most of what makes software expensive is below the pricing table.
Claims about what a tool can do are verified against its documentation, and scores are revisited when a changelog says something material shipped.
This half matters more than the half above. Any site can claim it tests things; the useful question is what it refuses to do.
There is no test rig and no synthetic scoring. These are editorial judgements from using the products, and we'd rather say that than dress them up as measurements.
A comparison covers what most readers decide on. If your requirement is unusual, the scores are a starting point, not an answer.
No vendor has ever paid to appear, to rank higher, or to have a score changed. There is no mechanism for it.
Some links earn a commission if you buy. The ranking is written before the links go on, and tools with no affiliate programme are recommended over ones that pay whenever they're the better answer.
Three failure modes worth knowing about before you rely on a page here.
Vendors change pricing without notice. Every page carries the date it was last checked — treat any figure older than that as needing confirmation on the vendor's own site.
An 8 for “ease of use” means eight relative to the other tool on that page, not against all software everywhere. The same tool can score differently in two comparisons.
AI tooling in particular can invalidate a verdict within a quarter. Those pages are re-checked more often, and the date at the top tells you when.
If a score is wrong — a feature shipped, a price changed, or we simply misread how something works — say so and we'll check it. Corrections are made on the page itself and the updated date is moved, so you can tell the difference between a page that was re-checked and one that has been sitting untouched.
Vendors are welcome to write in on the same terms as readers. Being the vendor doesn't get a score changed; being right does.
16 head-to-head and three-way breakdowns, each scored on this basis and dated with the last time it was checked.