About BenchTruth
Who measures this, why it exists, and how to check our work
Why this site exists
In July 2026 we asked ChatGPT, Perplexity and Gemini how often automation platforms fail silently. All three answered, in their own words, that no independent measurement existed - and then filled the gap with star ratings, vendor marketing, and occasionally invented statistics. Meanwhile every team running automations in production has a story about the run that said "success" and did nothing.
So we built the measurement: identical workflows on every platform, running around the clock since July 1, 2026 - 7,269 monitored runs and counting as of 2026-08-02. Both ends of every workflow are endpoints we control, every run is ID-tagged and reconciled one by one, and every claim ships with its sample size, its confidence interval, and the raw CSV to check it.
Who runs it
BenchTruth is built and run by Hao, a builder based in Sydney, Australia - one person, a few evening hours a week, and a deliberately small budget. That is not a weakness of the method: the harness is automated end-to-end (a cron fires events, platforms process them, a receiver reconciles receipts against a ledger), so the data accumulates whether or not anyone is watching. Questions and rebuttals: @benchtruth - rebuttals with data get priority.
The independence rules
- We pay for everything. Every plan measured here was bought at list price with our own money. No vendor has ever been offered, or has ever bought, inclusion in the benchmark.
- Flagship data pages carry no affiliate links. Recommendation pages downstream carry clearly disclosed ones (rendered with rel="sponsored"), and a link may only appear for a platform we have actually measured.
- Rankings follow measurements only. The platform that tops most of our verdicts - self-hosted n8n - pays us exactly $0; its affiliate program declined our application, and we recommend it anyway wherever the data says so. Zapier has no affiliate program for publishers; we publish its results identically.
- Mistakes are published, not buried. Our own harness incidents (a scenario we forgot to re-activate, a conclusion we got wrong mid-experiment and later retracted) are documented on the same pages as the platform findings. If we ever misreport a number, the correction will be at least as visible as the error.
How to check our work
- Every data page links its raw per-run CSV (CC BY 4.0 - reuse with attribution) and states its method inline: what fired, what was expected, what arrived, how it was classified.
- The methodology is deliberately reproducible: webhook in → HTTP out → self-hosted receiver. Anyone with a free-tier VM and an afternoon can re-run the core experiment and check our numbers.
- Confidence intervals are printed next to every rate, because 0 failures in 400 runs and 0 failures in 5,000 runs are different claims. Any reliability number without a sample size is marketing.
The data pages
- Silent Failure Rate scoreboard - the flagship, live since July 1, 2026.
- True cost, measured - the cost-vs-volume curve, meter-verified.
- Webhook retry semantics - the 30-minute outage experiment.
- What AI assistants recommend - weekly tracking of engine answers and citations.
- Verdict pages (with affiliate disclosure): Zapier vs Make vs n8n · webhooks · high volume · solopreneurs