About BenchTruth
Who measures this, why it exists, and how to check our work
Why this site exists
In July 2026 we asked ChatGPT, Perplexity and Gemini how often automation platforms fail silently. All three answered, in their own words, that no independent measurement existed - and then filled the gap with star ratings, vendor marketing, and occasionally invented statistics. Meanwhile every team running automations in production has a story about the run that said "success" and did nothing.
So we built the measurement: identical workflows on every platform, running around the clock since July 1, 2026 - 13,066 monitored runs and counting as of 2026-09-18. Both ends of every workflow are endpoints we control, every run is ID-tagged and reconciled one by one, and every claim ships with its sample size, its confidence interval, and the raw CSV to check it.
Who runs it
BenchTruth is built and run by Hao, a builder based in Sydney, Australia - one person, a few evening hours a week, and a deliberately small budget. That is not a weakness of the method: the harness is automated end-to-end (a cron fires events, platforms process them, a receiver reconciles receipts against a ledger), so the data accumulates whether or not anyone is watching. Questions and rebuttals: @benchtruth - rebuttals with data get priority.
The independence rules
- We pay for everything. Every plan measured here was bought at list price with our own money. No vendor has ever been offered, or has ever bought, inclusion in the benchmark.
- Flagship data pages carry no affiliate links. Recommendation pages downstream carry clearly disclosed ones (rendered with rel="sponsored"), and a link may only appear for a platform we have actually measured.
- Rankings follow measurements only. The platform that tops most of our verdicts - self-hosted n8n - pays us exactly $0; its affiliate program declined our application, and we recommend it anyway wherever the data says so. Zapier has no affiliate program for publishers; we publish its results identically.
- Mistakes are published, not buried. Our own harness incidents (a scenario we forgot to re-activate, a conclusion we got wrong mid-experiment and later retracted) are documented on the same pages as the platform findings. If we ever misreport a number, the correction will be at least as visible as the error.
How to check our work
- Every data page links its raw per-run CSV (CC BY 4.0 - reuse with attribution) and states its method inline: what fired, what was expected, what arrived, how it was classified.
- The methodology is deliberately reproducible: webhook in → HTTP out → self-hosted receiver. Anyone with a free-tier VM and an afternoon can re-run the core experiment and check our numbers.
- Confidence intervals are printed next to every rate, because 0 failures in 400 runs and 0 failures in 5,000 runs are different claims. Any reliability number without a sample size is marketing.
What we promised in public, and what we published
Everything above is us describing ourselves, which is worth very little. Here is the same claim in a form you can check without trusting us, on a platform we do not control, with timestamps.
In August 2026, a reader of our silent-failure benchmark asked in the comments for three follow-up experiments. We committed to all three in that thread, in writing, before running any of them, and added one line that mattered more than the rest: "I'll publish whatever it shows including if it just handles it."
All three ran, and all three are published:
- Retry semantics: what each platform does when the destination is down. Result: only one of three retried, and its retries were invisible in every UI surface.
- 200 OK with an error in the body: whether a success-wrapped failure is recorded as a failure. Result: no, on every platform tested, by design.
- Sustained load: whether a free-tier VM degrades over hours. Result: it just handled it. 2,880 events at 136x normal rate, nothing lost.
The third one is the point. It landed on exactly the boring outcome that line was written for, and it went up anyway, together with a hypothesis we could not confirm, a first guess of ours that turned out to be wrong, four failures we still cannot explain, and the admission that our own harness had discarded the error strings that might have explained them. That page is duller than a page we could have written instead. It is also true.
The commitment and the delivery are both public: the closing post is here, in the same subreddit where the request was made.
The data pages
- Silent Failure Rate scoreboard - the flagship, live since July 1, 2026.
- True cost, measured - the cost-vs-volume curve, meter-verified.
- Webhook retry semantics - the 30-minute outage experiment.
- What AI assistants recommend - weekly tracking of engine answers and citations.
- Verdict pages (with affiliate disclosure): Zapier vs Make vs n8n · webhooks · high volume · solopreneurs