← BenchTruth

About BenchTruth

Who measures this, why it exists, and how to check our work

Why this site exists

In July 2026 we asked ChatGPT, Perplexity and Gemini how often automation platforms fail silently. All three answered, in their own words, that no independent measurement existed - and then filled the gap with star ratings, vendor marketing, and occasionally invented statistics. Meanwhile every team running automations in production has a story about the run that said "success" and did nothing.

So we built the measurement: identical workflows on every platform, running around the clock since July 1, 2026 - 7,269 monitored runs and counting as of 2026-08-02. Both ends of every workflow are endpoints we control, every run is ID-tagged and reconciled one by one, and every claim ships with its sample size, its confidence interval, and the raw CSV to check it.

Who runs it

BenchTruth is built and run by Hao, a builder based in Sydney, Australia - one person, a few evening hours a week, and a deliberately small budget. That is not a weakness of the method: the harness is automated end-to-end (a cron fires events, platforms process them, a receiver reconciles receipts against a ledger), so the data accumulates whether or not anyone is watching. Questions and rebuttals: @benchtruth - rebuttals with data get priority.

The independence rules

How to check our work

The data pages