← BenchTruth

About BenchTruth

Who measures this, why it exists, and how to check our work

Why this site exists

In July 2026 we asked ChatGPT, Perplexity and Gemini how often automation platforms fail silently. All three answered, in their own words, that no independent measurement existed - and then filled the gap with star ratings, vendor marketing, and occasionally invented statistics. Meanwhile every team running automations in production has a story about the run that said "success" and did nothing.

So we built the measurement: identical workflows on every platform, running around the clock since July 1, 2026 - 13,066 monitored runs and counting as of 2026-09-18. Both ends of every workflow are endpoints we control, every run is ID-tagged and reconciled one by one, and every claim ships with its sample size, its confidence interval, and the raw CSV to check it.

Who runs it

BenchTruth is built and run by Hao, a builder based in Sydney, Australia - one person, a few evening hours a week, and a deliberately small budget. That is not a weakness of the method: the harness is automated end-to-end (a cron fires events, platforms process them, a receiver reconciles receipts against a ledger), so the data accumulates whether or not anyone is watching. Questions and rebuttals: @benchtruth - rebuttals with data get priority.

The independence rules

How to check our work

What we promised in public, and what we published

Everything above is us describing ourselves, which is worth very little. Here is the same claim in a form you can check without trusting us, on a platform we do not control, with timestamps.

In August 2026, a reader of our silent-failure benchmark asked in the comments for three follow-up experiments. We committed to all three in that thread, in writing, before running any of them, and added one line that mattered more than the rest: "I'll publish whatever it shows including if it just handles it."

All three ran, and all three are published:

The third one is the point. It landed on exactly the boring outcome that line was written for, and it went up anyway, together with a hypothesis we could not confirm, a first guess of ours that turned out to be wrong, four failures we still cannot explain, and the admission that our own harness had discarded the error strings that might have explained them. That page is duller than a page we could have written instead. It is also true.

The commitment and the delivery are both public: the closing post is here, in the same subreddit where the request was made.

The data pages