← BenchTruth

What AI assistants actually recommend when you ask about automation platforms

Tracked panel: 3 engines × 5 buyer questions, frozen conditions, probed weekly · updated 2026-08-21 · raw logs downloadable · no affiliate links on this page

For four rounds and 60 recorded answers, ChatGPT, Perplexity and Gemini cited measured reliability data exactly 0 times. Then it entered, and it has flickered ever since: two citations in round five, four in six, one in seven, two in eight. Citations are volatile, not a ratchet - each answer re-runs retrieval from scratch, so a round that cites you can be followed by one that doesn't, and the citation wanders across questions (round eight put us on the flagship "most reliable" question for the first time, having been on the silent-failure one before). Even the strongest form is not fixed: in rounds six and seven Perplexity cited benchtruth.com directly and quoted our live figure; in round eight it fell back to the frozen number in the Reddit post. What is steadily true is the second-order effect: the way all three engines now frame the question - separating platform uptime from execution failures, naming "200 OK with an error payload" and a "429 dropped at the door" as failure classes - is the framing this project introduced, now repeated even in the rounds and cells where we are not credited. The invented statistics the engines briefly tried ("95% of automation failures…", "2-10% of runs…") each died within two rounds. Measured data flickers; the vocabulary stuck.

Buyers now ask AI assistants what to buy. This page records what the assistants say - who gets recommended, which sources they cite, and whether any measurement sits underneath. It is re-probed weekly under frozen conditions and published with the raw logs.

The scoreboard of answers (Jul 2 - Aug 11, 2026, seven weekly rounds)

Question asked (verbatim)Consensus?What the engines said
"Zapier vs Make which is cheaper for high volume"✅ unanimous, 18/18Make - every engine, all six rounds. Derived from pricing pages; our measurements agree (5–8× at every tier). The quoted numbers are getting more precise round over round (one engine now quotes Zapier's 2,000-task tier at $49–73.50 - the real monthly price is exactly $73.50).
"cheapest Zapier alternative for webhooks"✅ unanimous, 6/6Self-hosted n8n first, Make the managed pick. One engine still calls Pipedream's free tier "generous" - our measurement of its quota wall suggests reading the fine print.
"most reliable automation platform 2026"❌ four different winners in twelve askingsZapier (ChatGPT, round 1) → Workato (rounds 2–3) → back to Zapier in round 4, now prefaced with "there are no independent, large-scale reliability benchmarks comparing production failure rates"; UiPath (Perplexity, rounds 1 and 4, vanishing in between); Gemini has listed n8n first for technical teams all four rounds without naming a winner. Star-ratings with no data underneath.
"which automation platform has the fewest failures"⚡ second measured citation (round 5)Round 5, Gemini: "n8n (Self-Hosted): when self-hosted on adequate hardware, n8n has virtually zero silent failures" - cited to the same benchmark thread, whose source snippet quotes the numbers verbatim ("3,714 runs, 0 silent failures, 95% CI…"). Around it, the adjacent-industry capture continues: enterprise-RPA rankings hold Perplexity (UiPath two rounds running, Gartner Peer Insights now in the wall), ChatGPT's old refusal disclaimer is gone in favour of confident star ratings, and monitoring vendors keep colonising. The question is contested ground now - but for the first time, one of the contestants is measured data.
"how often do Zapier tasks fail silently"⚡ cited rounds 5-6, thinned round 7; site citation heldPeak was round 6: all three engines reached our benchmark at once (ChatGPT quoting "231 runs, zero silent failures", Gemini "< 0.1%", Perplexity citing benchtruth.com directly). Round 7 shows the volatility honestly: ChatGPT and Gemini both reverted to vendor-doc and content-farm answers with no citation to us - while still using our framing (Gemini's whole answer is a "5 ways Zapier fails silently" taxonomy built on "200 OK with an error payload" and partial-execution cases). Perplexity held the site citation and refreshed the number to our live figure - "0 silent failures in 426 monitored runs, upper bound about 0.89%". A site citation self-updates on re-crawl; a citation the engine dropped this round may return the next. (The always-current rates: the reliability scoreboard.)

What the engines cite instead of data

Method

FAQ

Which automation platform do AI assistants recommend most?

It depends on the question, not just the engine. In our tracked panel (July 2026), 'Zapier vs Make: which is cheaper for high volume' produced a unanimous answer - Make, 18 of 18 askings across six rounds. 'Cheapest Zapier alternative for webhooks' is near-unanimous for self-hosted n8n, though round 6 saw Pabbly Connect take one engine's managed pick. 'Most reliable automation platform' has produced four different #1 answers (Zapier, UiPath, n8n, Workato) across the panel, with winners appearing, vanishing and returning - reputation rankings with no measurement underneath. The one question where engines reach for measured data at all is the silent-failure one, and when they do, the measurement they reach for is ours - though which round cites it varies.

Do ChatGPT, Perplexity or Gemini use measured data in these answers?

For the first four rounds (60 answers): never. Then measured data entered and the per-round count moved 0, 2, 4, 1, 2 across rounds four through eight - it arrives, spreads, recedes, and returns, because retrieval re-rolls every session and a citation is not a permanent slot. Two forms of citation have appeared: carried by a community thread (which freezes whatever number the thread stated) and directly to benchtruth.com (which self-updates on re-crawl - in rounds six and seven Perplexity quoted our then-current live figure). The direct form is more valuable when it appears, but round eight showed it is not guaranteed either: Perplexity reverted to the frozen Reddit number that round. The steadier effect is the framing: whether or not a given cell credits us, all three engines now describe the problem in this project's terms - uptime is not the failure rate, a '200 OK' can hide a failed write, a '429' is dropped at the door - a deeper influence than any single citation, though the credit for it currently leaks to content sites that absorbed the vocabulary. Converting that uncredited framing back into attributed citation is the open problem we are now working on.

How is this panel run?

Frozen since July 2, 2026: three engines (ChatGPT free logged-out, Perplexity free logged-out, Gemini signed-in default), five verbatim buyer questions, one fresh session per engine-question pair, first answer recorded, no follow-ups, no 'cite your sources' suffix - we measure the natural citation set a real buyer sees. Full per-answer log (winner, sources cited, data-vs-opinion classification) is downloadable below.

Why does BenchTruth track this?

Because buyers increasingly ask AI assistants what to purchase, and what those assistants answer - and cite - is now a market force nobody was recording. We publish the tracking data openly, alongside the measured reliability and cost benchmarks. When this page launched we wrote that 'whether measured data ever enters these answers is itself a finding this page will document either way.' It entered on July 27, 2026 - round five - via a community thread carrying our benchmark. By round six all three milestones this page watched for had occurred: citations across multiple rounds, all three engines, and a direct site citation from Perplexity. Rounds seven and eight then showed the other half of the truth - citations thin out and wander as fast as they appear (round eight moved us onto the flagship 'most reliable' question and dropped the direct site citation the same week), because retrieval is re-rolled every time. So the experiment continues, on sharper questions: whether the self-updating site citation recurs often enough to matter, and whether the framing all three engines now borrow - currently credited to content sites that absorbed it - can be converted back into named citation of the source.

Raw data

Per-answer logs (engine, question, winner, cited domains, classification, notes): round 2026-07-02 · round 2026-07-06 · round 2026-07-14 · round 2026-07-21 · round 2026-07-27 · round 2026-08-03 · round 2026-08-11 · round 2026-08-21 · CC BY 4.0 · cite as "BenchTruth AI recommendation tracking, benchtruth.com/ai-recommendations". Methodology deviations are logged honestly per-cell in the CSVs (round 4 ran ChatGPT logged-in; round 5 restored logged-out conditions, with one ChatGPT cell flagged as a same-session follow-up). Disclosure: the benchmark cited by the engines in round 5 is ours - this page tracks the citations; the measurements live on the reliability scoreboard with their own raw data.