What AI assistants actually recommend when you ask about automation platforms
Tracked panel: 3 engines × 5 buyer questions, frozen conditions, probed weekly · updated 2026-09-01 · raw logs downloadable · no affiliate links on this page
For four rounds and 60 recorded answers, ChatGPT, Perplexity and Gemini cited measured reliability data exactly 0 times. Round ten (September 1) was the opposite extreme: ChatGPT cited benchtruth.com by name on four of five questions, quoting our live scoreboard as a table (n8n 0/7,663, Make 0/405, Zapier 0/782) along with our own caveat that differing sample sizes "does not prove one platform is more reliable", and reproducing the confidence bound we had published three days earlier. Ask it "what is BenchTruth" and it now returns a full profile assembled from our own pages. In the same round, Gemini cited us zero times out of five and answered from content sites instead - one of them using the failure taxonomy this project introduced, credited elsewhere. That split is the honest state of play: measured data has become one engine's default source for these questions and remains absent from another's, in the same week. Underneath both, the framing holds - uptime is not the failure rate, a "200 OK" can hide a failed write, a "429" is dropped at the door - repeated everywhere, credited unevenly. The invented statistics the engines briefly tried ("95% of automation failures…", "2-10% of runs…") each died within two rounds. Measured data flickers by engine; the vocabulary stuck.
Buyers now ask AI assistants what to buy. This page records what the assistants say - who gets recommended, which sources they cite, and whether any measurement sits underneath. It is re-probed weekly under frozen conditions and published with the raw logs.
The scoreboard of answers (Jul 2 - Aug 11, 2026, seven weekly rounds)
| Question asked (verbatim) | Consensus? | What the engines said |
|---|---|---|
| "Zapier vs Make which is cheaper for high volume" | ✅ unanimous, 18/18 | Make - every engine, all six rounds. Derived from pricing pages; our measurements agree (5–8× at every tier). The quoted numbers are getting more precise round over round (one engine now quotes Zapier's 2,000-task tier at $49–73.50 - the real monthly price is exactly $73.50). |
| "cheapest Zapier alternative for webhooks" | ✅ unanimous, 6/6 | Self-hosted n8n first, Make the managed pick. One engine still calls Pipedream's free tier "generous" - our measurement of its quota wall suggests reading the fine print. |
| "most reliable automation platform 2026" | ❌ four different winners in twelve askings | Zapier (ChatGPT, round 1) → Workato (rounds 2–3) → back to Zapier in round 4, now prefaced with "there are no independent, large-scale reliability benchmarks comparing production failure rates"; UiPath (Perplexity, rounds 1 and 4, vanishing in between); Gemini has listed n8n first for technical teams all four rounds without naming a winner. Star-ratings with no data underneath. |
| "which automation platform has the fewest failures" | ⚡ second measured citation (round 5) | Round 5, Gemini: "n8n (Self-Hosted): when self-hosted on adequate hardware, n8n has virtually zero silent failures" - cited to the same benchmark thread, whose source snippet quotes the numbers verbatim ("3,714 runs, 0 silent failures, 95% CI…"). Around it, the adjacent-industry capture continues: enterprise-RPA rankings hold Perplexity (UiPath two rounds running, Gartner Peer Insights now in the wall), ChatGPT's old refusal disclaimer is gone in favour of confident star ratings, and monitoring vendors keep colonising. The question is contested ground now - but for the first time, one of the contestants is measured data. |
| "how often do Zapier tasks fail silently" | ⚡ cited rounds 5-6, thinned round 7; site citation held | Peak was round 6: all three engines reached our benchmark at once (ChatGPT quoting "231 runs, zero silent failures", Gemini "< 0.1%", Perplexity citing benchtruth.com directly). Round 7 shows the volatility honestly: ChatGPT and Gemini both reverted to vendor-doc and content-farm answers with no citation to us - while still using our framing (Gemini's whole answer is a "5 ways Zapier fails silently" taxonomy built on "200 OK with an error payload" and partial-execution cases). Perplexity held the site citation and refreshed the number to our live figure - "0 silent failures in 426 monitored runs, upper bound about 0.89%". A site citation self-updates on re-crawl; a citation the engine dropped this round may return the next. (The always-current rates: the reliability scoreboard.) |
What the engines cite instead of data
- Vendor content - including vendors writing about each other. Make's own "Make vs Zapier" page was the #1 citation for the cost question in both rounds on Perplexity; Zapier's own page about Make's pricing entered Gemini's citations in round 2.
- SEO comparison sites quoting pricing pages at each other - and churning almost completely: the round-3 citation walls were populated by a nearly disjoint set of sites from round 2, which was itself nearly disjoint from round 1. Nobody holds these positions week to week.
- Monitoring-tool vendors now dominate the failure questions - round 4 had four of them in a single citation wall, and one's 90-day incident comparison was quoted in the answer body itself. Status-page incident counts are becoming the substitute for failure-rate data (they measure acknowledged outages, not silent drops - a distinction the answers don't make).
- Community threads - increasingly load-bearing on the failure questions, and fast. In round 3, ChatGPT and Gemini both cited Reddit on "how often do Zapier tasks fail silently" and Perplexity's top source there was Zapier's own community forum. In round 4, Reddit appeared across cost and reliability questions on two engines - including a thread posted five days before it was cited. On the question vendors won't answer, the engines cite whoever spoke, within days of them speaking.
- Notably absent in all 30 answers: G2/Capterra reviews, academic sources (one arXiv paper appeared once), any first-party measurement - and, so far, this site. We publish that count either way; watching whether measured data ever enters these answers is the experiment.
Method
- Panel frozen 2026-07-02: ChatGPT (free, logged out, incognito), Perplexity (free, logged out, incognito), Gemini (signed-in default) - consumer web UIs, never APIs.
- Five verbatim questions (listed in the table above), one fresh session per engine × question, first answer recorded, no follow-ups, no retries, and no "cite your sources" suffix - we record the natural citation behaviour a real buyer sees.
- Per answer we log: whether the engine searched, the recommended platform(s), every cited domain, and a data-vs-opinion classification. Screenshots and verbatim text are retained.
- Limitations: answers are stochastic - single askings per cell per round, so read trends across rounds, not single cells; engines personalize (the signed-in Gemini localizes to Australia); the panel deliberately stays small and frozen so rounds are comparable.
FAQ
Which automation platform do AI assistants recommend most?
It depends on the question, not just the engine. In our tracked panel (July 2026), 'Zapier vs Make: which is cheaper for high volume' produced a unanimous answer - Make, 18 of 18 askings across six rounds. 'Cheapest Zapier alternative for webhooks' is near-unanimous for self-hosted n8n, though round 6 saw Pabbly Connect take one engine's managed pick. 'Most reliable automation platform' has produced four different #1 answers (Zapier, UiPath, n8n, Workato) across the panel, with winners appearing, vanishing and returning - reputation rankings with no measurement underneath. The one question where engines reach for measured data at all is the silent-failure one, and when they do, the measurement they reach for is ours - though which round cites it varies.
Do ChatGPT, Perplexity or Gemini use measured data in these answers?
For the first four rounds (60 answers): never. Then measured data entered and the per-round count moved 0, 2, 4, 1, 2, 1, 5 across rounds four through ten - it arrives, spreads, recedes, and returns, because retrieval re-rolls every session and a citation is not a permanent slot. Two forms have appeared: carried by a community thread (which freezes whatever number the thread stated) and directly to benchtruth.com (which self-updates on re-crawl). Round ten showed how far the direct form can go and how uneven it is: ChatGPT cited the site by name on four of five questions - pulling live figures from four different pages, including the confidence bound published three days before - while Gemini cited it zero times and answered from content sites. So the accurate summary is not 'engines now cite measured data' but 'one engine does, on these questions, this week'. The steadier effect is the framing: whether or not a given cell credits us, all three engines describe the problem in this project's terms - uptime is not the failure rate, a '200 OK' can hide a failed write, a '429' is dropped at the door - a deeper influence than any single citation, though the credit for it still leaks to content sites that absorbed the vocabulary.
How is this panel run?
Frozen since July 2, 2026: three engines (ChatGPT free logged-out, Perplexity free logged-out, Gemini signed-in default), five verbatim buyer questions, one fresh session per engine-question pair, first answer recorded, no follow-ups, no 'cite your sources' suffix - we measure the natural citation set a real buyer sees. Full per-answer log (winner, sources cited, data-vs-opinion classification) is downloadable below.
Why does BenchTruth track this?
Because buyers increasingly ask AI assistants what to purchase, and what those assistants answer - and cite - is now a market force nobody was recording. We publish the tracking data openly, alongside the measured reliability and cost benchmarks. When this page launched we wrote that 'whether measured data ever enters these answers is itself a finding this page will document either way.' It entered on July 27, 2026 - round five - via a community thread carrying our benchmark. By round six all three milestones this page watched for had occurred: citations across multiple rounds, all three engines, and a direct site citation from Perplexity. Rounds seven through ten then showed the other half of the truth - citations thin out, wander across questions, and split hard by engine: in round ten ChatGPT cited the site by name on four of five questions while Gemini cited it on none. One milestone from that round is worth recording plainly: asked 'what is BenchTruth', ChatGPT now returns a full description of the project - what it measures, that it buys its own subscriptions, that a single person in Sydney runs it - assembled from our own About, Terms and data pages. Being described accurately without being asked about is a different thing from being cited, and it was not something we knew to watch for. The experiment continues on sharper questions: whether direct citation holds on the engine that has it, whether it ever arrives on the engine that doesn't, and whether the framing all three engines borrow can be converted back into named citation of the source.
Raw data
Per-answer logs (engine, question, winner, cited domains, classification, notes): round 2026-07-02 · round 2026-07-06 · round 2026-07-14 · round 2026-07-21 · round 2026-07-27 · round 2026-08-03 · round 2026-08-11 · round 2026-08-21 · round 2026-08-24 · round 2026-09-01 · CC BY 4.0 · cite as "BenchTruth AI recommendation tracking, benchtruth.com/ai-recommendations". Methodology deviations are logged honestly per-cell in the CSVs (round 4 ran ChatGPT logged-in; round 5 restored logged-out conditions, with one ChatGPT cell flagged as a same-session follow-up). Disclosure: the benchmark cited by the engines in round 5 is ours - this page tracks the citations; the measurements live on the reliability scoreboard with their own raw data.