If your team built a vendor longlist with an answer engine this quarter, that longlist is an unsourced document until somebody opens the links. Two studies published today, using different methods on different questions, land on the same failure: the layer that grounds AI buying recommendations is not a source of record, and nothing in the output tells you which citations are load-bearing and which are furniture.
The practical change is small and it is procedural. The provenance check has to run before the shortlist, not after it. Right now most teams run it never.
Perplexity's Third-Most-Cited Domain Sells Demo Software
The third-most-cited domain behind Perplexity's software recommendations is a vendor's own marketing blog, and it does not sell the software it is being cited about. Trellner Research put 760 calls through perplexity/sonar and perplexity/sonar-pro across 380 buyer-intent software categories on 2 September 2026, kept 7,534 citations spanning 2,055 distinct domains, and ranked the domains by how often they appeared.
G2 came first at 291 citations. Reddit second at 261. Third, at 194 citations — 2.6% of the total — was guideflow.com — a demo automation platform that sells interactive product demos to sales and pre-sales teams. Trellner is unambiguous about the mismatch: it "is not a review site, a directory or a publisher, and it competes in none of the categories we asked about." Gartner came fourth at 158.
Guideflow appeared as a source in 96 of the 380 categories. Its blog carries titles like "7 best project profitability software for 2026," "8 best employee training software for 2026" and "6 best checklist management software for 2026" — 3,351 blog URLs, per Trellner's read of the sitemap, across categories the company does not compete in. This is not a fraud. It is content marketing, and it is extremely ordinary. It has simply become a primary input to enterprise vendor selection without anyone deciding that it should be.
The distribution underneath is the part that generalises. Tranco is a research ranking of the top million domains, averaged over 30 days across five providers and deliberately hardened against manipulation. Measured against it, 59.8% of Perplexity's citations pointed to domains ranked worse than #100,000, and 23.4% pointed to domains outside the top million entirely. The median Tranco rank of the ranked ones was 71,611.
A domain rank is not a quality score. Plenty of excellent, specific, low-traffic sites sit below 100,000, and a buyer's guide sourced only from the top thousand domains would be a buyer's guide sourced from Reddit and LinkedIn. But when six citations in ten come from below that line and the answer presents all of them in one undifferentiated list, the ranking is doing no work at all — and you have no cheap way to tell a specialist blog from a generated one.
Three Sites, 215,128 Guides, Nine Bylines for One Question
Three apparently coordinated sites have published 215,128 machine-generated software buying guides since December 2023, and Perplexity cites all three. Trellner counted gitnux.org, wifitalents.com and worldmetrics.org at between 103,578 and 107,083 sitemap URLs each, of which those 215,128 are /best/<something>-software/ pages. All three were registered through NameCheap between December 2023 and May 2024. They share a page template, and their DNS is delegated to the same Cloudflare nameservers.
The tells are worth naming, because they are the ones a human reviewer can actually spot in thirty seconds:
- They vouch for each other. WifiTalents publishes a page titled "Is Gitnux.org a Reliable Source?" concluding that Gitnux's statistical reports have been "referenced in over 3,000 high-quality magazines, newspapers, and blogs globally." WorldMetrics publishes "Is WifiTalents.com a Reliable Source?", calling it "one of the most reliable, accurate, and genuinely trustworthy data platforms on the internet" and instructing the reader to "cite WifiTalents with complete confidence."
- The bylines are staffing fiction. For the same question, Trellner found nine distinct named authors across the three sites — Kathryn Blake, Alexander Schmidt and Victoria Marsh on WorldMetrics; Diana Reeves, Helena Kowalczyk and Olivia Thornton on Gitnux; Ryan Gallagher, Isabella Rossi and Natasha Ivanova on WifiTalents. Each page also shipped an unrendered template variable in the byline line, reading "Within the next 26 days" on two of them and "Within the next 40 days" on the third.
- The audience is stated in the title. Fetched on 2 September 2026, Gitnux and WorldMetrics both returned an HTML title of the form
<Brand> — Facts & Grounding Page. Trellner's reading is blunt: "These pages are addressed, in their titles and descriptions, to the software that reads them." The Gitnux homepage claims 6,000+ verified statistics, 1,000+ reports and 72,001 best-lists across 102 categories, verified by "four model checks" using ChatGPT, Claude, Gemini and Perplexity. - They disagree with each other. For project estimation software, Gitnux's top-ranked tool does not appear anywhere in WorldMetrics' top five. Three independent expert rankings would converge somewhat. Three generated ones do not have to.
Be precise about the size of this, because the honest number is small: those three sites contributed 181 citations, 2.4% of the total. They are not drowning the corpus. What they demonstrate is that the ranking layer has no floor — a site can manufacture a hundred thousand pages, cross-endorse itself, staff itself with invented people, and land in the top ten sources behind enterprise software recommendations. The 2.4% is the part somebody bothered to trace. Nobody has traced the other 59.8%.
A Third of Numeric Citations Don't Contain the Number
Independently, Haus Research asked 310 factual questions about 210 technology companies — founding date, funding, headcount, entry price, headquarters, revenue, breaches, CEO, acquisitions, uptime SLA — then fetched every unique URL the two search models cited, 2,915 of them for sonar alone, and checked each one against the specific figure it sat next to.
34.7% of citation-to-figure pairs failed: 633 of 1,826 pointed to a page that would not open or that contained none of the numbers in the sentence it was attached to. Scored per claim rather than per marker, 14.4% of 872 claims failed outright.
The breakdown matters more than the headline, because the two halves need different fixes. Of sonar's cited URLs, 78.7% were live and readable, 16.1% were gated behind a paywall or login, 2.5% were client-rendered and unreadable to a fetcher, 1.3% were dead, and 1.4% were unreachable after three attempts. Then, among the pages that did open, 16.1% contained none of the cited figures at all.
So one failure mode is access and one is attribution, and only the second is a correctness problem. Haus is fair about the first: "That is not a fault of the source... but it is a fault of the citation. A footnote a reader cannot open is a claim of provenance with no way to test it."
Pass rates tracked how canonically a fact is written down somewhere. Headcount passed at 82.0%. CEO identity passed at 44.3% — the worst category, and one an executive would assume was trivially checkable.
There is a durability problem sitting behind all of it. 25.1% of 1,432 cited URLs had never been captured by the Wayback Machine — not once, at any point in the life of the page — rising to 39.3% for directory sites. A business case whose evidence lives only on a page that no archive has ever seen is a business case with an expiry date nobody set. Trellner found the same shape from the other end: the median first Wayback capture was 2020 for the unranked cited domains against 2011 for the ranked ones, and 16.6% of the archived unranked domains were first seen in 2025 or later.
Paying for Sonar Pro Buys 15x Output Price, Not Accuracy
The premium tier does not fix this, and the price gap between the tiers is large. Haus measured sonar passing 65.9% of numeric pairs (95% CI 62.8–68.9) and sonar-pro passing 64.7% (61.5–67.7). The intervals overlap comfortably. Citation density was effectively identical at 9.8 and 9.7 sources per answer. Haus's conclusion is one line: "On this measurement they are the same product."
Per Perplexity's pricing page, sonar is $1 per million input and output tokens with request fees of $5–$12 per thousand; sonar-pro is $3 per million input and $15 per million output, with request fees of $6–$14 per thousand. Three times the input price and fifteen times the output price, for a statistically indistinguishable rate of citations that support their claim.
The two tiers also disagreed about basic facts. Trellner found sonar-pro citing datadryad.org for the research data platform Dryad while sonar cited dryad.co, which redirects to an Indonesian online-gambling portal. For the data-quality vendor Monte Carlo, sonar gave montecarlodata.com and sonar-pro gave montecarlo.com, which redirects to a Monaco casino. Across 1,502 vendor homepages returned in the study, 17 (1.1%) were gone or unreachable and 92 (6.1%) redirected to a different registrable domain.
If your evaluation template has a "vendor website" column populated by a model, that column has roughly a one-in-fourteen chance of pointing somewhere the vendor does not control.
This Is Not New, and That Is the Point
The strongest counter-argument is that this is a known, aging problem being re-measured. It holds up. In March 2025 the Tow Center for Digital Journalism ran 1,600 queries across eight generative search tools against 20 news publishers and found they answered incorrectly more than 60% of the time. Perplexity was the best performer at a 37% error rate; Grok 3 was wrong 94% of the time, and 154 of 200 of its citations led to error pages.
Eighteen months later, two independent teams measure the same class of failure on a different task. That is not a bug report awaiting a patch. It is the steady-state behaviour of a system that retrieves from an open web where generating a hundred thousand plausible pages costs almost nothing.
Two more honest caveats. First, the comparison set is thin: Haus ran GPT-4.1 with a web plugin as a control, and it emitted no inline markers, so no claim-level check was possible on it — it also cited 2.0 sources per answer rather than 9.8, with 36.4% pointing at the subject company's own domain against Perplexity's 23.4%. Denser citation is better behaviour that makes the product more auditable, and it is why Perplexity is the one that gets audited. Second, the top-cited source is G2, whose community guidelines permit vendor-solicited reviews with incentives capped at $100 and require them to be labelled "Incentivized." That labelling exists on G2's own pages. It does not survive the trip into an answer.
Then apply this article's own rule to this article. Both reports are self-published, both landed on 2 September 2026, and neither names an individual author, an institutional affiliation or a funding source — Trellner ran its queries the same day it published them. Their methods are stated in full and the pages they cite are still live, and two teams reaching the same conclusion by different routes is worth more than either alone. But by the standard in point 4 below, each is commentary rather than a primary source, and you should read them the way this piece asks you to read anything else. Trellner is explicit about the limits of a single pass — "One prompt wording, one run per category, no repeat sampling" — and notes that in an earlier pilot the product shortlist moved noticeably on rewording while the citation mix moved much less. Guideflow leads Gartner by 36 citations, so treat the exact ordering as indicative and the shape of the distribution — six in ten below rank 100,000 — as the finding that survives a re-run. Trellner also measured Perplexity alone, and says so: there is "no reason to assume" other engines' retrieval mixes match.
And the specific product under test is a moving target: Perplexity has announced that Sonar endpoints "retire on September 27." If you built a longlist pipeline on the Chat Completions endpoint, you have a migration due this month regardless. The endpoint changes. The web it grounds against does not.
Put the Provenance Gate Before the Shortlist
The fix is not to stop using answer engines for discovery. They are good at discovery. The fix is to stop letting discovery output cross into evaluation without a gate.
This Week:
- Pull the last three vendor longlists your team produced and open every citation. Not the summary — the links. Score each one three ways: does the page load, does it contain the figure it was attached to, and is the domain the vendor's own, an independent publication, or neither. Haus's finding says roughly one pair in three fails. Get your own number before you argue about theirs.
- Add a domain allowlist to any longlist pipeline you run through the API. Perplexity's search_domain_filter takes up to 20 domains and runs in allowlist mode or denylist mode, not both at once. This is an API-level control — teams working through the Perplexity Enterprise Pro interface, or through a retrieval product like Glean pointed at the open web, need the equivalent rule written into the process instead. It supports TLD filtering (
.gov) and path filtering. Twenty is enough for the analyst firms, the standards bodies, the regulators and the four or five trade titles that cover your category — and an allowlist that small is itself a useful forcing function on what your team actually trusts. - Check every "official website" the model returned against the vendor's own footer or a filing. Six percent redirect somewhere else.
This Month:
- Write the provenance rule into the evaluation template, above the feature matrix. One line: no candidate advances to scoring unless at least one supporting citation is a primary source — the vendor's own pricing or docs page, a regulatory filing, a named reference customer, or a paper. Candidates whose entire evidence base is sub-top-100K commentary get dropped, not investigated.
- Archive the evidence at capture time. Submit each load-bearing URL to the Wayback Machine when you cite it in a business case. A quarter of these pages have never been archived, and a business case that cannot be re-verified at renewal is a business case you will have to rebuild from scratch.
- Stop paying for the premium search tier on the strength of assumed accuracy. If your reason for sonar-pro is sourcing quality, the measured difference is 1.2 points inside overlapping confidence intervals. Keep it for latency or output quality if those hold up in your own testing; re-justify it otherwise.
Before Your Next Renewal:
- Ask the incumbent's account team for the primary source behind every number in their renewal deck. The same 34.7% applies to a slide built by a human using the same tools. This is now a reasonable question to ask, and the answer tells you something either way.
The Bottom Line
Enterprise software buying has always run on sources of varying quality — the analyst quadrant paid for by the vendors in it, the peer review written for a $100 gift card, the reference customer hand-picked by the account team. Buyers learned to discount each one because they knew whose incentives were in play. The answer engine's contribution is not worse sources. It is the removal of the label. A G2 review, a Gartner note, a vendor's demo blog and a hundred thousand generated pages arrive in the same list, in the same typeface, with the same footnote marker.
You cannot outsource the question of who is talking. Open the links.
Continue Reading
- Best RAG Platforms for Regulated Industries: Permissions First
- Glean vs Copilot vs Dust: Buy the Permission Model
- Apple's On-Device Model Is 99% Sure. Sample It Five Times.
- Agents Averaged 73. Only 30% Were Usable. Grade Pass/Fail.
- CoCounsel's New Model Runs on Qwen. Go Read the Card.
- Agent Orchestration Platforms: Score Exit, Not Features
- Walmart and Amazon Both Said 40%. Neither Ran a Holdout.
- One Firm Ran Three Labs' Cyber Evals. Name Yours.
