URL to Any logoURL to Any
Back to blog
How AI Answer Engines Choose Sources: What 215,128 Manufactured Pages Reveal

How AI Answer Engines Choose Sources: What 215,128 Manufactured Pages Reveal

A new study found AI answers citing mass-generated "best software" pages. How AI answer engines choose sources, and how to make real pages extractable and citable.

Sep 3, 2026URL to Any

On September 2, a research team asked Perplexity’s models for the best software in 380 categories and kept every URL the answers were built on. Nearly a quarter of the 7,534 citations pointed at domains too small to make the top one million sites — and three of the most-cited domains turned out to be one operation that had mass-produced 215,128 “best software” pages, addressed not to readers but to the models themselves.

The uncomfortable lesson is not that AI search is broken. It is that AI answer engines choose sources differently from a search engine: they don’t rank the whole web, they retrieve documents that are easy to fetch, read, and quote. If your pages are genuinely useful but hard to extract, you lose citations to sites that are easier to read — including sites nobody human ever visits. The good news: the fix for an honest publisher is not to out-spam the spammers. It is to make real content extractable, and to leave the fingerprints that machines and readers can both verify.

Key takeaways

  • AI answer engines assemble answers through retrieval: they fetch a handful of documents and quote them. Document choice is driven by what can be fetched and parsed cleanly — which is gameable, as a new study demonstrates.
  • In a September 2026 study of 7,534 Perplexity citations, three apparently related sites had generated 215,128 category pages; a vendor’s marketing blog was cited more often than Gartner.
  • The manufactured route is fragile: the study’s examples contradict themselves — different rankings, invented staff, unrendered template variables — and are exactly what engines will learn to filter.
  • The durable play for real publishers: genuine value, extractable structure, truthful metadata, and a five-minute audit of how a machine actually sees your page.

The vocabulary: grounding, retrieval, extractability

Three terms do most of the work in this story.

Grounding is the step where an AI system fetches documents to condition its answer on, instead of answering from memory alone. When Perplexity or Google’s AI Mode answers “what’s the best CRM for a small agency,” something first retrieves candidate pages, then the model writes an answer based on them and attaches citations.

Retrieval is that fetch step. It behaves less like a ranked list of the whole web and more like a fast librarian: given a question, it grabs documents that look relevant and machine-readable. Google documents something similar for its AI features — a technique called query fan-out, where the system issues multiple related searches across subtopics and sources, then identifies more supporting links than a traditional results page shows.

Extractability is how cleanly a page’s content survives that trip. A retrieval system fetches your HTML and turns it into text — headings, paragraphs, tables — before the model ever reasons about it. Pages that render their substance as clean text and structure are easy to quote. Pages that bury their content in scripts, images of text, or layouts that parse badly are effectively invisible, no matter how good their content is.

This is also where GEO (generative engine optimization) parts ways with classic SEO. SEO competed for a ranked list position. GEO competes for a slot in the retrieval set that an answer is built from — a different selection layer with different weaknesses, as the study shows.

What the study actually measured

A small research outfit, Trellner, ran one of the more careful measurements of this layer to date. On September 2 it sent 380 buyer-intent software categories — from “CRM software” down to “museum collection management software” — to two Perplexity models through OpenRouter, one prompt per category per model, 760 calls in all, and kept every URL the models reported retrieving. That produced 7,534 citations across 2,055 domains.

The Trellner Research report page titled 'Three sites made 215,128 best software pages for AI. Perplexity cites them' with its summary statistics

Source: Trellner Research, report TR-2026-009 (trellner.com), September 2026. Accessed 2026-09-03.

The headline numbers:

  • 59.8% of citations pointed at domains ranked worse than #100,000 in the Tranco popularity list; 23.4% weren’t in the top million at all. The median cited domain that did rank sat at #71,611.
  • The most-cited domains were predictable — G2 (291 citations), Reddit (261) — but third place went to guideflow.com, a vendor blog about interactive product demos, cited 194 times across 96 of the 380 categories. It is not a review site and competes in none of the categories it was cited for. Wikipedia, for scale, was cited three times.
  • Domains outside the top million were also young: among archived cited domains, the median first capture was 2020 for unranked ones versus 2011 for ranked ones, and one in six of the archived unranked domains first appeared in 2025 or later.

Then there are the three that gave the report its title.

Homepage of worldmetrics.org presenting 'Business Intelligence on Software & Markets' with 'Rigorous editorial process' and 'Cited by Hundreds of Publications' claims

Source: worldmetrics.org homepage, accessed 2026-09-03. Its HTML title reads ‘Worldmetrics — Facts & Grounding Page’ — addressed to retrieval software, not readers.

wifitalents.com, worldmetrics.org, and gitnux.org were all registered within six months of each other in late 2023–mid 2024, share the same nameservers, and run the same page template. Their sitemaps list roughly 100,000 URLs each — together 215,128 auto-generated “best <category> software” buying guides, against six blog posts per site. Each homepage carries the HTML title “<Brand> — Facts & Grounding Page,” with near-identical meta descriptions describing “verified facts” in “one machine-readable record.” Grounding, remember, is the engine’s word, not a buyer’s. These pages speak directly to the software that reads them.

Why machine-made pages surfaced — and the tells

Why did retrieval favor them? The mechanism is mundane. Retrieval systems need documents they can fetch and quote, and these sites are built for exactly that: clean templates, consistent structure, JSON-LD rankings a parser can read without interpretation, and enormous surface area — 215,128 pages is a lot of lottery tickets for niche category queries. The report is careful to note nothing about the pages is technically deceptive; they are simply addressed to machines, in their own titles and descriptions. And the same study found a vendor’s ordinary content-marketing blog — guideflow’s — absorbed citations the same way, which tells you the selection layer rewards publishable, extractable volume from anyone.

The tells are in the content. Asked about “project estimation software,” the three sister sites produced three different top-five rankings. Their pages credited nine different staff names in total. One page labeled its results “AI-verified · Expert reviewed” — while another left an unrendered template variable in the byline, literally reading “Within the next 26 days.” And worldmetrics.org sells what the pages advertize for it: custom market research from €5,000, ready-made reports from €499, vendor selection from €2,500, stacked above the same generated Best Lists the models retrieve.

Two honest limits deserve their space here, and then we can move on. First, the study measures which documents form the evidence base — it did not test whether removing these sources would change any answer. The manufactured sites may be quoting reasonable products; the problem is what the evidence layer is made of. Second, the two Perplexity tiers turned out to share one retrieval layer (their citation lists were byte-identical in 289 of 380 categories), so this is one engine’s stack sampled twice, not a web-wide survey. Treat the study as a sharp case study of the mechanism, not a census of the internet.

The decision this hands you

If you publish a tool site, a documentation hub, or any genuinely useful content, the story poses one decision: chase the retrieval layer with volume, or make your real pages worth retrieving.

The volume route now has an entire industry around it — “get found by AI” consultancies were already a HN talking point within a day of the study. But the study itself shows why it’s a bad trade. Machine-addressed content has to fake the signals retrieval and readers both check for — authorship, editorial process, consistent verdicts — and fakes don’t scale cleanly: they leak template variables, contradict each other, and require inventing people. Engines have strong incentives to filter exactly these patterns, and Google’s own guidance for AI features points the other way: no special markup, no AI text files, no secret schema — just the standards, with important content “provided as text” and structured data that matches what’s visibly on the page.

The durable route is almost boring: publish less, mean it, and make sure a machine can read it. That’s not a compromise — it’s a moat, because the manufactured approach structurally cannot pass the checks your real site passes.

What that route rejects, concretely:

  • mass-generating category or comparison pages you can’t stand behind;
  • invented authors, fake “editorial processes,” fabricated credentials;
  • pages written for the crawler first and the reader second.

A five-minute audit for your own pages

Here is the practical half. Pick any page you actually care about — a tool page, a guide — and check what a machine sees when it fetches it.

1. Fetch your page the way a machine would. Paste your URL into URL to Markdown and look at what comes back. This is roughly the shape of text a retrieval system has to work from. Is your core content there? Or did it come back as navigation, cookie banners, and half a paragraph?

URL to Any's URL to Markdown tool showing a real webpage converted into clean Markdown text

2. Check the heading structure. The same Markdown output doubles as an outline: every # and ## is a heading a retrieval system will lean on. Read just those lines top to bottom. If they sketch a coherent answer to the page’s question, a quoted excerpt can too; if it’s a jumble of marketing lines, that’s what a citation will pull.

3. Read your own title and meta description like a machine would. The manufactured sites gamed exactly this layer — “<Brand> — Facts & Grounding Page.” A meta tags extractor shows what your page announces about itself. It should state plainly what the page is and for whom. Truthful metadata is also self-defense: when your title says exactly what you are, you don’t look like the sites whose titles say what they wish they were.

4. Confirm the substance is text. Tables, specifications, steps, prices — if they exist only as images or render behind scripts, extraction loses them. Google says it directly: keep important content as text.

5. Make structured data match the visible page. If your JSON-LD claims a rating, an author, or a review, the visible page must show the same thing. Consistency between markup and visible content is both a search-engine guideline and the exact test the manufactured sites fail.

For a deeper look at the fetching-and-parsing half of this — how agents negotiate Markdown, and a quick browser check — see our earlier piece on how AI agents read web pages.

The citation you want

The study’s authors close by saying they measured the evidence base, not the answers it produces. Fair — and the same humility helps publishers: no audit guarantees a citation, and no amount of structure substitutes for something worth citing. What the 215,128-page episode really shows is that the retrieval layer reads structure before it reads substance — so the winners will be the publishers who have both. Make your real content easy to fetch, honest about what it is, and verifiable by a machine and a human alike. That’s a citation no one has to manufacture.

Related Articles