fireflyoGet your free report
← all posts

July 18, 2026 · 3 min read

How AI engines choose which brands to recommend: the four filters

When an engine receives a commercial prompt, four filters run in sequence — retrieval, parsing, synthesis and citation. A brand can fail silently at any of them. Here is what each filter does and the lever that moves it.

mechanicsgeo-basics

TL;DR — Every commercial prompt runs candidate brands through four filters: retrieval (does the engine fetch pages that mention you?), parsing (do those pages survive being chunked and read?), synthesis (do independent sources agree about you?), and citation (are you the source the answer links to?). Filters 2 and 4 are technical and move in weeks. Filters 1 and 3 are reputational and move in months. Most GEO disappointment comes from vendors selling filter-2 tooling with filter-3 promises.

The pipeline, in order

When someone asks an AI engine “what’s the best X for Y,” the engine doesn’t consult a stored opinion. It runs a live pipeline, and a brand can drop out at any stage.

Filter 1 — Retrieval

The engine issues live searches (against its own index or partners’) and pulls candidate pages for the topic. If AI crawlers can’t fetch your pages — or you’re absent from the third-party pages that reliably get retrieved for your category — you never enter the candidate set at all.

The lever: indexability and crawl access for AI bots (GPTBot, ClaudeBot, PerplexityBot), ranking well enough to enter the candidate set, and presence on the third-party pages engines actually pull.

Filter 2 — Parsing

Retrieved pages are chunked and read by the model. Clearly structured, directly-answering passages survive; meandering prose is discarded. A page can be retrieved and still contribute nothing because no chunk of it directly answers the question.

The lever: answer-first paragraphs, comparison tables, schema/JSON-LD, and headings that mirror buyer questions. This is the “patchable” layer and the highest-leverage quick win.

Filter 3 — Synthesis and trust

The model weighs agreeing sources. Names supported by multiple independent pages are asserted confidently; one-source claims are dropped or hedged. This is why a brand with consistent descriptions across reviews, comparisons and community discussion beats a brand with one great page.

The lever: cross-source consensus — review platforms, comparison sites, community mentions that describe the brand consistently.

Filter 4 — Citation

The engine attributes its answer to a handful of links, which then receive whatever clicks exist. Cited pages get retrieved more in future runs, so presence compounds.

The lever: being the citable source — or dominating the pages that are.

The strategic implication

Filter Nature Timescale
1 · Retrieval Reputational Months
2 · Parsing Technical Weeks
3 · Synthesis Reputational Months
4 · Citation Technical Weeks

Filters 2 and 4 are fast to influence — that’s where productized tooling lives. Filters 1 and 3 are slower — that’s where content strategy and consensus building live. A credible GEO program is explicit about which filters it moves and on what timescale.

Where to start

Run a diagnostic before buying anything: for your top 20 buyer prompts, which filter is actually failing? A brand that’s retrieved but never parsed needs different work than a brand that never enters the candidate set. Fireflyo’s free report breaks your category down filter by filter.