Where an AI answer gets its sources.
Retrieve, rank, synthesise, cite. Four stages, and you can fail at any one of them, for four completely different reasons. Knowing which stage drops you is the difference between fixing the problem and rewriting a page nobody read.
· By the Askwords team
“The AI doesn’t cite us” is not one failure. It is four, stacked, and each has a different fix. Rewriting your page when the real problem is that no crawler ever fetched it is a month spent on the wrong stage.
You do not need the internals of any particular product to work this out. Every answer engine that cites the live web follows roughly the same shape, and that shape is enough to diagnose your own problem.
A question becomes an answer in four steps
Each stage narrows the field. Dozens of pages become a handful, a handful becomes one answer, and that answer carries two or three links. Your job is to work out which narrowing removed you.
Is your page even in the candidate set?
A search runs, sometimes several, rewritten from the user’s question, and a few dozen pages come back. If you are not in that set, nothing later can save you.
Fails when: Blocked crawler, not indexed, no page on the topic, content only rendered client-side.
Is your page near the top of that set?
Candidates are ordered by relevance, authority and freshness: largely the same machinery as classic search ranking, which is why SEO still matters here.
Fails when: Thin page, weak domain authority, stale content, a competitor answering the question more directly.
Does your page contain a liftable answer?
The model reads the top candidates and writes one answer. It needs self-contained passages that state something specific and checkable.
Fails when: Answer buried below 800 words of preamble, vague claims, marketing language with nothing to quote.
Did the answer attach your link to the claim?
Sources are selected for the statements actually made. Different answer engines cite very differently: some link almost everything, some link almost nothing.
Fails when: Your fact was used but attributed to an aggregator that repeated it more clearly.
The question you were asked is rarely the query that ran
Answer engines rewrite. A buyer typing “what should I use to invoice clients if I’m a one-person plumbing business” does not produce one search for that sentence. It produces several shorter ones: invoicing software for tradespeople, best invoicing app sole trader, and so on. And the union of those results forms the candidate set.
Two consequences follow, and both are useful. First, exact-phrase optimization is pointless, because your phrase is not what ran. Second, breadth of coverage beats precision: pages that address a topic from several angles get pulled into more of those sub-searches.
The stage-one checklist
This is the stage where SEO is still doing the work
Candidates get ordered before the model reads them, and that ordering leans on the same signals search has always used: topical relevance, link and brand authority, freshness, and whether the page appears to answer the question rather than merely mention it.
This is precisely why GEO does not replace SEO. Everything that makes you rank makes you a candidate; being a candidate is the price of admission to the stages where the newer work happens.
Relevance
One page per question beats one page trying to serve five. Specific pages win specific sub-searches.
Authority
Slow to build, and mostly earned off your own site. This is where third-party mentions pay off twice.
Freshness
Weighted heavily for anything with prices, versions or a year in it. A visible, honest update date matters.
The model needs something it can lift
Now the model reads the top candidates and writes an answer. It is looking for passages that stand on their own and say something definite. A paragraph that only makes sense after the three above it is hard to use; a paragraph that asserts a specific, checkable thing is easy.
This is the stage where good marketing pages lose to mediocre documentation. The marketing page is written to build to a conclusion. The documentation states the conclusion and moves on, and that is the shape that gets synthesised.
“As we saw above, this depends on your setup. Every business is different, which is why our team works closely with each client to find the right fit.”
“Sole traders pay £12 a month. VAT-registered businesses need the £29 tier for Making Tax Digital submissions. Below about ten invoices a month, a spreadsheet is genuinely fine.”
Being used and being credited are not the same
The final stage attaches sources to statements. It is possible, and common, for your page to inform an answer without being linked, because another page said the same thing more quotably, or because the answer engine simply cites sparsely.
Watch for the pattern where you are mentioned accurately and cited elsewhere. The model clearly knows your specifics and is crediting an aggregator that repeated them. The fix is usually to be the clearest statement of your own facts, not to publish more of them.
Four symptoms, four different weeks of work
Run your question, then work down this table until a row matches. The first match is your stage.
| What you observe | Stage | What to do |
|---|---|---|
| Not in normal search results either | Retrieve | Crawler access, indexing, then whether the page exists at all. |
| Ranks on page three, never cited | Rank | Ordinary SEO: depth, internal links, authority, a genuine update. |
| Ranks well, still never cited | Synthesise | Restructure. Answer in the first paragraph, specifics with numbers, self-contained sections. |
| Described accurately, credited elsewhere | Cite | Own your own facts. Put them plainly on a page that is unmistakably yours. |
One caution: check this on more than one answer engine before deciding. A retrieval failure on one engine and a synthesis failure on another is entirely possible, and it means two different jobs rather than one.
Fix the stage that dropped you, not the stage you enjoy working on.
Keep reading
Sources and further reading
- Google Search Central: AI features and your website
- Google Search Central: Optimizing your website for generative AI features
- Google Search Central: How Search works (crawling, indexing, serving)
Askwords measures a selected set of buyer questions against fresh answers from each engine. AI answers vary by model, time, context and personalization, so a check is evidence of the answers observed, not a universal judgement about a brand.
