Getting cited by AI means being retrievable, worth citing, and corroborated by third-party sources. The 3Bs Framework runs AEO strategy and execution as a flywheel. The 3Bs stand for benchmarking where you stand, breaking down why the gaps exist, then building what closes them.

A common problem my clients face is being ranked 1-5 on Google for their main category term but rarely being cited or mentioned on ChatGPT and other answer engines.
Ahrefs data shows that this isn’t a fluke; it’s not just impacting the handful of clients I work with. In March 2026, they analysed 863,000 SERPs and found that only 37.9% of URLs cited in Google AI Overviews featured in the first ten results (down 76% YoY). The rest split almost evenly between positions 11 to 100 and even pages not ranking in the top 100 at all.
That tells me that rankings and citations have become two different outcomes. A page can do everything Google rewards and still get skipped by AI engines, so doing more of the same SEO work won’t solve the problem.
Over the last few months, I’ve been developing my process for earning AI citations and mentions, built specifically around how these engines select and use sources. I call it the 3Bs Framework (Benchmark, Blueprint, and Build). This article details that process and other tips to get cited by AI.
Key Takeaways
- Getting cited by AI requires a consistent loop of diagnosing what is working and what isn’t, strategizing the next steps based on those findings, and executing; it’s not achievable through a one-time fix.
- Through trial and error, I’ve distilled this process into my proprietary approach — The 3Bs Framework. The 3Bs stand for “Benchmark,” “Blueprint,” and “Build.”
- Benchmark: Set up tracking tools and measure your citations, mentions, and other relevant metrics across ChatGPT, Gemini, Perplexity, and Google AI Overviews.
- Blueprint: Figure out why you’re not showing up before fixing anything. Being invisible, cited without being named, and described inaccurately are three different problems with three different solutions, for example.
- Build on-site first: Opening with direct answers, clear headings, consistent messaging, and original data make you eligible to be cited, but that alone might not be enough to be selected.
- Build off-site next: This can increase AI citations by up to 325% (per Stacker and Scrunch data) and includes coverage on review platforms and directories, earned media, and social media mentions.
- Know that citations and mentions are different and require different strategies.
How AI Assistants Decide Which Brands to Cite
AI assistants don’t rank pages. They retrieve passages. A model can only answer from what it fetches, chooses, and extracts, not everything that exists on your site or that is true about your brand.
This is because all answer engines, be it Google’s AI Overviews, ChatGPT, Gemini, Perplexity, or another, are bound by three constraints when retrieving information. That’s what I’ll explain here.
1. Answer Engines Can’t Search Your Whole Prompt
A long, conversational question gets broken down and translated into several query-shaped sub-searches before retrieval starts. This is query fan-out, and it is worth understanding in depth because it decides whether you are in the running before extraction comes into play.

A study of five million ChatGPT queries found the engine combines results from its sub-queries using Reciprocal Rank Fusion, so a source appearing across several sub-searches gets weighted more heavily than one that only shows up once.
Google describes a near-identical mechanic for AI Overviews and AI Mode, and Robby Stein, Google’s VP of Product for Search, has walked through a concrete example of it, researching home safes, where the system ran separate sub-searches on fire ratings and insurance implications before assembling one synthesised answer with specific product picks.
Perplexity‘s guide on its Search API describes scoring content at the sub-document level, meaning a single passage can be selected without the rest of the page being relevant at all.
And independent technical breakdowns of Perplexity’s Pro Search and Deep Research estimate three to five sequential sub-searches per complex query before it synthesises a cited answer.
🗝️ The practical result is that each sub-query behaves almost like an ordinary keyword search, matched against an index and ranked by relevance and freshness the same way a typed Google query always has. But your page ranking first for one keyword is only one input into one of several parallel searches running behind a single answer.
(I go further into this in my breakdown of AEO, GEO, AIO, and LLMO.)
2. Answer Engines Can’t Open the Whole Web
Out of everything a sub-search could fetch, the engine chooses a handful of candidates based on how each one presents itself before it has read a word of the actual content. This includes the title, the URL, the snippet, and sometimes a freshness date.
Think of this as your page’s cover. It’s the only information the engine sees before deciding whether to open the page at all, and a page with excellent content sitting behind a weak cover often never gets the chance to prove it.

3. Answer Engines Can’t Read the Whole Page
Once something is opened, the engine has to pull a usable chunk out of it rather than absorb the whole document, which is why passage-level structure matters.
Two independent studies point to where in a page that chunk usually comes from, and they land on the same pattern from different platforms:
- CXL’s analysis of 100 Google AI Overview citations found 55% come from the first 30% of a page and only 21% from the bottom 40%.
- Kevin Indig’s analysis of 18,012 verified ChatGPT citations found nearly the same shape on a different platform — 44.2% of citations from the first 30% of a document. He calls it the “ski ramp effect,” a steep drop after the first third followed by a long, flat tail.
What gets extracted is also short. VisibilityStack traced 2,422 AI-generated sentences back to their source pages and found only 24% could be matched to a specific passage at all. The rest were synthesised from the page rather than quoted, meaning your content shaped the answer without a single sentence surviving intact.
Where a passage was traceable, the median length across all five surfaces studied was about 25 tokens, roughly 19 words. Claude and Gemini stayed almost entirely under that ceiling. Google AI Overviews was the outlier, occasionally pulling passages up to 500 tokens.
Additionally, retrieval is text-first. So, an answer that lives only inside an image, chart, or video with no surrounding text usually has nothing for AIs to work with, however good it is.
⚠️ AI Answers Are Not Deterministic
Beyond the three constraints that shape how AI engines decide which brands to cite, it’s important to understand that AI answers are not deterministic.
Rand Fishkin of SparkToro and Patrick O’Donnell of Gumshoe.ai had 600 volunteers run 12 identical prompts through ChatGPT, Claude, and Google’s AI almost 3,000 times.
The odds of getting the same list of brands twice came in under 1%, meaning fewer than 1 in 100 prompt searches. The odds of the same list in the same order were closer to 1 in 1,000. Even the number of brands returned moved around, from two or three up to ten or more.
That does not make measurement pointless, but it does change what measurement has to look like. The same study found another important thing. Across many runs of many prompts, the rate at which a brand appears is stable even though any individual answer is not.
In the headphones category, for example, the same handful of names came back in 55 to 77% of responses regardless of phrasing.
The conclusion I derived from this is two-fold:
- Tracking prompts manually doesn’t work. We can only draw conclusions from volume — dozens of prompts run repeatedly across every engine that matters, and nobody produces that by typing questions into a chat window a few times and writing down what comes back. This is what tracking tools are for.
- While the same well-known brands get mentioned in 55-77% of answers, the extreme variation in output means that there are plenty of chances for smaller brands to be featured, meaning there’s a lot to gain from optimizing your brand for AI search.
With that in mind, let’s explore my framework for getting you cited in AI engines.
The 3Bs Framework: How to Get Cited by AI in Three Stages
The framework has three stages. Benchmark tells you where you stand. Blueprint tells you why and what to fix first. Build is the work itself, split between your own pages and everything that exists about you elsewhere.

It runs as a loop rather than a checklist, and it behaves like a flywheel. Each turn compounds the effects and makes the next one easier.
For example, in early cycles, the Build stage improves and/or produces content, which is necessary for AI systems to surface you at all. But in later stages, it focuses on getting you mentioned in third-party platforms, which is the strongest predictor of whether AI systems will surface you, according to an Ahrefs study.
What follows is an explanation of my three-step process. I explain it in more detail on this page: The 3Bs Framework for Getting Cited by AI Search Engines.
B1. Benchmark: Measure Your Current AI Visibility Across Engines
Benchmark means establishing a reliable baseline of how AI engines currently cite and describe you. That requires a tracking platform running a broad prompt set repeatedly across engines, because the run-to-run variance makes any hand-collected sample unreliable.
You cannot fix a gap you have not measured.
Define the Attributes First, Prompts Second
Start from what you want to be known for, meaning your category terms, verticals, use cases, and the specific capabilities you want associated with your name. Those attributes are the unit you will measure against. Individual prompts are just instruments for reading them.
Then build a spread of prompts per attribute rather than one canonical phrasing since phrasing moves results violently.
Real buyer prompts also carry far more context than keywords do. Semrush’s analysis of 80 million ChatGPT clickstream records found the average ChatGPT prompt runs around 23 words, against roughly four words for a typical Google keyword.
John-Henry Scherck of Growth Plays frames the alternative to one-time prompting as “PSUC,” meaning persona, segment, and use case, because a vendor that appears for a generic category prompt often disappears once the buyer adds industry, company size, or integration requirements.
Lastly, cover several prompt types, not just commercial ones:
- Discovery prompts test whether you make the list at all.
- Comparison prompts test whether you win a head-to-head.
- Validation prompts ask a direct yes-or-no about whether a capability is yours.
- That last type does specific diagnostic work in B2, which I’ll discuss later.
Use a Prompt Tracking Platform
This is the part where I disagree with many experts on this topic who advise you to run twenty or thirty prompts by hand and log what comes back. Given what the SparkToro data shows about variance (discussed above), that advice produces a number that’s closer to a coin flip than real evidence.
One prompt, on one platform, at one moment is a data point, not a read, and a team that has been spot-checking for months has usually built a narrative rather than a baseline.
What you need is a platform that runs a broad prompt set across the answer engines that matter for your brand repeatedly and aggregates the results.
Ahrefs Brand Radar, Semrush’s AI toolkit, Profound, and Scrunch all do a version of this at different price points. The specific tool matters less than the fact that something is sampling at a volume no person can replicate by hand.
When analyzing data within the platform, read two numbers:
- Visibility score is how consistently you appear, meaning the share of responses that mention you.
- Visibility rank is where you sit against every other brand appearing for that attribute.
High rank with a low score means the models put you first when they think of you but have not settled on you yet, which is a different problem from being included but ranking low consistently.
Then split citation rate from mention rate, by engine, because the two diverge sharply. In the Semrush and Kevin Indig dataset:
- ChatGPT cited sources in 87% of brand appearances but named the brand in only 20.7%.
- Gemini did almost the reverse, naming brands 83.7% of the time while citing them 21.4%.
Any report that averages across engines hides something you need to see.
Set Up Tracking in GA4 and GSC
Add an AI channel group in GA4 filtering for the AI engines that matter to your brand (e.g., chatgpt.com, perplexity.ai, claude.ai, and gemini.google.com), so AI referral traffic is separated from the rest.
Check Google Search Console’s own Generative AI performance report too, since it covers something GA4 can’t. It shows impressions, meaning how often your pages appeared inside AI Overviews and AI Mode, broken down by page, country, and device. There is no click data, no CTR, and no query-level detail, so far.
It is worth checking anyway, because it’s the only place that shows zero-click AI Overview visibility directly. GA4 referral tracking only catches AI traffic that clicks through.

Over time, the data from GA4 will help you understand which AI platforms generate the most traffic to your website, which will help you refine how you use your prompt tracking platform and set strategy.
Audit Whether You’re Fetchable and Extractable
Open your robots.txt and check for GPTBot, Google-Extended, PerplexityBot, ClaudeBot, and Applebot-Extended. Blocking these crawlers means opting out of the channel. This is a one-and-done step in the 3Bs Framework; the following two are ongoing across cycles.
Then review your top pages for extractability. Does each H2 answer one question in its opening sentence? If you don’t already have a website content inventory, build one first, because you can’t audit or prioritise what you haven’t mapped.
Lastly, check your entity consistency, meaning whether your own site, LinkedIn, G2, Crunchbase, and anywhere else you appear all describe you in the same category using the same language. Inconsistent descriptions make you ambiguous to AI engines.
B2. Blueprint: Diagnose the Gaps and Decide What to Prioritize
Blueprint means diagnosing why the gaps exist and turning that into a plan. Classify which of the three failure modes you have, map who is winning instead, then decide what deserves your attention first.
Skipping this stage is why so much AEO effort lands in the wrong place. AEO doesn’t just come with different tactics and KPIs from SEO; the way you analyze the data and strategize is also different. Let’s break it down.
Classify the Gap
There are three failure modes, each requiring a different response.
- The first is being invisible, meaning neither cited nor mentioned. That is either a content or eligibility problem.
- The second is being cited without being named. Your domain feeds the answer, but your brand never appears in it. These are ghost citations, and they are a positioning problem. Comparative content gets brands named while informational content feeds the machine anonymously.
- The third is being mentioned inaccurately or unfavourably. When AI systems lack clear, consistent information about you, they fill the space with whatever they can find.
Good prompt selection, which I discussed in B1, can help you dig even deeper when diagnosing the gap. For example, say you’re invisible. How do you know if it’s a content or eligibility problem?
Validation prompts help you answer this by directly asking a question about your brand, even if you’re absent from discovery for an attribute.
Does [brand] do X?
- If the engine says yes, it knows you have the capability and simply doesn’t recommend you, which is a trust and amplification problem that can’t be solved with on-page content.
- If it doesn’t know, you have a comprehension problem, and the fix is upstream in clearer descriptions and documentation.
Those are two completely different programmes of work, which will define the work you perform in B3, Build. And without the second prompt type, which was set up in B1, Benchmark, you wouldn’t be able to tell which one you are looking at.
Map Your Performance vs. Competitors
First, know that you should scope this analysis to discovery prompts, where no brand name appears in the question. Once your own name is in the prompt, engines pull more heavily from your own pages, and the picture skews toward owned sources rather than what the wider web actually says.
That said, this step produces three different lists.
1️⃣ The first list is actual competitors, meaning brands contesting the same commercial ground as you.
Pull this for your two or three closest rivals and track two numbers:
- Visibility score is how often you’re mentioned.
- Rank is your stack-ranked position against those same rivals for that attribute.
Rank is the more reliable of the two, since a decent-looking score can still mean a distant fourth or lower place. Pull this at the prompt level too, not just the topic average, because an average happily hides one prompt collapsing while another climbs, and the collapsing one is usually the more urgent problem.
The other two lists include everyone else in your citation mix who isn’t a direct competitor.
2️⃣ The second list includes earned media, publishers, review sites, and affiliate content.
This list is actionable, and whether one is genuinely independent or a competitor’s own placement is worth checking before you pitch it.
If sources like these dominate the mix for an attribute, that’s your outreach target list for B3, a specific, ranked set of publications you can pitch, place in, or partner with.
3️⃣ The third is for institutional sources, Wikipedia, government sites, and major reference bodies.
This one isn’t actionable in the same way. If these sources dominate, owned content faces a steep climb no matter what you publish, and the honest move is to put your effort into the earned and partnership route rather than trying to out-content an institution.
💡 One final tip: If an AI assistant is citing Reddit for a topic you should own, it usually means no brand has published a good answer.
Segment by Query Type
This is where the sorting from B1 pays off. In the Semrush data:
- Informational queries produced an 89.3% citation rate but only an 18% mention rate.
- Comparative queries named brands 43.3% of the time, roughly 2.4 times more often.
- Short conversational prompts produced 30 to 50 times more brand mentions than long structured ones.
Track these as separate numbers rather than one blended score. Rolling informational and commercial prompts into a single visibility figure hides exactly this pattern, since the two behave too differently for an average to describe either one honestly.
If your gap is concentrated in comparison prompts, comparison content is your first move.
Then Prioritise
Score opportunities on commercial value, gap size, and whether you can credibly claim authority on the topic. This is the same scoring logic I use when building a repeatable SEO strategy system, and the competitor gap analysis you run there transfers directly.
Where you start depends on where you already are:
- Strong existing SEO and solid content means go straight to off-site work.
- Thin or badly structured content means fix the foundations first, because off-site authority cannot rescue a page nobody can extract from.
B3. Build: Create the Content and Authority That Earn Citations
Build has two halves (plus maintenance work):
- On-site building makes you eligible to be cited.
- Off-site work is what gets you selected.
Build On-Site
➡️ Aim to lead every section with a direct answer in the first 40 to 60 words. The positional data supports this more strongly than any other structural finding. Nearly three-quarters of cited sentences come from the first half of a page, and the bottom quarter of a page earns 7.4% of citations. If your section opens with context setting, an assistant will skip past it and pull from whoever got to the point faster.
➡️ Write headings as the questions your readers actually ask. AirOps’ 2026 State of AI Search report, drawn from millions of data points, found that pages with sequential, well-organised headings were 2.8 times more likely to be cited, that 68.7% of pages cited in ChatGPT follow a logical heading hierarchy, and that 87% use a single H1.
➡️ Make sections standalone. If a passage stops making sense once removed from the page around it, it will not survive extraction. Use tables for comparisons, numbered lists for processes, short paragraphs carrying one idea each. Aim for quotable sentences in the six to seventeen word range.
➡️ Raise your factual density. The GEO paper by Aggarwal and colleagues, presented at KDD 2024, tested nine optimisation methods across 10,000 queries. Adding statistics, adding quotations, and citing authoritative sources were the three strongest, each producing roughly 30 to 40% relative improvement in visibility, with an upper bound around 40%.
The effect was larger for lower-ranked sources, and one analysis of the paper puts the gain from citing external sources at 115% for lower-ranked content. Smaller brands have more to gain here than established ones.
⚠️ Two caveats worth knowing:
- That figure is a maximum under favourable conditions, not an average. And a later benchmark, C-SEO Bench, tested a broader set of conversational SEO tactics and found most of them do not help, while plain source relevance keeps working. My take from experience is that structure and evidence do help.
- If you cite someone else’s statistic, the assistant credits the original source, not you. Original data is the only version of this tactic that earns you the mention rather than earning it for whoever you quoted. Most companies are sitting on usable first-party data already, in aggregate product usage, anonymised customer outcomes, support ticket patterns, or internal benchmarks they have never published.
➡️ Add Article, Organization, and HowTo schema. I want to be honest about the evidence here, because the effectiveness of this is generally oversold. Ahrefs tracked 1,885 pages that added JSON-LD between August 2025 and March 2026 against roughly 4,000 matched control pages, using a difference-in-differences design to strip out platform trends.
AI Mode moved 2.4% and ChatGPT 2.2%, both nearly statistically irrelevant. AI Overviews moved minus 4.6%, the only statistically significant result, and in the wrong direction. Note also that Google retired FAQ rich results in 2026, so FAQPage schema is worth less than it was.
However, the study has a real limitation. Every page tested already had 100 or more AI Overview citations, so it measured the marginal effect on pages engines could already see. It doesn’t prove schema is useless for a page starting from zero. My opinion is that schema can help make your content unambiguous to machines, but it’s not a growth lever.
➡️ Finally, get your content out from behind JavaScript, accordions, and read more toggles. Retrieval systems generally extract visible HTML during a live page fetch, so content that requires an interaction often does not exist as far as they are concerned.
Build Off-Site
The clearest evidence for building off-site comes from a controlled test. Stacker and Scrunch took eight articles and ran them in two conditions:
- Hosted only on the brand’s own domain
- Distributed across third-party news sites
Then, they measured 944 prompt and platform combinations across five leading LLMs.
- Brand-hosted content was cited 7.6% of the time.
- The same content distributed off the brand’s domain reached 34%, a ~325% increase. A March 2026 follow-up on a broader sample found a 239% median rise.
🎯 In roughly 1 in 5 cases, the engine cited the syndicated version and never cited the brand’s original at all.
In other words, your website makes you eligible, but other people talking about you is what gets you mentioned. That means:
- Claiming and actively maintaining your category review profiles on G2, Capterra, and Trustpilot.
- Earning coverage in publications your buyers genuinely read, prioritising relevance over raw domain authority.
- Participating honestly in the communities where your category gets discussed.
- And going multimodal, particularly on YouTube, which Ahrefs’ December 2025 follow-up found to be the single strongest individual signal it measured.
- Clean up your entity descriptions while you are here.
- Then think about PR carefully. For example, when you give a quote, include the exact positioning language you want assistants to repeat about you, because a vague quote gives them nothing to work with.
One place where practitioners genuinely disagree is backlinks. The Ahrefs correlation data points to backlinks not mattering much for AI visibility, putting backlinks at 0.218 against 0.664 for unlinked brand mentions.
My own view is that links still carry value for SEO, and traditional rankings still feed Google’s AI surfaces, so they are not irrelevant. But if you have limited hours, an editorial mention with no link in a publication your buyers read will do more for your AI visibility than a link from a site nobody in your category has heard of.
Finally, be realistic about time. While structural changes can show up in citation patterns within weeks, off-site signals compound over months.
Maintenance
Refresh your priority pages on roughly a 90-day cycle. AirOps found that pages not updated quarterly are three times more likely to lose citations, that 83% of citations for commercial and evaluation queries come from pages updated within the past twelve months, and that more than 60% come from pages refreshed within six.
Don’t change the dates cosmetically, though.
How the 3Bs Loop Closes and Compounds
The loop compounds because each cycle starts from a stronger position than the last. Carry three numbers forward: citation rate, mention rate, and share of voice, then run the stages again.
- Citation rate, meaning how often your domain appears as a source.
- Mention rate, meaning how often your brand is named in the answer.
- Share of voice, meaning your portion of total citations and mentions against competitors.
Track the pair rather than either alone. AirOps’ 2026 State of AI Search research shows that brands earning both a mention and a citation are 40% more likely to reappear in the next answer than citation-only brands, yet only about 28% of answers contain a brand with both signals. Persistence is rare — only around 30% of brands stay visible from one answer to the next, and fewer than 20% survive five consecutive runs.
Review monthly and reassess strategy quarterly. Because a tracking platform is sampling continuously, the temptation is to watch the dashboard daily, and that mostly means watching variance.
The compounding is the point. Your second cycle skips prompt set construction entirely. The publication list from B2 is already built and partly warm.
And because branded web mentions are the strongest correlate of AI visibility that anyone has measured, the mentions earned in B3 raise your baseline before you write a single new page during the following cycle, including on prompts you never targeted.
Citations vs. Mentions: Why the Distinction Changes Your Strategy
A citation is when an AI system links your domain as a source for its answer. A mention is when it names your brand in the answer text. They frequently happen independently of each other, and confusing them will send you after the wrong fix.
Commercially, they do different things. A citation is measurable, since it can send a click you can see in analytics. A mention is what actually influences a buyer, because it is your name appearing inside a recommendation.
The Semrush and Indig breakdown of 3,981 brand appearances quantifies how loosely coupled the two are:
- 61.7% were citations with no mention
- 25.1% were mentions with no citation
- Only 13.2% were both
You can be feeding answers across your entire category, generating no recognition whatsoever, and conclude from a citation report that things are going well.
That gap is also a strategy signal:
- Informational content earns citations at 89.3% but mentions at 18%.
- Comparative content earns mentions at 43.3%.
It tells you whether to prioritise informational content or comparative content, and for most brands the answer leans towards comparative, which is not where most content calendars are pointed. My framework optimises for both, but treats mentions as the harder and more valuable outcome.
How Citation Behaviour Differs by AI Platform
Citation behaviour varies enough between engines that averaged reporting hides the picture. ChatGPT cites often and names rarely, Gemini does the reverse, and Perplexity leans more on community sources.
Before the table, a caveat: the majority of what works is shared across every platform. Structure, factual density, and third-party corroboration move all of them. Platform-specific tactics are worth maybe the last fifth of your effort, and they are the wrong place to start.
The differences are real, though. Profound ran 100,000 prompts across both ChatGPT and Perplexity and found only 11% of cited domains overlapped between the two, with 37.4% cited exclusively by ChatGPT and 51.6% exclusively by Perplexity.
The source mixes below come from Profound’s separate 680-million-citation dataset unless noted.
| Platform | Source mix | Citation and mention behaviour | Practical implication |
|---|---|---|---|
| Google AI Overviews | Reddit around 21% of top sources. Roughly 43% of citations point to Google-owned properties, including YouTube | Only 37.9% of citations come from top ten organic results (Ahrefs data) | Existing SEO carries over more here than anywhere else, but 2 out of 3 citations come from outside page one |
| ChatGPT | Wikipedia dominant at 47.9% of top sources, Reddit far lower at 11.3% | Cites in 87% of brand appearances, names the brand in only 20.7% (Semrush) | Encyclopaedic and third party validation beat your own blog. Expect ghost citations |
| Perplexity | The most community-skewed engine. Reddit around 46.7% of top source share | Profound finds review aggregators such as G2 and Gartner prominent on commercial queries | Narrow, current, specific pages and strong review profiles do the work |
| Gemini | Blogs, news, and video, with YouTube mentions the strongest single correlate Ahrefs measured | Names brands 83.7% of the time but cites them only 21.4% (Semrush) | The inverse of ChatGPT. Optimise for being named, not for the link |
| Claude | Training data plus live retrieval, with a smaller published dataset than the others | Less independently measured than the other four | Not enough platform-specific guidance for a recommendation |
Platform-specific work becomes worth doing when your audience concentrates on one engine. A developer tooling company whose buyers live in Claude should be weighted differently than a consumer brand whose category gets researched in Gemini. Until you know that from your own B1 data, optimise for the shared fundamentals.
AI Citation Tactics That Don’t Work
The most common AEO mistakes are not errors of effort but of direction. Schema, publishing volume, cosmetic date changes, and single spot checks all consume time without moving citation rates.
Here are some tactics that have been proven not to work (at least not on their own):
- Treating schema as a citation lever: Cited pages correlate with structured data, so people assume adding it drives citations. The only controlled test we have found no lift on any platform. Implement it for clarity, then spend your effort elsewhere.
- Publishing more: In the Ahrefs correlation study, content volume was among the weakest signals measured, at roughly 0.194 against 0.664 for brand mentions. One well-structured page with original data will do more than a run of thin posts, and the thin posts cost you time you needed for B3.
- Optimising as if AI Overviews were featured snippets: Snippets have one winner. AI answers assemble from several sources at once, and query fan-out means a single answer may draw on several parallel retrievals, so chasing a single position misses how it works.
- Changing dates without changing content: Freshness appears to be assessed from the content itself rather than a timestamp, so this fools nobody and wastes a deploy.
- Publishing unedited AI-generated content: It tends to perform briefly and then decay, and it contains nothing a model cannot already produce. If an assistant can generate your answer without you, it has no reason to cite you. An editorial QA workflow that checks claims against primary sources and strips generic phrasing is what separates the two.
- Spot-checking prompts by hand: The most common version of this is a founder typing the company name into ChatGPT and declaring the situation fine or catastrophic. With identical prompts returning the same list under 1% of the time, a handful of manual runs tells you nothing you can act on, and worse, it feels like evidence.
- Skipping Diagnosis: In AEO, there’s a lot of emphasis on new KPIs and tactics (vs. SEO) but not on how this new data should be analyzed and inform strategy. This is what B2 of my proprietary framework focuses on.
Conclusion
Benchmark what is happening now. Create a blueprint by breaking down why it is happening and deciding what deserves your attention first. Build the content and the outside coverage that change it. Then measure again, because the answer will have moved.
The argument underneath the framework is simple enough. Citations go to sources that are genuinely worth citing and structurally easy to extract from. Neither half works alone. Perfect structure around content that adds nothing gets ignored, and genuine expertise buried in unextractable prose gets skipped for something worse that was easier to read.
If you do one thing today, write down the ten attributes you most want to be known for, in your buyers’ language rather than your product team’s. That list is what B1 measures against; it costs nothing, and most teams discover in the writing that they have never agreed on it.
FAQs
How do I earn AI citations?
Run three stages in a loop. Benchmark your current visibility using a tracking platform that samples a broad prompt set across the major AI engines repeatedly, since single checks are too unreliable to act on. Break down the results to identify which of the three failure modes you have and what to fix first. Build the on-site structure that makes you extractable and the third-party coverage that gets you selected. Then measure again and repeat.
How long does it take to get cited by AI?
Structural changes to existing pages tend to show up within roughly four to eight weeks as engines recrawl. Third-party authority building typically takes three to six months to move mention rate measurably. Expect variance either way, since AI responses fluctuate between identical runs.
How do I get cited in AI search results if my site has low authority?
Better than you might expect. An AEO study found structural improvements produced larger visibility gains for lower-ranked sources than for established ones, with one analysis putting the gain from adding citations at 115% for lower-ranked content. Focus on direct answer openers, original data, and specificity on narrow topics where you can credibly claim expertise, rather than competing on breadth.
What’s the difference between a citation and a mention in AI answers?
A citation links your domain as a source. A mention names your brand in the answer text. Semrush and Kevin Indig found the two overlap in only 13.2% of brand appearances, with 61.7% cited but unnamed and 25.1% named but uncited. Citations are easier to measure since they can drive traffic. Mentions influence buying decisions more directly.
Do I need to rank in Google to be cited by AI?
No, though it helps for Google’s own AI surfaces. Ahrefs found only 37.9% of AI Overview citations come from pages in the top ten, down from around 76% a year earlier, with the remainder split between positions 11 to 100 and pages outside the top 100 entirely. The gap is wider outside Google, where separate Ahrefs research puts the overlap between AI assistant citations and the organic top ten at around 12%. Strong rankings improve your odds but do not determine the outcome, and I unpack that split further in How SEO Is Changing in the Age of AI Search.
Does schema markup improve AI citation rates?
Not on its own, based on the only controlled test available. Ahrefs tracked 1,885 pages adding JSON-LD against roughly 4,000 controls and found no statistically significant lift on ChatGPT or AI Mode, and a small significant decline on AI Overviews. The test measured pages that already had 100 or more citations, so it does not settle the case for pages starting from zero. Implement schema for clarity, then invest elsewhere.
Do I need paid tools to track AI citations?
Realistically, yes. SparkToro and Gumshoe found identical prompts return the same brand list under 1% of the time, so a hand-collected sample of twenty or thirty answers is unreliable. Reliable measurement means sampling many prompts many times across engines, which is a volume problem tooling exists to solve. Options range widely in price, and the tool matters less than the sampling volume.
What’s the difference between AEO and GEO?
Answer Engine Optimisation (AEO) generally refers to being extracted as a direct answer, while Generative Engine Optimisation (GEO) refers more broadly to being cited across AI-generated responses. In practice, most practitioners use the terms interchangeably, and the tactics overlap almost entirely. The distinction matters far less than whether your content is extractable and corroborated. I make the fuller argument, including where AIO and LLMO fit, in AEO vs. GEO vs. AIO: Are They Actually Different?


