AI visibility requires being retrievable, worth citing, and corroborated by third-party sources. The 3Bs Framework runs AEO strategy and execution as a flywheel. The 3Bs stand for benchmarking where you stand, blueprinting a prioritized strategy, then building what closes the gaps.

A common problem my clients face is being ranked 1-5 on Google for their main category term but rarely being cited or mentioned on ChatGPT and other answer engines.
Ahrefs data shows that this isn’t a fluke; it’s not just impacting the handful of clients I work with. In March 2026, they analysed 863,000 SERPs and found that only 37.9% of URLs cited in Google AI Overviews featured in the first ten results (down 76% YoY). The rest split almost evenly between positions 11 to 100 and even pages not ranking in the top 100 at all.
That tells me that rankings and citations have become two different outcomes. A page can do everything Google rewards and still get skipped by AI engines, so doing more of the same SEO work won’t solve the problem.
Over the last few months, I’ve been developing my process for earning AI citations and mentions, built specifically around how these engines select and use sources. I call it the 3Bs Framework (Benchmark, Blueprint, and Build). This article details my process for helping brands improve their AI visibility.
Summary
AI visibility takes a consistent loop of measuring, planning, and executing, not a one-time fix. The three stages are Benchmark, Blueprint, and Build.
- Benchmark: Define the attributes you want to be known for, then measure how often engines cite and name you across them using a tracking platform rather than manual checks.
- Blueprint: Diagnose why the gap exists and strategize before fixing anything. Being invisible, being cited without being named, and being described inaccurately are three different problems with three different solutions.
- Build on-site to become eligible for citation, then build off-site to actually get selected. Off-site work is the bigger needle-mover.
Why I Run AI Visibility as a Loop
Answer engines change their source preferences continuously, so a one-time audit goes stale quickly. Running Benchmark, Blueprint, and Build as a repeating cycle means each pass starts from better information than the last.
The three stages divide the work cleanly:
| Stage | Question It Answers | Output |
| Benchmark | Where do we actually stand? | A baseline you trust |
| Blueprint | Why, and what do we fix first? | A prioritized plan |
| Build | What closes the gap? | Shipped work, on-site and off |
It behaves like a flywheel because each turn makes the next one easier and more fruitful.
The prompt set only gets built once (though it can later be improved or expanded). The outreach list from Blueprint is already warm by the second cycle. And because branded mentions are the strongest correlate of AI visibility anyone has measured, the coverage you earn in Build incrementally raises your baseline.
The stages also shift emphasis over time. Early cycles lean on Build‘s on-site half, because engines need something extractable to work with. Later cycles lean off-site, since that’s where the compounding actually happens.
One thing worth mentioning is that SEO in the AI search era remains foundational. The 3Bs Framework introduces a new measurement layer, diagnostic logic, and tactics, but the underlying fundamentals that make you findable, relevant, and trusted haven’t changed.
B1. Benchmark: Measure Where You Stand
Benchmark means establishing a reliable baseline of how AI engines currently cite and describe you. Run-to-run variance makes hand-collected samples unreliable, so this stage depends on tooling rather than spot checks.
You can’t fix a gap you haven’t measured. But measuring AI visibility badly is still an issue since you’re using those figures to make decisions.
Set Up AI Traffic Tracking in GA4 and Search Console
GA4 catches AI traffic that clicks through. Search Console’s Generative AI report catches impressions that don’t. Neither one alone shows the full picture.
In GA4, I add an AI channel group filtering for the engines that matter to the brand, so AI referral traffic separates from the rest. Over time, this tells you which platforms send people to your site, which then informs how you set up your prompt tracking software.
Search Console’s Generative AI performance report covers what GA4 structurally can’t. It shows impressions inside AI Overviews and AI Mode, broken down by page, country, and device. There’s no click data, no CTR, and no query-level detail so far, but it’s the only place that shows zero-click AI Overview visibility directly.

The reason to run both is that a page can appear inside AI answers constantly while sending you nothing. If you only track referrals, that page looks like a failure when it’s actually working.
Check Whether Engines Can Reach and Extract Your Pages
Before anything else, confirm engines can fetch your pages and pull clean passages out of them. This is a one-time setup check, unlike the remaining sections in B1, which require continuous measurement and/or tweaking.
In my guide on how to get cited by AIs, I explain this in detail, including how to fix what you find, but here’s a summarized version. Start by testing crawler access and rendering with Screaming Frog, configured specifically for this since it doesn’t run these checks by default.
➡️ For access, set a custom User-Agent under Configuration > HTTP Header and crawl as GPTBot, ClaudeBot, or PerplexityBot. A 403, 429, or 503 in the Response Codes tab means your CDN or WAF is blocking that bot, even though a browser sees the page fine.
➡️ For rendering, switch Configuration > Spider > Rendering to JavaScript and compare the Word Count Difference between raw and rendered HTML. A large gap means crawlers are seeing an empty page, since they read raw HTML rather than rendering JavaScript. Also check whether content sits behind a login, a form, a cookie wall, or is locked inside a PDF.
But being fetchable doesn’t guarantee selection. Crawlers weigh a page’s cover, its URL, title, snippet, and freshness, before deciding whether opening it is worth the computational resources, the same “chosen” gate I describe in my breakdown of AEO, GEO, AIO, and LLMO.

Deep, specific pages get cited. Homepages and broad category hubs almost never do, and Ahrefs analyzed 17 million citations and found that 76.4% of ChatGPT’s top-cited pages had been updated within the previous 30 days.
Finally, check entity consistency. Do your site, LinkedIn, G2, and Crunchbase (or any other review platform or directory) describe you in the same category, using the same language? Inconsistency means ambiguity for AI systems.
Select Attributes You Want to Be Known For, Then Write Prompts
Attributes are the unit you measure against. Prompts are just instruments for reading them. I start by defining category terms, verticals, use cases, and capabilities, then build prompts to test each one.
Working attribute-first stops you from optimising for phrasing instead of positioning. Ten well-chosen attributes beat a hundred scattered prompts.
Once the attributes exist, I build a spread of prompts per attribute rather than one canonical phrasing, because phrasing moves results violently. Change one word in a category label, and visibility can swing by tens of percentage points across the same tools in the same fortnight.
I structure prompts in two tiers:
- Coverage prompts are broad and commercial. They monitor whether you’re in the conversation for a category. Most of your set should be these.
- Depth prompts are compound, stacking modifiers onto the base term. An audience or ICP, an evaluation criterion, a budget constraint, a named comparison. These answer whether you’re the default recommendation once a buyer gets specific.
A coverage prompt like “Best RPC provider” is a different test from depth prompts like “best RPC provider for a DeFi protocol on Arbitrum” or “best RPC provider for a DeFi protocol needing sub-second finality at 500+ requests per second.”
I also cover three prompt types deliberately:
- Discovery prompts test whether you make the list. E.g.: “best RPC providers”
- Comparison prompts test whether you win a head-to-head. E.g.: “[brand] vs. Alchemy for high throughput”
- Validation prompts ask a direct yes-or-no about whether a capability is yours. E.g.: “Is [brand] an RPC provider?” or “Does [brand] support Arbitrum?”
That third type looks redundant, but you’ll understand its importance when you reach B2, where it does specific diagnostic work. Avoid brand names in discovery prompts, since including your own name skews the picture toward your own pages.
Set Up Your Prompt Tracking Platform (Don’t Run Prompts by Hand)
Identical prompts return the same brand list less than 1% of the time. A handful of manual checks produces a number that feels like evidence but is statistically irrelevant and can’t reliably inform decision-making.

This is where I disagree with a lot of advice on this topic, which tells you to run twenty or thirty prompts by hand and log the results. SparkToro and Gumshoe tested this directly, running identical prompts thousands of times, and found the same list came back under 1% of the time.
🎯 One prompt, on one platform, at one moment is a data point, not a read. A team that has been spot-checking for months has built a narrative, not a baseline.
What you need is something sampling at a volume no person can replicate. Ahrefs Brand Radar, Semrush’s AI toolkit, Profound, and Scrunch all do a version of this at different price points. The specific tool matters far less than the sampling volume behind it.
Still, there’s a setup decision that matters more than the tool choice. You should filter to the platforms that actually matter for your audience before you read anything. Including low-usage engines dilutes every number you look at.
The second important setup decision is tagging prompts by intent from the start. Informational and commercial prompts behave so differently that blending them into one aggregate score means the numbers aren’t painting the most accurate picture.
Build a Recurring Report Combining Platform, GA4, and Search Console Data
Every loop starts by pulling the same numbers into one place. The tracking platform tells you how AI engines see you. GA4 and Search Console tell you whether that visibility is worth anything.
Each source performs a different job:
| Source | What to pull | What it answers (in B2) |
| AI visibility platform | Visibility score and rank per attribute, citation share, mention rate, cited domains, all split by engine | Which gap you have, who’s beating you, and which third-party sources to target |
| GA4 | AI referral sessions by engine, landing pages receiving them, conversions, and your Direct traffic trend | Whether visibility is producing anything commercially |
| GSC | Generative AI impressions by page and branded searches | Where you appear in AI Overviews without earning a click, and whether AI exposure is driving people to search for you by name |
Regarding AI visibility, you should report at two levels:
- Roll up by attribute (i.e., topic or category), since that’s what you’re trying to win and what your gap list in B2 is organised around.
- But keep the prompt-level rows underneath, because that’s where decisions get made.
An attribute average will often hide one prompt growing while another decreases. The category data tells you where to look; the prompt tells you what happened. Put the previous period’s numbers beside the current ones at both levels, so movement is visible, and/or get percentage changes directly from the tracking tool for your desired comparison period.
Then, use GA4 and GSC to track Direct traffic, branded search volume, AI referrals, and generative AI impressions.

I look at these four metrics to try to piece together a somewhat accurate view of the traffic AI engines are generating because a large share of AI mentions never carry as a referrer. In fact, Kyle Poyar found 14.3% of his new subscribers self-reported first hearing about him through an LLM, roughly 10x what his click-based analytics showed.
- Someone who reads about you in ChatGPT and types your URL shows up as Direct.
- Someone who goes on to Google your name shows up in Search Console as impressions and clicks for branded terms.
Branded search is the cleaner of the two signals, since Direct also collects bookmarks and repeat visitors.
Automate the data pull if your platform supports scheduled exports. The weekly version only needs the deltas so you can spot large movements. The quarterly version is where you review the full picture, including the branded search trend, which typically moves too slowly to read week to week.
B2. Blueprint: Diagnose the Gaps and Decide What to Prioritize
Blueprint means working out why the gap exists, then turning that diagnosis into a prioritised plan. This is the stage most AEO advice skips entirely.
In this space, there’s a lot of emphasis on new AI search KPIs and new tactics. There’s very little on how the data should be read or how it should inform brand visibility strategy in AI engines. That’s the gap this stage fills. The goal is to ensure your AEO efforts target the real problems you face with effective solutions.
🗝️ Before we move further, know that one rule governs everything. Never diagnose from an aggregate. Averages hide opposing trends on two axes: one prompt decreasing while another climbs, and one AI engine performing well while another doesn’t. Both blend into a flat-looking number that tells you nothing about where to look.
Classify Which of the Three Visibility Gaps You Have
There are three types of visibility gaps, and each needs a different response:
| Gap | What You See | What to Fix |
| Invisible | Neither cited nor mentioned | Content or eligibility |
| Ghost Citations | Domain cited, brand never named | Positioning and comparative content |
| Misrepresented | Named, but inaccurately or unfavourably | Clarity and authority signals |
Once you’ve diagnosed the gap, you can start thinking about how to fix it and add it to your action plan to complete in B3.
According to Semrush and Kevin Indig’s study, comparative queries named brands 43.3% of the time, against 18% for informational queries, roughly 2.4x more often. So, if you’re getting cited but not mentioned (ghost citations), creating comparative content on your website can help. Off-site, aim to get named explicitly inside third-party comparisons.
Whether you decide to focus on owned, earned, or paid comparative content, ensure your brand is explicitly named next to the attributes you want to be known for, so AI engines associate the two.
Write “[Brand] delivers sub-second finality on Arbitrum at 500+ requests per second,” not “our infrastructure delivers sub-second finality on […].” The second version is citable, but the engine might not attach your name to the attributes, meaning your brand is less likely to get mentioned.
Meanwhile, if your brand is being misrepresented, it could be due to an outdated or biased source driving the narrative. In this case, create new content somewhere authoritative enough to outweigh it. If several independent sources are making similar claims, it’s likely a product problem, and no amount of writing can overcome accurate criticism.
Lastly, if you’re invisible, two actions can help, but we need to drill deeper into the data to know which one will be effective for your particular case. That’s what I’ll discuss in the next section.
Use Validation Prompts to Separate a Trust Problem From a Comprehension Problem
If you’re invisible for an attribute, a direct validation prompt tells you whether the engine doesn’t know about your capability or knows and doesn’t recommend you. Those need completely different work.
This is where the prompt types from B1 become key. Ask the engine directly: “Does [brand] do X?”
- If it says yes, the engine knows you have the capability and simply doesn’t recommend you. That’s a trust and amplification problem, and no amount of on-page content will touch it. The work is off-site.
- If it doesn’t know, you have a comprehension problem. The fix is on-site, in clearer descriptions, documentation, and positioning language.
Without validation prompts, both cases look identical in your dashboard. You’d see “not visible” and reasonably conclude you need more content, which could be the wrong call.
Read Visibility Score Against Rank to Judge Your Competitive Position
Visibility score is how consistently you appear. Rank is where you sit against competitors for that category or prompt. Rank is arguably the more reliable read, because a healthy-looking visibility score can still mean fifth place or lower.
That’s why it helps to sort by rank first (if your tracking platform supports it) and treat score as supporting context. Now, I said “arguably” above because the statistical reliability of the rank value depends on how many times the tool ran the prompt, given how much AI responses vary between runs.
Assuming the rank value is reliable, its combination with the visibility score tells you something neither number does alone:
- High rank, low score means that models put you first when they think of you, but haven’t settled on you yet. To turn occasional into reliable selection, strengthen third-party mentions and consistent entity signals.
- Low rank, decent score means you’re consistently included and consistently beaten. Look at the sources the engine cites for the prompt you’re analyzing to find out whether net new content, an update, or outreach is the best strategy.
- Rank moving sharply means that something changed. Large movement is worth investigating before small movement, regardless of direction. Check for a new source that recently entered the source list and other changes.
Sort Cited Sources Into Competitors, Outreach Targets, and Institutions
The sources cited instead of you fall into three groups that need three different responses. Treating them as a single list of competitors is a common and expensive mistake.
Scope this analysis to discovery prompts, where no brand name appears in the question. Once your name is in the prompt, engines pull more heavily from your own pages, and the picture skews toward owned sources.
1. Actual Competitors
Brands contesting the same commercial ground. Track rank against your two or three closest rivals per attribute. These are who you measure share of voice against.
If competitor content dominates a prompt’s citation list, that’s good news for you. Engines are treating brand-owned pages as satisfactory sources for that prompt, so your own content has a real chance of breaking in. Create new content or update what you have.
2. Earned Media, Publishers, Review Sites, and Affiliates
This list is actionable. If these dominate an attribute’s citation mix, that’s your outreach target list. Prioritise domains with rising citation share, since a trending source is one engines are trusting more. Check whether a source is independent or a competitor’s own placement before pitching.
3. Institutional Sources
These include Wikipedia, government sites, and major reference bodies. Not actionable in the same way. If these dominate, owned content faces a steep climb no matter what you publish.
If these dominate, it usually means the query wants a strictly informational answer rather than a vendor’s take, so you’re better off deprioritising that attribute and putting the effort where brand content has a route in.
4. Forums or Social Media Platforms
If an engine cites Reddit or a similar platform for a topic you should own, it usually means no brand has published a good answer yet. That’s an opening.
💡 Alongside who gets cited, check what format wins. If every top-cited page for a prompt is a listicle, engines have decided listicles are the citable format for that question. Match the shape of what’s already being fetched.
Build a Gap List With Causes and Fixes Before You Prioritize
A complete diagnosis isn’t “visibility went down.” It names what dropped, where it dropped, why, and what the fix options are. Answer the following four questions in order.
1. What dropped?
Go to the prompt level, since a topic average blends everything together. Trend direction is what matters most here:
| What you see | What it means | What to do |
| Increasing prompt | Momentum | Find what’s working and accelerate it |
| Decreasing prompt | Emergency | Investigate separately and urgently |
2. Where did it drop?
Check whether the decline is uniform across engines or split, since three platforms down and one up blends into a moderate-looking decline that misrepresents both.
3. Why did it drop?
Reading the prompt text usually generates the hypothesis. Say the three slipping prompts are “cheapest RPC provider for high-volume dApps,” “most cost-effective Arbitrum RPC at scale,” and “RPC providers with predictable pricing.”
The angle is pricing, not RPC generally, so competitors are winning the cost conversation rather than you having a broad visibility problem. The most common smoking gun is a new source that recently surged in citation share.
4. What to do about it?
Analyze the sources that the AI engine prefers and decide whether to:
- Create new content built for that prompt
- Update an existing page that’s underperforming
- Pitch the publishers already being cited
- Pursue a paid or affiliate placement
- Fix an entity or positioning inconsistency
- Deprioritise the attribute because it isn’t winnable.
More than one action can apply to the same gap.
Decide What to Fix First Based on Value, Gap Size, and Winnability
I score opportunities on commercial value, how large the gap is, and whether we can credibly claim authority on the topic. Where you start depends on where you already are.
The prioritisation logic transfers directly from traditional SEO. It’s the same scoring I use when building a repeatable SEO strategy system. The starting point is what differs:
- If you lack a reliable baseline, finish the setup described in B1 before touching anything.
- If you have thin or badly structured content, fix the foundations first. Off-site authority can’t rescue a page nobody can extract from.
- If you have strong existing SEO and solid content, go straight to off-site work. Your pages are already extractable, so the constraint is corroboration.
At minimum, the working rhythm I recommend is one prompt, one fix, once a week. AEO rewards consistent iteration far more than it rewards large infrequent actions, partly because the ground keeps moving.
B3. Build: Create the Content and Authority That Earn Citations
Build has two halves plus maintenance. On-site work makes you eligible to be cited. Off-site work is what actually gets you selected and mentioned. Most teams do the first and stop.
Build On-Site: Make Pages Extractable and Worth Quoting
On-site work is about being extractable. Lead with direct answers, write headings as questions, keep passages self-contained, and add data (preferably data that nobody else has).
Here are the on-page practices that matter most to earn AI citations and mentions:
- Write headings as the questions readers actually ask.
- Make sections standalone. If a passage stops making sense once removed from the page, it won’t survive extraction.
- Use tables, numbered lists, and short paragraphs carrying one idea each.
- Publish original data and/or back your claims with data from reputable sources.
- Get content out from behind JavaScript, accordions, and read-more toggles.
- Answer in the opening 40 to 60 words of each section.
Engines retrieve passages, not whole pages, so each section has to resolve on its own. Indig’s analysis of 18,012 ChatGPT citations found that cited passages cluster where a question-format heading is answered directly beneath it, so write headings as the questions readers ask and answer them immediately.
On schema, I’ll be honest about the evidence. Ahrefs tracked 1,885 pages adding JSON-LD against matched controls and found no positive effect on any platform. It makes your content unambiguous to machines, which is worth doing, but it isn’t a growth hack.
Decide Whether to Optimise an Existing Page or Create a New One
Optimise when a page covers the right topic but has structural problems. Create something new when your existing page is too broad for the specific prompt winning citations.
More specifically, optimise the existing page when:
- It covers the right topic but buries the answer
- It tries to serve two purposes at once, like being both a fair comparison and a product showcase
- The information is all there, but its structure prevents an engine from extracting it easily
I want to highlight the second case because it’s common and often overlooked. For example, a comparison article that gives one product dramatically more space, or shifts from editorial to promotional language partway through, reads as an untrustworthy source for the question it should be answering. Engines have cleaner comparison pages available and will use those.
You should create something new when:
- Your page is too broad for the prompt that’s winning
- No top-cited page answers the specific question directly
- Nothing on your site covers that attribute
A broad “best RPC providers” page can’t double as a precise “best Arbitrum RPC providers” page. The supply chain rewards specificity, so build for the question rather than bolting a subsection onto something general.
Build Off-Site: Earn the Coverage That Gets You Selected
Off-site work is the most effective strategy. The same article earned roughly 325% more citations when distributed to third-party sites instead of living only on the brand’s own domain.
That figure comes from a controlled test by Stacker and Scrunch, which took eight articles, ran them in two conditions, and measured 944 prompt and platform combinations across five LLMs. In roughly one in five cases, the engine cited the syndicated version and never cited the brand’s original.
Your website makes you eligible. Other people talking about you is what gets you mentioned. In practice, that means you should:
- Claim and maintain category review profiles on G2, Capterra, and Trustpilot.
- Earn coverage in publications your buyers read, prioritising relevance over raw domain authority.
- Participate honestly in the communities where your category gets discussed.
- Go multimodal, particularly on YouTube, which Ahrefs found to be the strongest individual signal it measured.
- Clean up entity descriptions across every profile you control.
Approach PR with the citation mechanic in mind. When you give a quote, include the exact positioning language you want engines to repeat, because a vague quote gives them nothing usable.
On backlinks, practitioners genuinely disagree. Ahrefs’ correlation data puts backlinks well below unlinked brand mentions. My own view is that links still carry SEO value, and traditional rankings still feed Google’s AI surfaces, so they aren’t irrelevant. But with limited hours, an unlinked editorial mention in a publication your buyers read beats a link from a site nobody in your category has heard of.
Review Priority Pages at Least Quarterly; Refresh Only What’s Dropping
Freshness might matter more here than in classic SEO, but a fixed rewrite schedule is the wrong instrument. Put priority pages at least on a quarterly review, then rewrite the ones losing citations and/or mentions.
The distinction that matters is that the cadence governs the review, not the rewrite. AirOps found pages not updated quarterly are around three times more likely to lose citations, which is a good reason to look every quarter (or more frequently if you can). It isn’t a reason to rewrite everything you look at.
The two rules I hold to are:
- Leave pages that are still earning citations alone. Judge by performance, not age, as rewriting it risks losing what’s working.
- Don’t change dates without changing content. Freshness appears to be assessed from the content itself, so this fools nobody.
Scope the review to your priority attributes rather than your whole library. Trying to review everything quarterly is how teams end up making cosmetic updates.
How the 3Bs Loop Closes and Compounds
Each cycle starts from a stronger position than the last. Carry citation rate, mention rate, and share of voice forward, then run the stages again with better information.
The three numbers I carry between cycles are:
- Citation rate: How often your domain appears as a source.
- Mention rate: How often your brand is named in the answer.
- Share-of-voice (SoV): Your portion of total citations and mentions against competitors.
Track the pair rather than either alone. AirOps found brands earning both a mention and a citation are 40% more likely to reappear in the next answer than citation-only brands, so healthy mix matters.
Meanwhile, two cadences run in parallel, doing different jobs:
| Cadence | What it’s for |
| Weekly | Scan for large movements, update the gap list, and ship at least one fix, one topic at a time |
| Quarterly | Reassess attributes and priorities, review priority pages |
Weekly works, and I prefer it, because these engines change constantly and a tight loop teaches you faster. But only if you act on large movements and ignore small ones. Small weekly swings are usually variance.
The compounding is the point. Your second cycle skips prompt set construction, and the publication list from Blueprint is already built and partly warm. And because branded mentions are the strongest correlate of AI visibility anyone has measured, coverage earned in Build raises your baseline, including on prompts you never targeted.
Citations vs. Mentions: Why the Distinction Changes Your Strategy
A citation links your domain as a source. A mention names your brand in the answer text. They happen independently, and confusing them sends you after the wrong fix.
Commercially, they do different things. A citation is measurable, since it can send a click you’ll see in analytics. A mention is what actually influences a buyer, because it’s your name appearing inside a recommendation.
The Semrush and Indig’s analysis of 3,981 brand appearances quantifies how loosely coupled the two are:
- 61.7% were citations with no mention
- 25.1% were mentions with no citation
- Only 13.2% were both
You can be feeding answers across your entire category, generating no recognition whatsoever, and conclude from a citation report that things are going well.
That gap is also a strategy signal. Informational content earns citations far more readily than it earns mentions. Comparative content is where brands actually get named. For most brands, the answer leans toward comparative.
There’s a related diagnostic worth running. If your visibility is healthy but your citation share is low, third-party content is carrying your presence rather than your own pages. Whether that’s a problem depends on your category. A brand that wouldn’t naturally publish “best in category” guides has no reason to expect otherwise. A company with active blogs, guides, and product docs showing the same gap has a content performance problem.
My framework optimises for both, but treats mentions as the harder and more valuable outcome.
AI Visibility Tactics That Don’t Work
The most common AEO mistakes are not errors of effort but of direction. Schema, publishing volume, cosmetic date changes, and single spot checks all consume time without moving citation rates.
Here are some tactics that have been proven not to work (at least not on their own):
- Treating schema as a citation driver: Cited pages correlate with structured data, so people assume adding it drives citations. The only controlled test we have found no positive effect on any platform. Implement it for clarity, then spend your effort elsewhere.
- Publishing more: In the Ahrefs correlation study, content volume was among the weakest signals measured, at roughly 0.194 against 0.664 for brand mentions. One well-structured page with original data will do more than a run of thin posts, and the thin posts cost you time you needed for B3.
- Optimising as if AI Overviews were featured snippets: Snippets have one winner. AI answers assemble from several sources at once, and query fan-out means a single answer may draw on several parallel retrievals, so chasing a single position misses how it works.
- Changing dates without changing content: Freshness appears to be assessed from the content itself rather than a timestamp, so this fools nobody and wastes a deploy.
- Publishing unedited AI-generated content: It tends to perform briefly and then decay, and it contains nothing a model cannot already produce. If an assistant can generate your answer without you, it has no reason to cite you. An editorial QA workflow that checks claims against primary sources and strips generic phrasing is what separates the two.
- Spot-checking prompts by hand: The most common version of this is a founder typing the company name into ChatGPT and declaring the situation fine or catastrophic. With identical prompts returning the same list under 1% of the time, a handful of manual runs tells you nothing you can act on, and worse, it feels like evidence.
- Skipping Diagnosis: In AEO, there’s a lot of emphasis on new KPIs and tactics (vs. SEO) but not on how this new data should be analyzed and inform strategy. This is what B2 of my proprietary framework focuses on.
Conclusion: Start by Writing Down Your Top Attributes
Benchmark what’s happening. Blueprint why it’s happening and what deserves attention first. Build the content and coverage that change it. Then keep repeating the loop to compound its effects.
The argument underneath the framework is simple. Citations go to sources that are worth citing and structurally easy to extract from. Neither half works alone. Perfect structure around content that adds nothing gets ignored, and real expertise buried in unextractable prose gets skipped for something worse that was easier to read.
If you do one thing today, write down the top attributes you most want to be known for, in your buyers’ language rather than your product team’s. That list is what Benchmark measures against, and that shapes the whole process.
If you’d rather run this with help, that’s what I do. Get in touch if you need help improving your brand’s AI visibility.
FAQs
What is the 3Bs Framework?
The 3Bs Framework is my methodology for improving AI visibility, run as a repeating loop rather than a one-time fix. Benchmark measures where you stand across AI engines, Blueprint diagnoses the gaps and what to prioritize, and Build creates the work that closes them.
How often should you check your AI visibility?
Weekly and quarterly, doing different jobs. Weekly, scan for large movements, update your gap list, and ship at least one fix. Quarterly, reassess attributes and priorities and review priority pages. Ignore small weekly swings, since AI responses vary enough that minor movement is usually insignificant.
Why is my brand cited by AI but never mentioned by name?
That’s known as a ghost citation, which is typically a positioning problem rather than a content problem. Your domain feeds the answer while your brand name isn’t attached to the fact being extracted. Comparative content helps, on your site and in third-party comparisons, with your brand named beside the attribute.
How many prompts do I need to track AI visibility accurately?
Enough to cover every attribute you want to be known for, with several phrasings each, rather than a fixed number. Ten well-chosen attributes beat a hundred scattered prompts.
How do I track ChatGPT traffic in Google Analytics?
Create an AI channel group in GA4 filtering for the engines that matter to your brand, so AI referral traffic separates from everything else. Pair it with Search Console’s Generative AI report, which shows AI Overviews impressions that never produce a click.
Do backlinks still matter for AI search visibility?
Less than they do for traditional SEO. Ahrefs’ correlation data puts backlinks well below unlinked brand mentions. My view is that links still carry SEO value and rankings still feed Google’s AI surfaces, but a relevant unlinked mention beats a link from an irrelevant site.
How do I find out why my AI visibility dropped?
Work through four questions in order. What dropped, at the prompt level rather than the topic average. Where it dropped, by AI engine. Why it dropped, usually visible in the growth/decline of cited sources. Then what to do about it.



