<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[Windgrove's Blog]]></title><description><![CDATA[Expert perspectives on AEO strategy, AI search optimization, and what it takes to get your brand recommended by ChatGPT, Perplexity, Gemini, and Claude.]]></description><link>https://windgrove.hashnode.dev</link><image><url>https://cdn.hashnode.com/uploads/logos/6a6513b3e2d2908b72dc5cf1/8e3a99a2-afeb-4d85-8943-23d04a36ff1d.png</url><title>Windgrove&apos;s Blog</title><link>https://windgrove.hashnode.dev</link></image><generator>RSS for Node</generator><lastBuildDate>Sun, 20 Sep 2026 09:55:27 GMT</lastBuildDate><atom:link href="https://windgrove.hashnode.dev/rss.xml" rel="self" type="application/rss+xml"/><language><![CDATA[en]]></language><ttl>60</ttl><item><title><![CDATA[AEO vs SEO: Which Should B2B Companies Invest In?]]></title><description><![CDATA[TL;DR
SEO and AEO are not competing budget lines. SEO keeps you visible in traditional search. AEO makes you citable inside AI-generated answers from ChatGPT, Perplexity, Google AI Overviews, and Gemi]]></description><link>https://windgrove.hashnode.dev/aeo-vs-seo-investment</link><guid isPermaLink="true">https://windgrove.hashnode.dev/aeo-vs-seo-investment</guid><dc:creator><![CDATA[Spencer Duke]]></dc:creator><pubDate>Fri, 11 Sep 2026 17:30:00 GMT</pubDate><content:encoded><![CDATA[<h2>TL;DR</h2>
<p>SEO and AEO are not competing budget lines. SEO keeps you visible in traditional search. AEO makes you citable inside AI-generated answers from ChatGPT, Perplexity, Google AI Overviews, and Gemini. For most B2B companies in 2026, the right answer is both — sequenced correctly, with budget weighted toward AEO if your buyers are already using AI tools to research vendors.</p>
<p>The short version of the decision matrix:</p>
<p><strong>B2B SaaS:</strong> Weight toward AEO now. Your buyers use AI to shortlist vendors before they ever visit your site.</p>
<p><strong>Fintech:</strong> Balanced investment. Regulatory context and trust signals matter equally in both channels.</p>
<p><strong>Health tech:</strong> AEO first for discovery. SEO for compliance-sensitive content.</p>
<p><strong>Professional services:</strong> AEO is the new word-of-mouth. Prioritise it.</p>
<hr />
<p>Here is the uncomfortable reality for most B2B marketing leaders: your SEO is probably working. Your content ranks. Your domain authority is solid. And you are still invisible to a growing share of your buyers — because they never opened a search results page.</p>
<p><a href="https://www.gartner.com/en/articles/gartner-predicts-25-of-enterprise-search-volume-will-shift-to-ai-chatbots-by-2026">Gartner projects</a> that 25% of enterprise search volume will shift to AI chatbots by 2026. <a href="https://sparktoro.com/blog/we-analyzed-332-million-google-searches-heres-what-we-learned-about-zero-click-search-in-2024/">SparkToro and Datos research</a> found that nearly 60% of Google searches already end without a click. The traffic model that SEO was built on is eroding from two directions at once: zero-click search on one side, AI-generated answers on the other.</p>
<p>This is not a prediction. It is the current state of B2B discovery.</p>
<p>The question revenue and marketing leaders are now asking is the right one: where does the next pound of budget go? Into the channel that has been working for a decade, or into the channel where your buyers are increasingly making decisions?</p>
<p>This guide gives you a direct answer. No hype. No vague frameworks. A clear definition of each discipline, what each one does well, and a decision matrix you can apply to your specific vertical.</p>
<h2>Defining the Terms: SEO, AEO, and GEO</h2>
<p>These three terms get conflated constantly. They are not the same thing, and the differences matter when you are allocating budget.</p>
<h3>SEO (Search Engine Optimization)</h3>
<p>SEO is the practice of making your content visible in traditional search engine results pages (SERPs), primarily Google and Bing. It works through a combination of technical signals (site speed, crawlability, structured data), content signals (keyword relevance, topical authority, content depth), and authority signals (backlinks, domain reputation).</p>
<p><strong>What SEO produces:</strong> ranked pages. When someone searches a query, SEO determines whether your page appears in the list of results and how high it sits.</p>
<h3>AEO (Answer Engine Optimisation)</h3>
<p>AEO is the practice of making your brand citable inside AI-generated answers. The target engines are ChatGPT, Perplexity, Google AI Overviews, Gemini, and Bing Copilot. These engines do not rank pages in a list. They synthesise a single, confident answer and name the brands, sources, and solutions they consider authoritative.</p>
<p><strong>What AEO produces:</strong> citations. When someone asks an AI "what is the best spend management tool for agencies," AEO determines whether your brand gets named in the response.</p>
<p><strong>Key signals for AI citation:</strong> entity clarity (does the AI know exactly who you are and what you do), citation footprint (do third-party sources reference you), content structure (is your content written so an LLM can extract a direct answer), and platform consistency (is your brand described consistently across LinkedIn, directories, Reddit, and your own site).</p>
<h3>GEO (Generative Engine Optimisation)</h3>
<p>GEO is a newer term, used largely in academic and technical circles, that describes optimisation for generative AI outputs more broadly. In practice, GEO and AEO describe the same discipline. The distinction is mostly semantic. AEO is the term used by practitioners; GEO appears more often in research papers.</p>
<blockquote>
<p><strong>The core distinction:</strong> SEO earns you a place on a list. AEO earns you a place in an answer. Those are different surfaces, driven by different signals, requiring different execution.</p>
</blockquote>
<p>For B2B buyers who now begin their vendor research inside ChatGPT or Perplexity, the answer is the only surface that matters.</p>
<h2>What Traditional SEO Still Does Well</h2>
<p>Before making the case for AEO investment, it is worth being honest about what SEO still delivers. Abandoning SEO in favour of AEO is not the right move for most companies. Here is why SEO remains valuable.</p>
<h3>High-intent transactional queries</h3>
<p>When someone searches "buy project management software" or "HR platform pricing," they are ready to act. Traditional search results still capture this intent reliably. Google's SERP remains the dominant surface for transactional and navigational queries, and ranking well there converts.</p>
<h3>Long-tail informational content</h3>
<p>Long-tail SEO content builds topical authority over time. Articles that rank for specific, niche questions signal relevance to both Google and AI engines. That authority is also a signal that AI engines use when deciding which sources to cite. A strong SEO content library is not wasted; it feeds both channels.</p>
<h3>Compounding organic traffic</h3>
<p>A well-optimised page can drive traffic for years with minimal ongoing investment. That compounding return is one of SEO's most defensible advantages. AEO content can compound too, but the measurement model is newer and less predictable.</p>
<h3>Local and map pack visibility</h3>
<p>For professional services firms with physical locations, Google Maps and local pack rankings are still the primary discovery surface for location-based queries. AEO is less developed for local intent. SEO wins here.</p>
<h3>What SEO cannot do</h3>
<p>SEO cannot make you visible inside a ChatGPT answer. A page that ranks #1 on Google is not automatically cited by AI engines. The signals are different. A study by <a href="https://www.seerinteractive.com/">Seer Interactive</a> found that Google's top-ranked pages were cited in AI Overviews only a fraction of the time. Ranking and citation are not the same outcome.</p>
<p><strong>The bottom line:</strong> SEO is not broken. It is incomplete. It covers the channels it was built for. Those channels now represent a shrinking share of where B2B buyers start their research.</p>
<h2>What AEO Adds That SEO Cannot</h2>
<p>AEO is not a rebrand of SEO with a new acronym. The execution is meaningfully different, and so are the outcomes it produces.</p>
<h3>Entity clarity</h3>
<p>AI engines build a model of who you are based on how consistently your brand is described across the web. Your website, LinkedIn company page, Crunchbase profile, directory listings, and third-party mentions all feed this model. If your description is inconsistent across those surfaces, the AI's confidence in citing you drops.</p>
<p>SEO does not require this level of entity consistency. AEO depends on it.</p>
<h3>Citation-worthiness</h3>
<p>To be cited in an AI-generated answer, your brand needs external validation. An AI engine will not confidently recommend a company that only exists on its own website. Third-party sources, including industry directories, Reddit threads, review platforms like G2, and media mentions. These are the citation signals that give AI engines confidence to name you.</p>
<p>Building this citation footprint is a core AEO activity. It has no direct equivalent in traditional SEO.</p>
<h3>Answer-formatted content</h3>
<p>AI engines extract answers from content. They look for direct, structured responses to questions, ideally within the first 100 words of a page, supported by FAQ schema markup that tells the engine exactly where the question and answer are. Content written for Google, which often buries the answer in the body of a long article, does not get pulled as an AI citation.</p>
<p>Reformatting existing content for answer extraction is one of the highest-leverage AEO activities. It often produces citation improvements within weeks.</p>
<h3>Platform tracking and AI visibility measurement</h3>
<p>SEO is measured through Google Search Console, keyword rankings, and organic traffic. AEO requires a different measurement layer: tracking how often your brand appears across ChatGPT, Perplexity, Gemini, and Google AI Overviews for your target prompts.</p>
<p>Platforms like <a href="https://www.getsearchable.com/">Searchable</a> track AI visibility score, share of voice against competitors, and prompt-level citation data across all major AI engines. Without this measurement layer, you are running AEO blind.</p>
<h3>Demand capture at the AI layer</h3>
<p><a href="https://www.perplexity.ai/">Perplexity reported</a> that it was handling over 100 million queries per week as of early 2025. ChatGPT's user base crossed 400 million weekly active users in February 2025. These are not edge-case audiences. They are the buyers your pipeline depends on, and they are making shortlist decisions inside AI interfaces before your sales team ever gets a call.</p>
<p>AEO captures that demand. SEO alone does not.</p>
<h2>The Decision Matrix: AEO vs SEO by B2B Vertical</h2>
<p>The right budget split depends on your vertical, your buyers' research behaviour, and where you are in your current content maturity. This matrix gives you a starting point.</p>
<table>
<thead>
<tr>
<th>Vertical</th>
<th>SEO Priority</th>
<th>AEO Priority</th>
<th>Rationale</th>
</tr>
</thead>
<tbody><tr>
<td>B2B SaaS</td>
<td>Medium</td>
<td>High</td>
<td>Buyers use AI to shortlist vendors before opening a browser tab</td>
</tr>
<tr>
<td>Fintech</td>
<td>High</td>
<td>High</td>
<td>Regulatory trust signals matter in both channels equally</td>
</tr>
<tr>
<td>Health tech</td>
<td>Medium</td>
<td>High</td>
<td>AI discovery is the new referral. Compliance content stays SEO-led</td>
</tr>
<tr>
<td>Professional services</td>
<td>Low to Medium</td>
<td>High</td>
<td>AI recommendation is the new word-of-mouth for services</td>
</tr>
<tr>
<td>Technology (infrastructure/dev tools)</td>
<td>High</td>
<td>Medium</td>
<td>Developer audiences still use search. AI citation growing fast</td>
</tr>
</tbody></table>
<h3>B2B SaaS</h3>
<p>Your buyers are already using ChatGPT and Perplexity to research software categories. A <a href="https://www.forrester.com/blogs/ai-in-b2b-buying/">2024 Forrester report</a> found that the majority of B2B buyers now use AI tools during the purchase research phase, with adoption accelerating fastest among technology and SaaS buyers. When they ask "what is the best CRM for a 50-person sales team," they get a named list. If you are not on it, you do not exist in that buyer's consideration set.</p>
<p><strong>Recommended split:</strong> 40% SEO, 60% AEO. Maintain your SEO foundation. Accelerate AEO investment now, before your competitors do.</p>
<h3>Fintech</h3>
<p>Fintech buyers are sophisticated and trust-sensitive. They cross-reference AI answers with regulatory databases, review sites, and industry publications. Both channels matter. SEO keeps you visible in the compliance and technical query layer. AEO gets you cited in the broader "best fintech for X" category answers.</p>
<p><strong>Recommended split:</strong> 50% SEO, 50% AEO. Run them in parallel. Neither can be neglected.</p>
<h3>Health Tech</h3>
<p>Discovery in health tech is increasingly AI-driven, particularly for practice management software, patient engagement tools, and clinical workflow platforms. Compliance-sensitive content (anything touching HIPAA, clinical evidence, or regulatory claims) should remain SEO-led, where you control the page. Category-level discovery belongs to AEO.</p>
<p><strong>Recommended split:</strong> 35% SEO, 65% AEO. Prioritise AEO for discovery; keep SEO for compliance content.</p>
<h3>Professional Services</h3>
<p>Accountants, lawyers, consultants, and advisors live and die by trust and recommendation. AI engines are replacing the referral conversation for a growing share of buyers. When someone asks ChatGPT "who is the best M&amp;A lawyer in Toronto," the AI gives a confident answer. Being in that answer is the new word-of-mouth.</p>
<p><strong>Recommended split:</strong> 25% SEO, 75% AEO. This is the vertical where AEO ROI is most immediate.</p>
<h3>Technology (Infrastructure and Developer Tools)</h3>
<p>Developer audiences still use Google heavily. Stack Overflow, GitHub, and technical documentation dominate search. But AI-assisted development tools and LLM-native research workflows are shifting this fast. Invest in AEO now to build citation authority before the shift accelerates.</p>
<p><strong>Recommended split:</strong> 55% SEO, 45% AEO. Maintain SEO strength; begin AEO foundation building immediately.</p>
<h2>Why Integrated SEO + AEO Is Usually the Strongest Model</h2>
<p>The framing of "AEO vs SEO" is useful for budget conversations, but it is slightly misleading in practice. The strongest model is integrated. Here is why.</p>
<h3>SEO content feeds AEO citation</h3>
<p>Content that ranks well on Google tends to be authoritative, well-structured, and topically comprehensive. Those are the same qualities AI engines look for when deciding what to cite. A strong SEO content library is not wasted on AEO; it is a foundation to build from. The difference is in the last mile: reformatting that content for direct answer extraction, adding FAQ schema, and front-loading the answer.</p>
<h3>AEO signals reinforce SEO authority</h3>
<p>A robust citation footprint, the third-party mentions, directory listings, and Reddit presence that AEO requires, also builds the domain authority and backlink signals that SEO depends on. The two disciplines share an underlying logic: earn trust from external sources, and both engines will reward you.</p>
<h3>The sequencing matters more than the split</h3>
<p>For most B2B companies starting AEO from scratch, the right sequence is:</p>
<p><strong>Month 1:</strong> Fix the technical foundation. Ensure AI crawlers can read your site. Implement schema markup. Resolve entity inconsistencies across your web presence.</p>
<p><strong>Month 2:</strong> Reformat existing content for answer extraction. Build the citation footprint. Launch Reddit and directory presence.</p>
<p><strong>Month 3:</strong> Produce new AEO-native content mapped to your highest-value buyer prompts. Begin tracking AI visibility against competitors.</p>
<p>SEO runs in parallel throughout. You do not pause SEO to do AEO. You add AEO execution to an existing SEO foundation.</p>
<blockquote>
<p>The companies that will own AI search in their categories are not the ones who abandoned SEO. They are the ones who extended their SEO investment into the new surface before their competitors noticed the shift.</p>
</blockquote>
<p>This is a narrow window. AI visibility in most B2B categories is still wide open. Early citation authority is harder to displace than late entry.</p>
<h2>How to Start: Three Questions to Answer Before Allocating Budget</h2>
<p>Before splitting budget between SEO and AEO, answer these three questions. They will tell you where to weight your investment.</p>
<p><strong>Question 1: Are your buyers already using AI tools to research vendors?</strong></p>
<p>If you sell to marketing leaders, technology executives, or founders at growth-stage companies, the answer is almost certainly yes. Ask your last ten closed-won customers how they first heard about you. If any of them mention ChatGPT or Perplexity, AEO investment is already overdue.</p>
<p><strong>Question 2: What is your current AI visibility score?</strong></p>
<p>If you have never checked how often your brand appears in AI-generated answers for your category's key prompts, you are making a budget decision without data. Tools like <a href="https://www.getsearchable.com/">Searchable</a> give you a baseline AI visibility score, share of voice against competitors, and prompt-level citation data. Run this audit before committing to any budget split.</p>
<p><strong>Question 3: What is your content maturity on SEO?</strong></p>
<p>If your SEO content library is thin or your domain authority is low, building that foundation first will also accelerate AEO. If your SEO is strong, the incremental investment required to extend it into AEO is lower than you might expect. Much of the work is reformatting and restructuring what already exists.</p>
<p>The budget question is not really AEO vs SEO. It is: how much of your current SEO investment is reaching buyers who have already moved to AI-first research? That is the number worth knowing.</p>
<p>If you want to find out where you stand across ChatGPT, Perplexity, Google AI Overviews, and Gemini, <a href="https://windgrove.ai">Windgrove</a> runs AI visibility audits for B2B companies and builds the execution plan to close the gap.</p>
]]></content:encoded></item><item><title><![CDATA[AEO for B2B SaaS: Should You Build In-House, Hire a Consultant, or Work with a Specialist Agency?]]></title><description><![CDATA[TL;DR
In-house AEO requires a full content and technical team, a six-month minimum to show results, and ongoing investment in specialist tooling. It makes sense at scale.
An AEO consultant gives you a]]></description><link>https://windgrove.hashnode.dev/aeo-b2b-saas-build-in-house-hire</link><guid isPermaLink="true">https://windgrove.hashnode.dev/aeo-b2b-saas-build-in-house-hire</guid><category><![CDATA[B2B SaaS]]></category><dc:creator><![CDATA[Spencer Duke]]></dc:creator><pubDate>Wed, 09 Sep 2026 13:30:00 GMT</pubDate><content:encoded><![CDATA[<h2>TL;DR</h2>
<p><strong>In-house AEO</strong> requires a full content and technical team, a six-month minimum to show results, and ongoing investment in specialist tooling. It makes sense at scale.</p>
<p><strong>An AEO consultant</strong> gives you a strategy and a roadmap. Execution is still yours.</p>
<p><strong>A specialist AEO agency</strong> does the work end to end: technical infrastructure, content, citation building, and funnel attribution. For most B2B SaaS companies under $10M ARR, this is the fastest path from invisible to cited to booked demos to revenue.</p>
<p>The question is not which option sounds most appealing. It is which one will generate qualified pipeline in the next 90 days.</p>
<hr />
<p>Most B2B SaaS marketing teams discover AEO the same way. A founder notices a competitor getting recommended by ChatGPT on a category-level query. The team searches for their own product. Nothing. Then someone asks: do we build this capability internally, or do we bring in outside help?</p>
<p>It's a reasonable question. But also the wrong starting point.</p>
<p>The right starting point is the outcome. Not "how do we get AI mentions?" but "how do we get the right buyers into our pipeline through AI search?" That reframe changes every decision that follows.</p>
<p>This guide answers the in-house versus agency versus consultant question specifically for B2B SaaS companies. It covers what each path actually requires, when each one makes sense, and how to evaluate a specialist partner if you go that route.</p>
<h2>Why AEO Is a Different Problem for B2B SaaS</h2>
<p>B2B SaaS buyers already use AI search differently from consumers. They ask compound, evaluative questions: "best project management software for remote agencies under 50 people," "which CRM integrates with HubSpot and has a free trial," "top spend management tools for marketing teams." These are not informational queries. They are buying queries. The person asking has budget, a problem, and is shortlisting vendors.</p>
<p><a href="https://www.forrester.com">Research from Forrester</a> confirms that AI-assisted research now influences the majority of B2B technology purchase decisions, with buyers using tools like ChatGPT and Perplexity to generate initial shortlists before visiting a single vendor website. If your product is not in that shortlist, you are not in the consideration set.</p>
<p><strong>The uncomfortable truth:</strong> most B2B SaaS companies are invisible at exactly the moment a buyer is ready to evaluate options.</p>
<p>This is not a content volume problem. It is a citation mechanics problem. AI engines pull from structured sources: third-party directories, review platforms like G2 and Capterra, Reddit discussions, and content that directly answers the query in its first 100 words. Most SaaS blogs are built for Google's ranking algorithm, not for LLM citation patterns. The gap between "we publish a lot" and "we get cited" is wide.</p>
<h3>The Funnel Stage That Gets Ignored</h3>
<p>Most SaaS companies have some top-of-funnel AI presence. Category awareness content exists. What is almost universally missing is bottom-of-funnel visibility: the moment a buyer asks "how do I get started with [category] software" or "what is the onboarding process for [vendor]."</p>
<p>That is where deals are won and lost. That is where AEO work needs to go first.</p>
<h2>The Three Paths: What Each One Actually Requires</h2>
<p>Before choosing a path, be honest about what you're actually buying. Each option has a different cost structure, execution burden, and time to results.</p>
<h3>Path 1: Build In-House</h3>
<p>Building AEO capability internally is viable. It is also expensive and slow to compound.</p>
<p>To execute AEO properly in-house, a SaaS team needs:</p>
<p><strong>People required:</strong> A technical SEO or web developer to handle schema markup, robots.txt, canonical URL resolution, and llms.txt configuration. A content strategist who understands LLM citation mechanics (not just keyword strategy). A content writer producing four or more AEO-optimized articles per week. Someone managing third-party citation building: directory placements, Reddit, G2 and Capterra profiles, and external publication outreach.</p>
<p><strong>Tooling required:</strong> An AI visibility tracking platform (such as <a href="https://searchable.ai">Searchable</a>) to measure prompt-level citation performance across ChatGPT, Perplexity, Google AI Overviews, and Gemini. Standard SEO tooling does not measure this.</p>
<p><strong>Realistic timeline:</strong> Six to nine months before citation patterns shift meaningfully. The first two months are infrastructure and content foundation. Months three through six are authority building. Results compound after that.</p>
<p>In-house makes sense when you have an existing content team, a technical resource who can own schema and infrastructure, and the patience for a longer ramp. It is the right call for companies above $15M ARR with a dedicated growth function.</p>
<h3>Path 2: Hire an AEO Consultant</h3>
<p>An AEO consultant gives you a strategy, a prompt audit, and a prioritized roadmap. The execution stays with your team.</p>
<p>This works well when your team has the capacity to execute but lacks the specialised knowledge to know where to start. A good consultant will identify your highest-value prompt gaps, map your technical gaps, and tell you what to build and in what order.</p>
<p><strong>The limitation:</strong> consultants do not do the work. If your team is already stretched, a roadmap without execution adds to the backlog rather than reducing it.</p>
<h3>Path 3: Hire a Specialist AEO Agency</h3>
<p>A specialist agency handles execution end to end. Technical infrastructure, content production, citation building, directory placements, Reddit presence, and funnel attribution reporting. The client's job is to provide access and show up to a monthly call.</p>
<p>This is the right path when speed matters, when your team does not have AEO-specific expertise in-house, and when the goal is qualified pipeline within a defined timeframe rather than building an internal capability over 12 months.</p>
<table>
<thead>
<tr>
<th></th>
<th>In-House</th>
<th>Consultant</th>
<th>Specialist Agency</th>
</tr>
</thead>
<tbody><tr>
<td>Execution burden</td>
<td>High (your team)</td>
<td>High (your team)</td>
<td>Low (agency handles it)</td>
</tr>
<tr>
<td>Time to first results</td>
<td>6-9 months</td>
<td>4-6 months</td>
<td>60-90 days</td>
</tr>
<tr>
<td>Upfront cost</td>
<td>High (hiring)</td>
<td>Medium (retainer)</td>
<td>Medium (retainer)</td>
</tr>
<tr>
<td>Ongoing cost</td>
<td>High (salaries)</td>
<td>Low-medium</td>
<td>Predictable monthly</td>
</tr>
<tr>
<td>Pipeline attribution</td>
<td>Requires setup</td>
<td>Requires setup</td>
<td>Built into delivery</td>
</tr>
<tr>
<td>Best for</td>
<td>$15M+ ARR with growth team</td>
<td>Teams with capacity, no expertise</td>
<td>Sub-$15M ARR, speed required</td>
</tr>
</tbody></table>
<h2>How to Evaluate an AEO Agency: Five Questions That Matter</h2>
<p>If you decide to bring in a specialist partner, the evaluation criteria are different from a traditional content or SEO agency. Most agencies can produce content. Fewer can connect that content to AI citation. Fewer still can connect citations to pipeline.</p>
<p>Here are the five questions that separate a real AEO partner from a rebranded content shop.</p>
<h3>1. How Do You Measure AI Visibility?</h3>
<p>The answer should be specific. A credible AEO agency uses a platform that tracks prompt-level citation data across multiple AI engines. They should be able to show you: which prompts you currently appear in, what your share of voice is versus named competitors, and how that changes week over week.</p>
<p>If the answer is "we track rankings" or "we monitor Google AI Overviews," that is an SEO agency with a new label.</p>
<h3>2. What Does Your Technical Audit Cover?</h3>
<p>AI engines cannot cite content they cannot read. A specialist agency should audit your robots.txt, XML sitemap, schema markup implementation, canonical URL configuration, and llms.txt file before writing a single word of content. Technical infrastructure is Month 1 work. Any agency that skips it is building on an unstable foundation.</p>
<h3>3. Can You Show Funnel Attribution, Not Just Visibility Scores?</h3>
<p>Visibility is a leading indicator. Pipeline is the outcome. Ask to see how they connect AI citation improvements to actual leads entering the funnel. This means "how did you hear about us" attribution at the first touchpoint, CRM pipeline data, and a clear line from prompt performance to booked demos.</p>
<p>An agency that cannot show this is optimizing for a metric that does not pay your sales team.</p>
<h3>4. Do You Have SaaS-Specific Case Studies?</h3>
<p>AEO for a SaaS company looks different from AEO for a local service business or a healthcare provider. The prompt structure is different, the buyer journey is different, and the citation sources that matter are different (G2, Capterra, Reddit's SaaS-focused subreddits, and technology publications). Ask for examples where the work drove pipeline for a B2B software company specifically.</p>
<h3>5. What Does the Engagement Model Look Like?</h3>
<p>A credible agency should be able to tell you exactly what gets delivered in month one, month two, and month three. Deliverables, not intentions. If the answer is "it depends" without further specificity, that is a red flag.</p>
<h2>What a Done-for-You AEO Engagement Looks Like for B2B SaaS</h2>
<p>For SaaS companies that decide to work with a specialist partner, the engagement model matters as much as the strategy. Here is what a properly structured 90-day AEO engagement delivers, using Windgrove's B2B SaaS model as a reference.</p>
<h3>Month 1: Technical Foundation</h3>
<p>No content goes live until the site can be properly read and cited by LLMs. Month 1 covers:</p>
<p><strong>Technical deliverables:</strong> Schema markup implementation across all key pages (FAQ schema, Organisation schema, Article schema with author attribution). Canonical URL resolution. robots.txt and llms.txt configuration. XML sitemap audit and update. Baseline Searchable report showing current prompt-level visibility across ChatGPT, Perplexity, Google AI Overviews, and Gemini. Funnel attribution setup so LLM-sourced leads are tracked from first touchpoint through to booked demo.</p>
<p>Existing blog content is also reformatted in Month 1: direct answers front-loaded into the first 100 words, FAQ sections added with schema, publication dates and author attribution applied throughout.</p>
<h3>Month 2: Content and Citation Authority</h3>
<p>With the technical foundation in place, Month 2 focuses on content production and third-party citation building:</p>
<p><strong>Content deliverables:</strong> Four new AEO-optimised articles per week, each mapped to a specific tracked prompt. Comparison pages for the client's top competitor pairs. Bottom-of-funnel content covering "how to get started," "what the onboarding process looks like," and "how to book a demo," structured so AI engines can read and recommend it.</p>
<p><strong>Citation deliverables:</strong> G2 and Capterra profile optimisation. Directory placements on relevant authority sites. Reddit engagement in SaaS-specific subreddits where target buyers research options.</p>
<h3>Month 3: Visibility and Pipeline Attribution</h3>
<p>Month 3 shifts toward dominating the conversations target buyers are having. Emerging topic capture. Additional comparison pages. A full 90-day Searchable report showing visibility movement, share of voice changes, and LLM-attributed leads entering the pipeline.</p>
<blockquote>
<p><strong>Target outcomes at 90 days:</strong> 5-10% increase in AI visibility score. 3-5% increase in share of voice above 10%. Ten or more LLM-attributed leads entering the funnel is a great place to start, sometimes at a minimum.</p>
</blockquote>
<h2>Who This Is the Right Fit For</h2>
<p>Not every B2B SaaS company is at the right stage for a specialist AEO partner. Here is the profile where this engagement model generates the clearest return.</p>
<p><strong>Good fit:</strong> A B2B SaaS company with a defined ICP, an existing product with paying customers, and a marketing team that is stretched across too many channels to add AEO execution internally. Revenue between $1M and $15M ARR. A sales cycle where buyers research options using AI tools before booking a demo. A founder or marketing lead who understands that AI search is changing the discovery layer but does not have the bandwidth to own the execution.</p>
<p><strong>Not a good fit:</strong> Pre-product or pre-revenue companies without a defined ICP. Enterprise SaaS with 12-month procurement cycles where AI-search influence is harder to attribute. Companies that want to be involved in every content decision. Teams that have already hired an AEO-specific content strategist and technical resource internally.</p>
<p><strong>The honest question to ask yourself:</strong> if a qualified buyer asked ChatGPT which tool to use in your category right now, would your product appear? If the answer is no, or you do not know, that is the gap this engagement closes.</p>
<p>AEO is not a long-term brand play. For B2B SaaS, it is a pipeline lever. The window where most categories are still unclaimed is narrowing. The companies that move now will be the ones AI engines recommend by default six months from now.</p>
<p><a href="https://windgrove.ai">Learn more about Windgrove's AEO engagement model for B2B SaaS</a>.</p>
]]></content:encoded></item><item><title><![CDATA[How B2B Companies Get Cited in ChatGPT, Perplexity, and Google AI Overviews]]></title><description><![CDATA[The truth is your buyers are using AI to shortlist vendors. They type a question into ChatGPT, Perplexity, or Google AI Overviews and get a confident, well-structured answer that names three or four c]]></description><link>https://windgrove.hashnode.dev/b2b-ai-citations</link><guid isPermaLink="true">https://windgrove.hashnode.dev/b2b-ai-citations</guid><category><![CDATA[citations]]></category><dc:creator><![CDATA[Spencer Duke]]></dc:creator><pubDate>Mon, 07 Sep 2026 12:30:00 GMT</pubDate><content:encoded><![CDATA[<p>The truth is your buyers are using AI to shortlist vendors. They type a question into ChatGPT, Perplexity, or Google AI Overviews and get a confident, well-structured answer that names three or four companies. Yours isn't one of them.</p>
<p>This is not a traffic problem. It is not a content volume problem. It is a citation problem with a specific, fixable cause.</p>
<p>Most B2B companies have spent years optimizing for Google results: keyword density, backlinks, meta descriptions, page speed. That work is by no means wasted, but it doesn't transfer cleanly to AI engines. The mechanics are fundamentally different. Google ranks pages. AI engines synthesize answers. And the signals that drive a citation inside ChatGPT are not the same signals that drive a ranking in Google Search.</p>
<p>This guide covers exactly what those signals are, how they differ across platforms, and what B2B companies need to do to get cited consistently.</p>
<h2>TL;DR</h2>
<p>Getting cited in ChatGPT, Perplexity, and Google AI Overviews requires a different playbook than traditional SEO. AI engines retrieve from sources they trust, not just sources that rank. To get cited, your brand needs six things in place: technical access (LLM crawlers can actually read your site), entity consistency (your brand is described the same way everywhere), first-party proof (your own content directly answers the questions buyers are asking), answer-ready structure (your pages are formatted for extraction, not just reading), external corroboration (third-party sources confirm your authority), and a measurement system (so you know what is working). Most B2B companies have none of these. The ones that do are getting recommended. The gap is still wide open. The land is ripe for taking.</p>
<h2>Key Takeaways</h2>
<ul>
<li><p><strong>Technical access comes first.</strong> If LLM crawlers cannot read your site, nothing else matters. Fix robots.txt, add llms.txt, and resolve canonical conflicts before touching content.</p>
</li>
<li><p><strong>Entity consistency is the trust signal most teams miss.</strong> Inconsistent descriptions across LinkedIn, Crunchbase, G2, and your own site tell AI engines you are not a verified entity worth citing.</p>
</li>
<li><p><strong>A mention is not a citation.</strong> Appearing in an AI response is not the goal. Being cited as the source of a claim is. The gap between the two is structural, not promotional.</p>
</li>
<li><p><strong>External corroboration is non-negotiable.</strong> A brand that only exists on its own website will be mentioned at best. Reddit threads, trade publications, and directory listings are what elevate a mention to a citation.</p>
</li>
<li><p><strong>Decision-stage prompts are worth ten times awareness-stage prompts.</strong> Most B2B companies chase easy, broad queries. The citations that drive deals happen at the bottom of the funnel.</p>
</li>
<li><p><strong>This is not a campaign.</strong> AI citation visibility compounds with consistent execution and decays with neglect. The brands building this system now are establishing an advantage that gets harder to close every month.</p>
</li>
</ul>
<h2>How ChatGPT, Perplexity, and Google AI Overviews Actually Retrieve Sources</h2>
<p>The first mistake most B2B teams make is treating all AI platforms the same. They aren't. Each one retrieves information differently, weights sources differently, and produces citations in a different format. Understanding the distinctions is not academic, it changes where you focus your efforts.</p>
<h3>ChatGPT: Training data plus live retrieval</h3>
<p>ChatGPT operates in two distinct modes depending on whether web search is enabled. In its base form, it draws from its training data: a massive collection of web content, books, and structured data collected up to a specific cutoff date. Brands that were well-represented in that corpus, through high-authority publications, Wikipedia entries, Reddit threads, and widely-cited articles, have a big head start.</p>
<p>When <a href="https://help.openai.com/en/articles/8077698-how-do-i-use-chatgpt-browse-with-bing">ChatGPT's web search</a> is active, it performs a live query and retrieves current pages before generating its answer. In this mode, the citation mechanics shift closer to traditional search: recency matters, crawlability matters, and the structure of the retrieved page determines how cleanly it can be extracted and quoted.</p>
<p><strong>The practical implication:</strong> you need to be in both places. Training-data presence means building a citation footprint across the open web: publications, directories, Reddit, forums. Live-retrieval presence means your actual site pages need to be structured so they can be extracted cleanly when ChatGPT fetches them.</p>
<h3>Perplexity: Real-time retrieval with visible citations</h3>
<p>Perplexity is a retrieval-first engine. Every answer it generates is grounded in a live web search, and it shows its sources explicitly as numbered citations. This is the platform where traditional SEO signals matter most, because Perplexity is actively fetching pages and pulling content from them in real time.</p>
<p><a href="https://www.perplexity.ai/">Perplexity's retrieval system</a> prioritises pages that are crawl-able, clearly structured, and directly answer the query being asked. It favours sources that match the intent of the question, not just the keywords. A page that ranks well on Google but buries its answer in paragraph five will underperform relative to a page that leads with a direct, extractable answer.</p>
<p>Perplexity also weights third-party sources heavily. A mention in a credible publication, a well-upvoted Reddit thread, or a structured directory listing can generate a citation even if the brand's own website is not the primary source.</p>
<h3>Google AI Overviews: Search index plus trust signals</h3>
<p>Google AI Overviews (formerly Search Generative Experience) draws from Google's existing search index, which means the signals that drive traditional Google rankings: domain authority, quality backlinks, structured data, E-E-A-T (Experience, Expertise, Authoritativeness, Trustworthiness) all carry significant weight here. But ranking in the index is not sufficient. Google's AI layer has to determine that your content is synthesis-worthy: specific enough, authoritative enough, and structured enough to be pulled into a generated answer.</p>
<p><a href="https://developers.google.com/search/docs/appearance/ai-overviews">Google's documentation on AI Overviews</a> confirms that the system uses retrieval-augmented generation, meaning it actively fetches and synthesizes content rather than simply surfacing ranked pages. FAQ schema, clear heading hierarchies, and author attribution all improve the likelihood that a page gets pulled into an Overview rather than simply ranked below one.</p>
<p><strong>The critical difference from traditional SEO:</strong> ranking on page one does not guarantee inclusion in an AI Overview. A page ranked fifth with strong structured data and a direct answer in the first paragraph can outperform a page ranked first with a generic introduction. David can really beat Goliath within AEO.</p>
<h3>What this means across all three platforms</h3>
<table>
<thead>
<tr>
<th>Platform</th>
<th>Primary retrieval method</th>
<th>Citation format</th>
<th>Key differentiator</th>
</tr>
</thead>
<tbody><tr>
<td>ChatGPT</td>
<td>Training data + optional live search</td>
<td>Named in response, sometimes linked</td>
<td>Web-wide footprint + crawl-able site</td>
</tr>
<tr>
<td>Perplexity</td>
<td>Real-time web retrieval</td>
<td>Numbered citations, always visible</td>
<td>Page structure + third-party mentions</td>
</tr>
<tr>
<td>Google AI Overviews</td>
<td>Search index + retrieval-augmented generation</td>
<td>Inline snippets with source cards</td>
<td>E-E-A-T signals + structured data</td>
</tr>
</tbody></table>
<p>The common element across all three: AI engines cite sources they trust, and trust is built through a combination of web-wide presence, on-site structure, and third-party corroboration. None of those factors is accidental. All of them are buildable.</p>
<h2>The Difference Between a Brand Mention and a Citation</h2>
<p>This distinction matters more than most B2B teams realize, and conflating the two leads to a false sense of progress.</p>
<p><strong>A brand mention</strong> is when an AI engine includes your company's name somewhere in a generated response. It might appear in a list of five options, in a passing aside, or as a brief reference. Mentions are better than nothing. They are not the goal.</p>
<p><strong>A citation</strong> is when an AI engine actively attributes a claim, recommendation, or answer to your brand and pulls from your content as the source. A citation carries authority. A citation means the AI engine trusts your content enough to use it as evidence. A citation is what drives buyer behaviour. When ChatGPT or Perplexity says "according to [your brand]," the reader's confidence in both the answer and your company increases significantly.</p>
<h3>Why the gap between mention and citation is so large</h3>
<p>Most B2B brands that appear in AI responses at all appear as mentions. They show up because the model has encountered their name in training data or in a web result. But the model does not have enough confidence in their authority to cite them as a source. It names them, but does not back up that naming with structured evidence.</p>
<p>The factors that elevate a mention to a citation are specific:</p>
<p><strong>Extractable content:</strong> The AI engine needs to be able to pull a specific claim, answer, or data point from your content. If your page is a wall of prose with no clear structure, the model cannot extract a citable unit. Pages with direct answers in the first paragraph, FAQ sections with schema markup, and clearly headed subsections generate citations. Pages without those elements generate mentions at best.</p>
<p><strong>Third-party corroboration:</strong> A brand that only exists on its own website will be mentioned if the model has encountered it, but it will not be cited with confidence. Citation confidence comes from cross-referencing. When a model sees your brand mentioned on Reddit, in a trade publication, in a G2 review, and on your own site, it treats your brand as a verified entity worth citing. One source produces a mention. Multiple corroborating sources produce a citation.</p>
<p><strong>Entity clarity:</strong> AI engines build an internal model of what your company is, what it does, and who it serves. If your LinkedIn says one thing, your website says another, and your Crunchbase profile is outdated, the model's confidence drops. Inconsistent entity data produces vague mentions. Consistent, cross-platform entity data produces confident citations.</p>
<blockquote>
<p>The <strong>goal is not</strong> to appear in AI answers. The goal is to be <strong>cited as the authority</strong> in AI answers. Those are two different outcomes, and they require two different levels of execution.</p>
</blockquote>
<p>Most B2B companies are optimizing for the wrong outcome. They celebrate appearing in a ChatGPT response without asking whether they were cited as a source or simply named as an option. The difference determines whether that AI response drives a buyer toward your company or merely acknowledges that you exist.</p>
<p>This can be the difference between <strong>booking 30 sales calls</strong> from AI referrals or none at all.</p>
<h2>The Six-Pillar Citation-Readiness Framework</h2>
<p>Citation readiness is not a single fix. It is a system. The six pillars below are ordered by dependency: the first two are foundational and must be in place before the others compound. Skipping ahead produces diminishing returns.</p>
<h3>Pillar 1: Technical access (Core 4)</h3>
<p>AI engines cannot cite what they cannot read. This sounds obvious. It is also the most common failure point. Windgrove calls these the "Core 4".</p>
<p>The technical access layer covers everything that determines whether LLM crawlers can discover, fetch, and parse your site's content. The checklist is short but the consequences of getting it wrong are significant.</p>
<p><strong>robots.txt:</strong> Many B2B sites have robots.txt configurations that were set up for Google's crawler and inadvertently block other crawlers. Review yours. Any disallow rule that blocks broad crawler categories may be preventing AI engines from accessing your content entirely.</p>
<p><strong>llms.txt:</strong> A newer file, analogous to robots.txt but specifically for LLM crawlers. It signals which content is authoritative and citable. Setting this up correctly tells AI engines where to look and what to prioritize. <a href="https://llmstxt.org/">The llms.txt specification</a> is publicly documented and straightforward to implement.</p>
<p><strong>XML sitemap:</strong> Ensure your sitemap is current, accurate, and includes every page you want cited. Pages not in the sitemap are harder for AI crawlers to discover, particularly for sites with complex URL structures.</p>
<p><strong>Canonical URL consistency:</strong> If your site resolves at both www and non-www, or if URL parameters create duplicate versions of the same page, AI engines may treat your site as multiple separate entities. This dilutes citation authority. Resolve canonical conflicts before building anything else on top of them.</p>
<p>None of these require a developer for more than a few hours. They are the foundation. Everything else is built on them.</p>
<h3>Pillar 2: Entity consistency</h3>
<p>AI engines do not just read your website. They cross-reference your brand across the entire web and build a composite model of who you are. If the data is inconsistent, the model's confidence in citing you drops.</p>
<p>Entity consistency means your brand name, description, category, and key claims are identical across every platform where you appear:</p>
<p>LinkedIn company page. Google Business Profile. Crunchbase. G2. Capterra. Yelp. Apple Maps. Bing Places. Any industry-specific directories relevant to your category. Author bios on external publications. Job listings. Podcast descriptions.</p>
<p>Every one of these is a data point the model uses to verify your entity. A company that describes itself as "AI-powered project management" on LinkedIn but "workflow automation" on Crunchbase and "team collaboration software" on G2 is sending three conflicting signals. The model hedges. It may mention you but will not cite you with confidence.</p>
<p>Run an entity audit before you publish a single new piece of content. The audit takes less time than a week of blog writing and has more impact on citation frequency.</p>
<h3>Pillar 3: First-party proof</h3>
<p>Your own content is the primary signal that tells AI engines what you know, what you do, and who you serve. But most B2B content is written for Google, not for AI citation. The structural difference is significant.</p>
<p>Google rewards content that demonstrates topical depth over time. AI engines reward content that directly answers a specific question in the first paragraph.</p>
<p><strong>The citation-ready content structure looks like this:</strong></p>
<p>The first 100 words of any article or page should contain a direct, extractable answer to the question that page targets. Not a teaser. Not a hook. An answer. If someone asked ChatGPT the question your page is about, the model should be able to pull those first 100 words and use them as a cited response.</p>
<p>After the direct answer: supporting evidence, context, and depth. Subsections with clear H3 headings. An FAQ section at the end, structured with <a href="https://developers.google.com/search/docs/appearance/structured-data/faqpage">FAQ schema markup</a> so AI engines can pull individual question-answer pairs as standalone citations.</p>
<p>Author attribution on every piece. Publication dates. Internal links to related content. These are not cosmetic additions — they are signals that tell AI engines your content is maintained, authoritative, and part of a coherent knowledge base.</p>
<p><strong>The volume question:</strong> one well-structured article is worth more than ten generic ones. But coverage matters. If your category has 30 questions buyers commonly ask, you need content that answers all 30. The brands that own a category in AI answers are the ones that have answered every question a buyer might ask, at every stage of the buying journey.</p>
<h3>Pillar 4: Answer-ready formatting</h3>
<p>This pillar is distinct from first-party proof because it is about structure, not substance. You can have excellent content that is formatted in a way that AI engines cannot extract cleanly.</p>
<p>Answer-ready formatting means every page passes what can be called the extraction test: if an AI engine fetches this page and reads only the first 150 words plus the heading structure, does it have enough to generate a cited response?</p>
<p>The structural elements that drive extraction:</p>
<p><strong>H1 to H3 heading hierarchy.</strong> A clear, logical heading structure tells the model how your content is organised and what each section covers. Models use headings as anchors for extraction. A page with five paragraphs and no subheadings gives the model nothing to anchor to.</p>
<p><strong>FAQ sections with schema.</strong> Individual question-answer pairs that can be pulled as standalone citations. Each answer should be 40 to 80 words: long enough to be substantive, short enough to be extractable without truncation.</p>
<p><strong>Tables for comparative information.</strong> When you are comparing options, pricing tiers, features, or approaches, a markdown table is far more extractable than prose. AI engines parse tables cleanly and can cite specific rows.</p>
<p><strong>Short paragraphs.</strong> Paragraphs longer than 80 words create what might be called muddy embeddings — the model has difficulty identifying the single claim it should extract. Keep paragraphs focused. One idea per paragraph.</p>
<h3>Pillar 5: External corroboration</h3>
<p>This is the pillar most B2B companies ignore entirely, and it is the one that most directly determines whether you get cited or merely mentioned.</p>
<p>External corroboration means your brand is discussed, referenced, and recommended by sources other than yourself. AI engines treat external mentions as verification. They confirm that your brand is real, that other people find it valuable, and that it is worth recommending to someone who has not encountered it before.</p>
<p>The corroboration sources that carry the most weight across AI platforms:</p>
<p><strong>Reddit.</strong> ChatGPT, Perplexity, and Google all pull heavily from Reddit because it represents real human opinion at scale. A well-upvoted Reddit thread where your brand is recommended organically is worth more than a dozen press releases. This is a long-game strategy — Reddit accounts need warmup time before they can post authoritatively — but the citation half-life of a strong Reddit thread is measured in months, not days.</p>
<p><strong>Third-party publications.</strong> A mention in a trade publication, a feature in a relevant newsletter, or a quote in an industry roundup creates a citation source that AI engines can reference independently of your own site. The publications that matter most are the ones that AI engines are already citing for your category. Identifying those publications is more valuable than targeting publications by domain authority alone.</p>
<p><strong>Directories and review platforms.</strong> G2, Capterra, Trustpilot, Clutch, and vertical-specific directories are citation sources in their own right. A detailed G2 profile with reviews is something AI engines can pull from directly. An empty or outdated profile is a missed citation opportunity.</p>
<p><strong>Podcast appearances and YouTube content.</strong> AI engines increasingly reference video and audio content, particularly for "who is [person]" and "how does [concept] work" queries. A single podcast appearance can generate a transcript, a blog post, and a YouTube clip: three corroborating citation sources from one piece of work.</p>
<p>The goal is not to be mentioned everywhere. The goal is to be mentioned consistently, by credible sources, in contexts that match the queries you want to be cited for.</p>
<h3>Pillar 6: Measurement</h3>
<p>You cannot improve what you cannot measure. This is not a platitude in the context of AI citation — it is a practical constraint. AI visibility does not show up in Google Search Console. It does not appear in your GA4 dashboard. Without a dedicated measurement layer, you are flying blind.</p>
<p>The metrics that matter for B2B AI citation:</p>
<p><strong>AI visibility score:</strong> The percentage of tracked prompts where your brand is mentioned across AI platforms. This is your headline number. It should be tracked weekly, not monthly, because visibility can shift quickly when new content is published or a competitor makes a move.</p>
<p><strong>Share of voice:</strong> How often you appear versus named competitors across the same prompt set. Visibility score tells you how you are doing in absolute terms. Share of voice tells you how you are doing relative to the market.</p>
<p><strong>Prompt-level data:</strong> Which specific questions are generating citations, which are generating mentions, and which are generating nothing. This tells you where to focus content and corroboration efforts.</p>
<p><strong>Platform distribution:</strong> Whether your citations are concentrated on one platform or spread across ChatGPT, Perplexity, and Google AI Overviews. Concentration is a risk. If 80% of your AI citations come from Perplexity and Perplexity changes its retrieval algorithm, your visibility can drop overnight.</p>
<p><strong>Funnel-stage coverage:</strong> Are you being cited at the awareness stage, the consideration stage, or the decision stage? Most B2B companies have some awareness-stage visibility and almost no decision-stage citations. The decision stage is where deals are won.</p>
<p>Tools like <a href="https://searchable.ai/">Searchable</a> are built specifically to track these metrics across platforms. Without this layer, you are guessing. With it, you can connect visibility changes to specific content and corroboration work, and eventually to revenue. That connection is what turns AEO from a branding exercise into a growth system.</p>
<h2>The Most Common Mistakes B2B Companies Make</h2>
<p>The AEO category is new enough that most of the advice circulating online is either incomplete or wrong. These are the mistakes that consistently appear in B2B companies that are doing some of this work but still not getting cited.</p>
<h3>Mistake 1: Relying solely on schema markup</h3>
<p>Schema markup is valuable. It is not sufficient. A common pattern: a B2B team reads that FAQ schema helps with AI Overviews, adds schema to their existing content, and waits for citations to appear. They do not.</p>
<p>Schema tells AI engines how to parse your content. It does not tell them your content is worth citing. A page with perfect schema markup and a generic answer that does not directly address the query will not be cited. The schema has to be paired with content that is actually extractable and authoritative.</p>
<p>Schema is a signal amplifier. It amplifies the quality of what is already there. If what is already there is thin, schema makes thin content easier to parse — which is not the same as making it citable.</p>
<h3>Mistake 2: Publishing generic AI content</h3>
<p>The volume of generic "AI-generated" content on the web has increased dramatically. AI engines know this, and they are increasingly adept at identifying content that is semantically fluent but substantively empty. Publishing 20 articles that cover a topic at surface level does not build citation authority. It dilutes it.</p>
<p>The content that generates citations is specific, opinionated, and demonstrates genuine expertise. It answers questions that require real knowledge to answer well. It includes data, examples, and positions that a generic content generator cannot produce. If your content reads like it could have been written by anyone about anything, it will not be cited as an authority by anyone.</p>
<h3>Mistake 3: Building the citation footprint before fixing the foundation</h3>
<p>A common sequencing error: a company invests in Reddit presence, third-party publications, and directory listings before fixing their robots.txt, resolving canonical conflicts, or restructuring their on-site content. The external corroboration builds. The citations do not appear.</p>
<p>The reason is straightforward. External sources can confirm your brand exists and is discussed. But when an AI engine follows those signals back to your site to verify the claims, it finds content it cannot extract cleanly. The corroboration is there. The extractable authority is not. Both have to be in place for citations to follow.</p>
<h3>Mistake 4: Targeting the wrong prompts</h3>
<p>Not all AI queries are equal. A B2B company that gets cited for a broad awareness-stage query ("what is account-based marketing") is not in the same position as one that gets cited for a decision-stage query ("best account-based marketing platforms for mid-market B2B"). The second citation is worth ten times the first in terms of buyer impact.</p>
<p>Most B2B teams, when they start tracking AI visibility, focus on the prompts they can win easily. These tend to be broad, low-competition queries at the top of the funnel. Decision-stage prompts are harder to win and more valuable to win. The brands that own a category in AI answers are the ones that have systematically worked through the entire funnel, from awareness to purchase intent.</p>
<h3>Mistake 5: Treating AI visibility as a one-time project</h3>
<p>AI citation is not a campaign. It is an ongoing system. The web changes. Competitors publish new content. AI engines update their retrieval algorithms. A brand that achieves strong AI visibility in Q1 and stops working on it will see that visibility erode by Q3.</p>
<p>The companies that maintain strong AI citation over time are the ones that have built a continuous execution model: new content published regularly, entity profiles kept current, external corroboration actively managed, and visibility metrics reviewed weekly. The compounding effect of consistent execution is significant. The decay rate of neglected AI visibility is equally significant.</p>
<h2>The B2B AI Citation Execution Checklist</h2>
<p>This checklist maps directly to the six-pillar framework. Use it to audit where you are today and identify the highest-priority gaps.</p>
<h3>Foundation (do these first)</h3>
<p><strong>Technical access</strong></p>
<p>Audit your robots.txt for unintended crawler blocks. Create or update your llms.txt file. Verify your XML sitemap is current and includes all pages you want cited. Resolve any canonical URL conflicts between www and non-www, or between paginated and canonical versions. Confirm AI crawlers (GPTBot, PerplexityBot, Googlebot) are not blocked in your server configuration.</p>
<p><strong>Entity consistency</strong></p>
<p>List every platform where your brand has a profile. Compare your brand name, description, category, and core value proposition across all of them. Update any that are inconsistent, outdated, or missing. The platforms to prioritise: LinkedIn, Google Business Profile, Crunchbase, G2, and any vertical-specific directories in your category.</p>
<h3>Content and structure (do these second)</h3>
<p><strong>First-party proof</strong></p>
<p>Identify the 20 to 30 questions your buyers ask most frequently at each stage of the funnel. Map each question to an existing page or identify the gap. For each existing page, check whether the first 100 words contain a direct, extractable answer to the question that page targets. Rewrite any that do not. Add author attribution and publication dates to all content.</p>
<p><strong>Answer-ready formatting</strong></p>
<p>Audit your top 10 pages for heading hierarchy (H1, H2, H3 structure). Add or restructure subheadings where they are missing. Add an FAQ section with schema markup to every article and service page. Break up any paragraphs longer than 80 words. Add tables to any pages that compare options, pricing, or features.</p>
<h3>Corroboration and measurement (do these third)</h3>
<p><strong>External corroboration</strong></p>
<p>Identify the publications, forums, and directories that AI engines are currently citing for your category. Search for your category on ChatGPT and Perplexity and note which sources appear in the citations. Prioritise building a presence on those specific sources. For Reddit: identify the three to five subreddits most relevant to your buyers and begin genuine engagement before any brand-adjacent posting.</p>
<p><strong>Measurement</strong></p>
<p>Set up a tracking system for the 20 to 30 prompts most relevant to your category. Run each prompt across ChatGPT, Perplexity, and Google AI Overviews weekly. Record whether your brand is cited, mentioned, or absent. Track changes over time. Connect visibility improvements to specific content or corroboration work so you understand what is driving results.</p>
<h3>The 90-day sequence</h3>
<table>
<thead>
<tr>
<th>Month</th>
<th>Focus</th>
<th>Key deliverables</th>
</tr>
</thead>
<tbody><tr>
<td>Month 1</td>
<td>Foundation</td>
<td>Technical fixes, entity audit, existing content restructured for extraction</td>
</tr>
<tr>
<td>Month 2</td>
<td>Authority</td>
<td>New citation-ready content, Reddit engagement, directory placements</td>
</tr>
<tr>
<td>Month 3</td>
<td>Visibility</td>
<td>Comparison pages, emerging topic capture, full funnel coverage, measurement review</td>
</tr>
</tbody></table>
<p>The sequence matters. Skipping Month 1 and going straight to content production is the most common execution mistake. The foundation determines how much of your content investment compounds into citations.</p>
<h2>Good News: The Window Is Still Open</h2>
<p>Most B2B companies have not started this work. The AI visibility landscape in most categories is still wide open. The brands that move now will establish citation authority before their competitors understand why it matters.</p>
<p>That advantage compounds. A brand that is cited consistently today is building the training data footprint and retrieval trust that will make it harder to displace six months from now. A brand that waits is not just losing citations today. It is ceding the compounding advantage to whoever moves first.</p>
<p>The six pillars are not complicated. They are sequential, and they require consistent execution. Technical access, entity consistency, first-party proof, answer-ready formatting, external corroboration, measurement. In that order. With that cadence.</p>
<p>The buyers are already asking AI engines for vendor recommendations. The question is whether your brand is in the answer.</p>
]]></content:encoded></item><item><title><![CDATA[Why Your Business Isn't Showing Up in ChatGPT (And How to Fix It)]]></title><description><![CDATA[Most businesses aren't showing up in ChatGPT because they've built their digital presence for Google, and Google and ChatGPT are fundamentally different engines with different credibility signals. The]]></description><link>https://windgrove.hashnode.dev/chatgpt-visibility</link><guid isPermaLink="true">https://windgrove.hashnode.dev/chatgpt-visibility</guid><category><![CDATA[aeo]]></category><category><![CDATA[chatgpt]]></category><dc:creator><![CDATA[Spencer Duke]]></dc:creator><pubDate>Sat, 25 Jul 2026 22:19:54 GMT</pubDate><content:encoded><![CDATA[<p>Most businesses aren't showing up in ChatGPT because they've built their digital presence for Google, and Google and ChatGPT are fundamentally different engines with different credibility signals. The fix isn't more content. It's building the right kind of presence in the right places.</p>
<hr />
<p><strong>Key Takeaways</strong></p>
<ul>
<li><p>94.7% of brands receive zero mentions in ChatGPT recommendations, according to research by Profound Strategy.</p>
</li>
<li><p>Ranking on Google does not translate to appearing in AI-generated answers. The two systems use different signals entirely.</p>
</li>
<li><p>The four most common reasons for AI invisibility are: vague content that doesn't answer real questions, weak entity signals across the web, technical barriers blocking AI crawlers, and no third-party validation.</p>
</li>
<li><p>The fastest wins come from restructuring existing content to lead with direct answers, not from publishing more of it.</p>
</li>
<li><p>AI visibility compounds over time. The brands dominating ChatGPT results in most categories got there by starting early, not by outspending anyone.</p>
</li>
</ul>
<hr />
<p>Open ChatGPT right now and type: <em>"What's the best [your type of business] in [your city]?"</em></p>
<p>For most business owners, the result is the same: a confident, well-written answer that names competitors, cites sources, and makes recommendations without mentioning you once. Not because you've done anything wrong. Because you've been optimizing for the wrong engine entirely.</p>
<p><strong>The uncomfortable truth:</strong> <a href="https://www.profoundstrategy.com">According to research by Profound Strategy</a>, 94.7% of brands receive zero mentions in ChatGPT recommendations. That's not a rounding error. That's nearly every business on the internet, invisible to a tool that millions of people now use to make buying decisions.</p>
<p>This isn't a Google ranking problem. It's a different problem, with different causes, and a different set of solutions. Here's what's actually going on, and where to start.</p>
<h2>ChatGPT Doesn't Work Like Google</h2>
<p>To understand why you're invisible, you first need to understand what ChatGPT is actually doing when it answers a question.</p>
<p>Google crawls your website and ranks it based on hundreds of signals: keywords, backlinks, page speed, user engagement. You show up when someone searches. ChatGPT does something fundamentally different. It synthesizes an answer from what it already knows, drawing on training data, real-time web sources (in the case of tools like Perplexity and ChatGPT with browsing enabled), and signals about which sources are credible and authoritative.</p>
<p>The result is that ranking on Google does not translate to appearing in ChatGPT. A brand can sit at position one on Google for a competitive keyword and be completely absent from AI-generated answers on the same topic.</p>
<blockquote>
<p><strong>The core distinction:</strong> Google ranks pages. AI engines cite sources they trust. These are two different credibility systems, and most businesses have only built for one of them.</p>
</blockquote>
<h3>How AI Engines Decide What to Cite</h3>
<p>When an AI model constructs an answer, it is effectively asking: <em>"What do I know about this, and which sources have established authority on the topic?"</em> The signals it weighs include:</p>
<ul>
<li><p><strong>Entity recognition:</strong> Is this brand a clearly defined, verifiable entity with consistent information across the web?</p>
</li>
<li><p><strong>Content structure:</strong> Is the content formatted in a way that makes it easy for a language model to extract a clean, citable answer?</p>
</li>
<li><p><strong>External validation:</strong> Is this brand mentioned, recommended, or linked to by sources the AI already trusts?</p>
</li>
<li><p><strong>Topical depth:</strong> Does this source demonstrate genuine expertise on the subject, or does it only mention it in passing?</p>
</li>
</ul>
<p>None of these map neatly to traditional SEO tactics. A well-optimized title tag does nothing for AI visibility. A page stuffed with keywords may actually work against you if it lacks the directness and structure an AI needs to extract a useful answer.</p>
<h2>The Four Reasons You're Not Showing Up</h2>
<p>Most businesses fall into at least one of these gaps. Many fall into all four.</p>
<h3>1. Your Content Answers the Wrong Questions</h3>
<p>AI engines are built to respond to natural-language queries. When someone asks ChatGPT <em>"What's the best accounting software for a small restaurant?"</em>, it looks for content that directly answers that kind of question. Not content that says <em>"We offer industry-leading financial solutions for the hospitality sector."</em></p>
<p>Marketing language is invisible to AI. Specificity is what gets cited.</p>
<p>If your website is full of positioning statements and feature lists but lacks direct, plain-language answers to the questions your customers actually ask, you have a content structure problem. Not a content volume problem.</p>
<h3>2. You Don't Exist as a Verified Entity</h3>
<p>This is the gap most business owners don't expect. AI models build a picture of your brand from dozens of signals across the web: your <a href="https://business.google.com">Google Business Profile</a>, your LinkedIn page, mentions in industry publications, your schema markup, directory listings, and more. When these signals are inconsistent, incomplete, or absent, the model can't confidently identify you as a real, trustworthy entity.</p>
<p><strong>Think of it this way:</strong> if a journalist had never heard of your company and tried to verify it existed in 10 minutes using only public sources, what would they find? For most small businesses, the answer is: not much. AI models face the same challenge.</p>
<h3>3. Your Site Has Technical Barriers AI Can't Get Past</h3>
<p>Even if your content is excellent, it may be inaccessible to AI crawlers. Common blockers include:</p>
<table>
<thead>
<tr>
<th>Issue</th>
<th>What Happens</th>
</tr>
</thead>
<tbody><tr>
<td>JavaScript-rendered content</td>
<td>AI sees a blank page instead of your text</td>
</tr>
<tr>
<td>Robots.txt blocking AI crawlers</td>
<td>Your site is explicitly telling AI not to read it</td>
</tr>
<tr>
<td>Gated content (forms, paywalls)</td>
<td>The AI simply can't access what's behind the gate</td>
</tr>
<tr>
<td>Missing structured data (schema)</td>
<td>The AI can read your content but can't interpret it efficiently</td>
</tr>
</tbody></table>
<p>These are fixable, but they require a technical audit. Many businesses unknowingly block AI crawlers while allowing Google's bots through, assuming the two are equivalent. They are not.</p>
<h3>4. No One Else Is Talking About You</h3>
<p>Third-party validation is one of the strongest signals AI engines use to determine authority. If your brand is only mentioned on your own website, that's a weak signal. If it's mentioned in industry publications, cited in forum discussions, referenced in <a href="https://www.reddit.com">Reddit threads</a>, and linked to from authoritative sources, that tells the AI your brand is real, established, and worth recommending.</p>
<p>This is the "digital footprint" problem. A business can be genuinely excellent and completely unknown to an AI because the web hasn't validated it yet.</p>
<h2>Where to Start: The First Moves That Actually Matter</h2>
<p>The good news is that AI visibility is still early enough that the bar isn't impossibly high. The brands dominating ChatGPT results in most industries got there not because they had massive budgets, but because they started earlier and built for the right signals.</p>
<p>Here's where to begin.</p>
<h3>Run a Visibility Audit First</h3>
<p>Before fixing anything, you need to know where you actually stand. Open ChatGPT, Perplexity, and Google's AI Overview, then search for the questions your customers are most likely to ask. Not your brand name. The actual questions.</p>
<ul>
<li><p><em>"What's the best [your service] for [your customer type]?"</em></p>
</li>
<li><p><em>"How do I choose a [your category] provider?"</em></p>
</li>
<li><p><em>"Who are the leading [your industry] companies in [your region]?"</em></p>
</li>
</ul>
<p>Note which brands appear, which sources are cited, and whether you show up at all. This gives you a baseline and tells you exactly who you're competing against in the AI layer.</p>
<h3>Restructure Your Content for Direct Answers</h3>
<p>The most impactful change most businesses can make is rewriting their key pages to lead with direct, specific answers. Not preamble. Not brand story. The answer, in the first sentence.</p>
<p>A page that starts with <em>"Founded in 2015, we've been helping businesses grow..."</em> will not be cited. A page that starts with <em>"[Business name] provides [specific service] for [specific customer type] in [location], specializing in [specific outcome]"</em> gives an AI exactly what it needs to cite you confidently.</p>
<p><strong>Structure your content like this:</strong></p>
<ol>
<li><p>Direct answer to the question the page targets (1-2 sentences)</p>
</li>
<li><p>Why it matters / context (1 short paragraph)</p>
</li>
<li><p>Supporting evidence, examples, or data</p>
</li>
</ol>
<p>This isn't just good for AI. It's better for human readers too.</p>
<h3>Build Your Entity Footprint</h3>
<p>Audit every place your brand should appear and make sure the information is consistent and complete:</p>
<ul>
<li><p><strong>Google Business Profile:</strong> Claimed, verified, fully filled out with accurate categories and description</p>
</li>
<li><p><strong>LinkedIn company page:</strong> Active, with a clear description of what you do and who you serve</p>
</li>
<li><p><strong>Schema markup:</strong> At minimum, <a href="https://schema.org/Organization">Organization schema</a> on your homepage with your name, address, phone, and description</p>
</li>
<li><p><strong>Industry directories:</strong> Listed with consistent NAP (name, address, phone) data</p>
</li>
<li><p><strong>Third-party mentions:</strong> Press releases, guest articles, podcast appearances, or any content that puts your brand name in contexts AI can find and trust</p>
</li>
</ul>
<p>Each of these signals, individually, is small. Together, they tell AI models that your business is a real, verifiable entity worth recommending.</p>
<h3>Don't Ignore the Platforms AI Already Trusts</h3>
<p>AI engines heavily weight content from platforms they already consider authoritative. <a href="https://www.youtube.com">YouTube</a>, LinkedIn, Reddit, and industry publications are consistently cited in AI responses. If your brand has a presence on these platforms with content that addresses the questions your customers ask, you're building visibility in the places AI is already looking.</p>
<p>This is why a single well-written LinkedIn article or a YouTube video with a detailed description can sometimes drive more AI visibility than a dozen new blog posts on your own site.</p>
<h2>The Window Is Still Open, But It Won't Be Forever</h2>
<p>AI search is not a future trend to plan for. It's where your customers are finding answers right now. <a href="https://www.gartner.com/en/newsroom">Gartner projects</a> that 25% of traditional search traffic will shift to AI-generated answers by the end of 2026. The businesses that show up in those answers are being chosen as the default recommendation before a buyer ever visits a website.</p>
<p>The brands dominating ChatGPT results in most categories got there by acting early, not by outspending anyone. That window is still open in most industries, but it closes as more competitors wake up to the shift.</p>
<p><strong>The first step is simply knowing where you stand.</strong> Run the audit. Search for your category in ChatGPT and see who's being recommended instead of you. That answer tells you everything about the gap you need to close.</p>
<p>AEO isn't about gaming an algorithm. It's about building the kind of digital presence that AI models recognize as authoritative, consistent, and worth recommending. The fundamentals: clear content, verified entity signals, technical accessibility, and third-party validation, are the same things that make a brand genuinely trustworthy. The difference is that now, getting those fundamentals right determines whether AI recommends you or your competitors.</p>
<h2>What AI Visibility Looks Like in Practice</h2>
<p>To make this concrete, consider what happens when a high-intent buyer searches for a service like yours in ChatGPT. The AI doesn't return a list of links. It returns a synthesized recommendation, confident, specific, and sourced. The brands it names are the ones that have done the work to be recognizable, verifiable, and citation-ready.</p>
<p>The screenshot below illustrates a typical ChatGPT response to a business-category query. Notice how it cites specific sources and names specific brands, not the ones with the biggest ad spend, but the ones with the clearest, most accessible, most authoritative presence.</p>
<img src="https://imagedelivery.net/mGEFJEtXKBGNgaFU-kVEbQ/5bc2a1f5-5498-4bd9-e7e3-c91a45c66900/public" alt="ChatGPT responding to a business category search query, showing how AI engines cite specific brands and sources in their answers" style="display:block;margin:0 auto" />

<p>This is the new front page of the internet. And unlike Google's first page, there are only a handful of spots, and no paid placements.</p>
]]></content:encoded></item><item><title><![CDATA[How AI Search Engines Recommend B2B Companies (And Why Most Are Invisible)]]></title><description><![CDATA[Your next B2B customer is probably asking an AI right now which vendor to use. Not searching Google. Not scrolling LinkedIn. Asking ChatGPT, Perplexity, or Claude a question like: "What's the best CRM]]></description><link>https://windgrove.hashnode.dev/ai-search</link><guid isPermaLink="true">https://windgrove.hashnode.dev/ai-search</guid><category><![CDATA[aeo]]></category><category><![CDATA[ai search]]></category><dc:creator><![CDATA[Spencer Duke]]></dc:creator><pubDate>Sat, 25 Jul 2026 20:22:05 GMT</pubDate><content:encoded><![CDATA[<p>Your next B2B customer is probably asking an AI right now which vendor to use. Not searching Google. Not scrolling LinkedIn. Asking ChatGPT, Perplexity, or Claude a question like: <em>"What's the best CRM for a 50-person SaaS company?"</em> or <em>"Which marketing automation platform integrates with Salesforce?"</em></p>
<p>If your company isn't in the answer, you don't exist for that buyer.</p>
<p><strong>The shift is already here.</strong> According to a <a href="https://www.hubspot.com">HubSpot survey of B2B buyers</a>, 48% now use AI search while evaluating vendors. A separate survey by Responsive found that 80% of tech buyers rely on generative AI at least as much as traditional search to research vendors. These aren't early adopters anymore — this is mainstream buying behaviour.</p>
<p>The problem is that most B2B companies are still optimizing for Google while their prospects have already moved on. And the rules for showing up in AI answers are fundamentally different from the rules for ranking in search.</p>
<p>This article explains how AI search engines actually decide which companies to recommend — and what you can do to make sure yours is one of them.</p>
<h2>AI Search Is Not Just Faster Google</h2>
<p>Most founders assume AI search works like Google with a chatbot interface on top. It doesn't. The underlying mechanics are completely different, and that distinction determines whether your company gets recommended or ignored.</p>
<p>Google ranks pages. AI engines synthesize answers.</p>
<p>When a buyer asks Google "best project management software for agencies," they get a list of ten blue links and decide where to click. When they ask ChatGPT the same question, they get a personalized, conversational recommendation — often with specific vendor names, reasons why, and comparisons — all generated from the AI's internal understanding of the market.</p>
<p>That internal understanding is built from everything the model has been trained on: your website, yes, but also third-party publications, review platforms, directories, community forums, and industry databases. The AI is not crawling the web in real time for most queries. It is drawing on a synthesized model of your brand's reputation across the entire digital ecosystem.</p>
<blockquote>
<p><strong>The key insight:</strong> Your website is your resume. AI search is calling your references. What third parties say about you matters far more than what you say about yourself.</p>
</blockquote>
<p>There is also a critical overlap problem. Research shows that <a href="https://searchengineland.com">only about 12% of URLs cited by AI engines sit in Google's top 10</a> for the same query. That means strong SEO rankings and strong AI visibility are largely independent outcomes. A company can dominate Google and be invisible to AI — and increasingly, that is exactly what is happening.</p>
<h2>How LLMs Actually Decide Who to Recommend</h2>
<p>When a B2B buyer asks an AI assistant to recommend a vendor, the model runs through a layered decision process. Understanding each layer is the first step to influencing the outcome.</p>
<h3>Step 1: Entity Recognition</h3>
<p>AI models think in terms of entities and relationships, not keywords. Before recommending your company, the model needs to "know" you exist as a distinct, clearly defined entity in your category. This means your brand name, what you do, who you serve, and how you relate to other known entities (competitors, integrations, categories) must be consistently represented across the web.</p>
<p>Inconsistent naming, vague positioning, or a thin digital footprint causes the model to either ignore your brand or misrepresent it. One common failure mode: a company appears in AI answers but is described in the wrong category or with outdated information because the model has conflicting signals.</p>
<h3>Step 2: Authority and Trust Signals</h3>
<p>Once the model recognizes your entity, it evaluates authority. This is where third-party signals dominate:</p>
<ul>
<li><p><strong>Earned media:</strong> Coverage in respected industry publications</p>
</li>
<li><p><strong>Review platforms:</strong> G2, Capterra, Trustpilot, and similar sites</p>
</li>
<li><p><strong>Directories and databases:</strong> Crunchbase, LinkedIn, industry-specific listings</p>
</li>
<li><p><strong>Community mentions:</strong> Reddit threads, forums, and discussion boards</p>
</li>
<li><p><strong>Backlinks from authoritative domains:</strong> Signals that others vouch for your expertise</p>
</li>
</ul>
<p>The model weights these signals to determine how confidently it can recommend you. A company with deep third-party coverage gets recommended with conviction. A company that only has its own website gets skipped.</p>
<h3>Step 3: Relevance to the Buyer's Query</h3>
<p>AI assistants personalize answers based on the specific context of the query. A buyer asking "best CRM for a bootstrapped SaaS startup" gets a different answer than one asking "enterprise CRM with HIPAA compliance." The model matches your brand to queries based on how clearly your content addresses specific use cases, company sizes, industries, and pain points.</p>
<p>This is why broad, generic positioning hurts AI visibility. If your content doesn't clearly signal who you're for and what problems you solve, the model can't confidently place you in the right answers.</p>
<h3>Step 4: Freshness and Recency</h3>
<p>AI models are regularly updated and increasingly use retrieval-augmented generation (RAG) to pull in current web content for time-sensitive queries. Brands that publish consistently, earn ongoing media coverage, and maintain active profiles across platforms signal to the model that they are current and relevant. Brands that went quiet two years ago may have strong historical training data but declining real-time signals.</p>
<h2>Each AI Engine Has Different Priorities</h2>
<p>Not all AI search engines work the same way. A B2B company visible in ChatGPT may be invisible in Perplexity, and vice versa. Understanding each platform's bias helps you prioritize where to focus.</p>
<p><strong>The practical implication:</strong> A brand that appears consistently across all five platforms is far more likely to be recommended regardless of which AI a buyer happens to use. Platform-specific optimization is a second-order concern; broad authority is the foundation.</p>
<p>One important nuance: Perplexity runs a live web search for every query, which means it can surface newer or smaller companies that have strong recent coverage, even if they lack years of accumulated authority. For B2B companies that are earlier in their growth, this is a meaningful opportunity.</p>
<h2>What B2B Companies Can Do About It</h2>
<p>Understanding the mechanics is useful. Knowing what to do about them is what matters. AI visibility is not a one-time fix — it is a system that compounds over time. Here is where to start.</p>
<h3>Build Topical Authority, Not Just Pages</h3>
<p>AI engines favour brands that comprehensively own a topic, not brands that have one good blog post. Pick three to five core topics that directly map to your buyers' questions and build deep, interconnected content around each one. The model needs to be able to identify your brand as the authoritative source on the problems you solve.</p>
<h3>Earn Third-Party Coverage Systematically</h3>
<p>Your own website is the weakest signal you can send. Industry publications, analyst mentions, review platforms, and community discussions are what AI models treat as social proof. One well-placed feature in a respected industry publication generates more AI citation weight than twenty blog posts on your own site.</p>
<h3>Establish Clear Entity Signals</h3>
<p>Use your exact brand name consistently across every platform — your website, LinkedIn, Crunchbase, G2, industry directories, and press mentions. Inconsistency creates ambiguity. Ambiguity means the model either skips you or misrepresents you.</p>
<h3>Structure Your Content for AI Extraction</h3>
<p>AI engines extract specific passages from your content to use in answers. Each section of your content should open with a clear, direct answer to a specific buyer question — not a preamble. Content structured around questions that buyers actually ask in AI prompts ("What is the best [category] for [use case]?") is far more likely to be quoted directly.</p>
<h3>Monitor Your Visibility Across Platforms</h3>
<p>Most B2B companies have no idea how they appear in AI answers right now. Running your ten most important buyer queries through ChatGPT, Perplexity, Claude, and Google AI Overviews monthly is the minimum baseline. Track whether you appear, how you are described, and which competitors are being recommended instead of you.</p>
<blockquote>
<p><strong>The first-mover advantage is real.</strong> LLM perception drift, the month-over-month shift in how AI models position brands in a category, is already reshaping B2B markets. <a href="https://searchengineland.com/why-llm-perception-drift-will-be-2026s-key-seo-metric-465676">Research from Search Engine Land</a> shows that brands like Atlassian gained significant AI visibility scores in a single month while established competitors dropped — purely based on which brands had stronger semantic anchoring in the model's training data.</p>
</blockquote>
<p>Companies that build AI visibility now are compounding an advantage that will be very difficult to close in two years.</p>
<h2>The Window Is Open — For Now</h2>
<p>AI search is not a future trend to monitor. It is an active buying channel where your prospects are building shortlists today. The mechanics favour companies that move early: AI models develop brand associations over time, and those associations are sticky. Brands that establish strong entity signals and third-party authority now will be the default recommendations in their categories by 2027.</p>
<p>The good news for B2B companies is that most of your competitors haven't started. The AI visibility landscape in most B2B categories is still wide open. That is a narrow window.</p>
<p>If you want to understand where your company stands right now — which queries you appear in, how you're described, and what it would take to move up — <a href="https://windgrove.ai">Windgrove AI</a> offers AI visibility audits specifically built for B2B companies. It's the fastest way to see your current position and build a roadmap to own it.</p>
<h2>Related reading</h2>
<ul>
<li><a href="https://windgrove.ai/blog/what-is-aeo">What Is AEO (Answer Engine Optimization) and How Does It Work?</a></li>
<li><a href="https://windgrove.ai/blog/chatgpt-visibility">Why Your Business Isn't Showing Up in ChatGPT (And How to Fix It)</a></li>
<li><a href="https://windgrove.ai/blog/ai-citation-sources">The Sites That Actually Influence AI Answers in Your Industry</a></li>
</ul>
]]></content:encoded></item><item><title><![CDATA[What Is AEO (Answer Engine Optimization) and How Does It Work?]]></title><description><![CDATA[Search has always been about getting found. For decades, that meant ranking on Google, earning clicks, and driving traffic to your website. But that model is changing faster than most businesses reali]]></description><link>https://windgrove.hashnode.dev/what-is-aeo</link><guid isPermaLink="true">https://windgrove.hashnode.dev/what-is-aeo</guid><category><![CDATA[aeo]]></category><dc:creator><![CDATA[Spencer Duke]]></dc:creator><pubDate>Sat, 25 Jul 2026 20:12:16 GMT</pubDate><content:encoded><![CDATA[<p>Search has always been about getting found. For decades, that meant ranking on Google, earning clicks, and driving traffic to your website. But that model is changing faster than most businesses realize, and the companies that adapt earliest will hold an outsized advantage over those that wait.</p>
<p><strong>The core shift:</strong> AI-powered answer engines like ChatGPT, Perplexity, Google AI Overviews, Gemini, and Claude are no longer just tools people use to draft emails. They are becoming the first stop for business research, vendor evaluation, and purchasing decisions. When a potential client types a question into one of these platforms, they receive a synthesized answer, often without ever visiting a website. If your brand is not part of that answer, you are invisible at the most critical moment in the buyer journey.</p>
<p>This is where Answer Engine Optimization (AEO) comes in.</p>
<p>This guide covers what AEO is, how it works mechanically, how it differs from traditional SEO, and why the window to act is narrower than most businesses think.</p>
<h2>What Is AEO (Answer Engine Optimization)?</h2>
<p>Answer Engine Optimization (AEO) is the practice of structuring and positioning your content so that AI-powered answer engines cite your brand when responding to relevant queries. Rather than optimizing for a position on a search results page, AEO optimizes for inclusion in a synthesized AI response, the kind of direct, conversational answer that ChatGPT, Perplexity, Google AI Overviews, and similar platforms deliver to users.</p>
<p>The distinction is significant. Traditional search optimization earns you a link. AEO earns you a citation inside the answer itself, meaning your brand becomes part of the information a user receives, not just an option they might click on.</p>
<blockquote>
<p><strong>Key definition:</strong> AEO is not a replacement for SEO. It is an extension of it, applied to a new class of search interface where the engine synthesizes answers rather than listing links.</p>
</blockquote>
<h3>What "Answer Engines" Actually Are</h3>
<p>An answer engine is any AI system that interprets a natural language query and generates a direct, synthesized response. This includes:</p>
<ul>
<li><p><strong>ChatGPT</strong> (OpenAI): the dominant <a href="https://windgrove.ai/blog/ai-search">AI search</a> platform, processing over 2.5 billion prompts per day</p>
</li>
<li><p><strong>Perplexity AI</strong>: a search-native AI tool that cites sources inline and is growing rapidly among research-oriented users</p>
</li>
<li><p><strong>Google AI Overviews (AIO)</strong>: Google's AI-generated summaries that now appear at the top of search results for over 25% of queries</p>
</li>
<li><p><strong>Gemini</strong> (Google): Google's conversational AI assistant, deeply integrated with Google Workspace</p>
</li>
<li><p><strong>Claude</strong> (Anthropic): increasingly used for research, writing assistance, and professional queries</p>
</li>
<li><p><strong>Microsoft Copilot</strong>: embedded across Microsoft 365 and Bing, making it significant for enterprise B2B audiences</p>
</li>
</ul>
<p>Each of these platforms pulls from different data sources, weights authority signals differently, and serves a slightly different user base. Effective AEO accounts for all of them, not just one.</p>
<h2>How the Search Market Is Shifting</h2>
<p>The numbers behind this shift are significant enough that they warrant attention from any business that relies on inbound discovery.</p>
<p><strong>60% of searches in the US and EU now result in zero clicks</strong> due to AI Overviews and featured snippets answering the query directly on the results page. The average click-through rate for a website ranking first on Google dropped from 0.73 to 0.26 between March 2024 and March 2025, a 64% reduction in less than 12 months. Meanwhile, <a href="https://www.conductor.com/academy/aeo-geo-benchmarks-report/">according to Conductor's 2026 AEO Benchmarks Report</a>, AI referral traffic is growing by roughly 1% of total web visits per month across major industries.</p>
<p>That growth rate may sound modest until you consider the quality of the traffic it represents.</p>
<h3>AI Traffic Converts Differently</h3>
<p>Traffic arriving from AI citations is not equivalent to organic search traffic. It is materially better qualified. When an AI engine recommends your brand in response to a specific query, the user arrives having already received a pre-qualified endorsement. They are not browsing. They are evaluating.</p>
<p><strong>The conversion rate difference is stark:</strong> research shows AI-referred traffic converts at approximately 14.2% on average, compared to 2.8% from traditional Google organic search. That is a 5x difference in conversion performance from the same discovery moment.</p>
<p>For B2B service providers, this matters enormously. A prospective client asking an AI "which agencies specialize in AEO for professional services firms" and receiving a recommendation is a fundamentally different lead than someone clicking a generic search result. The intent is more specific, the trust is higher, and the sales cycle is shorter.</p>
<h3>The Buyer Journey Has Compressed</h3>
<p><a href="https://blog.hubspot.com/marketing/answer-engine-optimization-trends">HubSpot's Consumer Trends Report</a> found that 72% of consumers plan to use AI-powered search for purchasing decisions more frequently. Answer engines are compressing the traditional buyer journey by delivering synthesized comparisons, recommendations, and vendor shortlists in a single response. For B2B buyers, who already conduct extensive pre-purchase research before engaging a vendor, this compression is accelerating the moment of shortlist formation.</p>
<p>The practical implication: the brands that appear in AI-generated answers are being shortlisted before a buyer ever reaches a company website. Those that do not appear are being screened out at the same stage.</p>
<h2>How AEO Works: The Core Mechanics</h2>
<p>AEO is not a single tactic. It is a system of overlapping signals that, together, tell AI engines that your brand is a credible, citable authority on a given topic. Understanding the mechanics helps explain why some brands appear consistently in AI answers while others with strong traditional SEO rankings do not.</p>
<h3>1. Content Structure and Answer-First Formatting</h3>
<p>AI engines extract information differently from how Google's crawler indexes it. They look for clear, direct answers to specific questions, not pages that bury the answer in the fifth paragraph. Content formatted for LLM extraction is <a href="https://www.jacklimebear.com/post/state-of-answer-engine-optimization-aeo-2026">three times more likely to be cited</a> than content that is not.</p>
<p>Practically, this means:</p>
<ul>
<li><p>Leading each section with a concise, direct answer (40-60 words) before expanding on the detail</p>
</li>
<li><p>Using clear heading structures that mirror how people phrase questions</p>
</li>
<li><p>Writing self-contained sections that make sense when extracted in isolation, without needing surrounding context</p>
</li>
<li><p>Including structured FAQ content that maps directly to the queries your audience asks</p>
</li>
</ul>
<h3>2. Schema Markup and Structured Data</h3>
<p>Schema markup is code added to a webpage that tells AI systems and search engines exactly what the content represents. For AEO, schema is a critical trust signal. Brands with comprehensive schema markup see <a href="https://aeoengine.ai/blog/state-of-ai-search-complete-guide">57% more AI Overview triggers</a> on long-tail queries compared to those without it.</p>
<p>Relevant schema types for AEO include Article, FAQ, HowTo, Organization, and Person. Each provides structured context that makes it easier for an AI to confidently cite your content without risk of misrepresenting it.</p>
<h3>3. Domain Authority and Third-Party Citations</h3>
<p>There is a strong positive correlation (0.65 linear correlation) between a website's overall authority and how frequently it appears in AI citations. This is not accidental. AI engines are trained to minimize hallucinations and errors, so they naturally favour sources that have demonstrated credibility across multiple platforms.</p>
<p>Third-party citations matter significantly:</p>
<ul>
<li><p>47.9% of ChatGPT referrals come from Wikipedia</p>
</li>
<li><p>Reddit appears in 21% of Google AI Overview citations</p>
</li>
<li><p>LinkedIn accounts for 13% of Google AI Overview citations</p>
</li>
<li><p>YouTube appears in 13.9% of Perplexity citations</p>
</li>
</ul>
<p>This tells us that AEO is not only about your own website. It is about building a presence across the platforms AI engines trust most.</p>
<h3>4. Entity Consistency</h3>
<p>AI engines build a model of who and what your brand is based on how consistently you appear across the web. If your company name, description, and area of expertise are described differently across your website, LinkedIn, industry directories, and press mentions, AI systems struggle to confidently include you in answers. Entity consistency means maintaining a coherent, unified description of your brand across every platform where you have a presence.</p>
<h3>5. Topical Authority</h3>
<p>AI engines favour brands that demonstrate deep, consistent expertise in a specific domain over generalists that cover everything shallowly. Publishing a cluster of well-structured, interconnected content around a specific topic signals to AI systems that your brand is a reliable authority on that subject, making you a safer citation choice.</p>
<h2>AEO vs. SEO: What Is the Difference?</h2>
<p><a href="https://windgrove.ai/blog/aeo-vs-seo-where-should-a-b2b-tech-company-invest-in-2026">AEO and SEO</a> are complementary disciplines, not competing ones. A strong SEO foundation, particularly domain authority and high-quality content, directly supports AEO performance. But the goals, tactics, and success metrics are distinct.</p>
<table>
<thead>
<tr>
<th>Dimension</th>
<th>SEO (Search Engine Optimization)</th>
<th>AEO (Answer Engine Optimization)</th>
</tr>
</thead>
<tbody><tr>
<td><strong>Primary goal</strong></td>
<td>Rank on a search results page</td>
<td>Be cited in an AI-generated answer</td>
</tr>
<tr>
<td><strong>Success metric</strong></td>
<td>Rankings, organic clicks, impressions</td>
<td>AI citations, brand mentions in answers</td>
</tr>
<tr>
<td><strong>Content format</strong></td>
<td>Keyword-optimized pages</td>
<td>Answer-first, structured, self-contained sections</td>
</tr>
<tr>
<td><strong>Key signals</strong></td>
<td>Backlinks, page authority, technical health</td>
<td>Entity consistency, schema, topical authority, third-party citations</td>
</tr>
<tr>
<td><strong>Traffic type</strong></td>
<td>Click-through from SERP</td>
<td>Direct referral from AI answer</td>
</tr>
<tr>
<td><strong>Buyer stage</strong></td>
<td>Varies widely</td>
<td>Typically high intent, mid-to-late evaluation</td>
</tr>
<tr>
<td><strong>Ranking required?</strong></td>
<td>Yes, by definition</td>
<td>No. Pages outside the top 10 can appear in AI answers</td>
</tr>
</tbody></table>
<p>The last row is worth emphasizing. Contrary to common assumption, a high Google ranking is not a prerequisite for appearing in AI-generated answers. AI systems surface the clearest, most contextually relevant answer regardless of traditional search position. A well-structured page on page two of Google can outperform a page-one result in AI citations if it is better formatted and more authoritative on the specific question being asked.</p>
<p>This creates a meaningful opportunity for businesses that have not yet built dominant SEO authority. AEO is, in many ways, a more level playing field, at least for now.</p>
<h2>Why Businesses Need to Act Now</h2>
<p>The AEO landscape is nascent, and that is precisely why timing matters. Early movers in any new visibility channel build compounding advantages that become increasingly difficult to displace.</p>
<p><strong>The first-mover window is real.</strong> In traditional SEO, it took years for brands to accumulate the domain authority and backlink profiles that dominate today's search results. AEO is following a similar trajectory. The brands that establish topical authority, structured content, and third-party citation networks now will be the default recommendations AI engines serve to buyers in 12 to 24 months.</p>
<h3>The Cost of Waiting</h3>
<p>Consider what happens to a B2B service provider that delays AEO investment while a competitor does not. That competitor begins appearing in AI-generated answers to queries like "best [service type] firms for [industry]" or "how to choose a [service category] provider." Buyers receive that recommendation before they ever search Google, before they visit any website, and before a sales conversation begins. The shortlist is formed without the waiting brand ever having a chance to compete.</p>
<p>This is not a hypothetical future scenario. It is already happening across industries. <a href="https://www.conductor.com/academy/aeo-geo-benchmarks-report/">Conductor's cross-industry analysis</a> of over 3.3 billion web sessions found that AI referral traffic is growing by approximately 1% of total web traffic per month. The IT industry already sees 2.8% of all traffic arriving from AI referrals, with that figure rising monthly.</p>
<h3>The Measurement Shift</h3>
<p>AEO also requires a fundamental change in how businesses measure visibility. Traditional KPIs, including keyword rankings and organic impressions, do not capture AI citation performance. Brands investing in AEO need to track:</p>
<ul>
<li><p>How often their brand is cited in AI-generated answers for target queries</p>
</li>
<li><p>The volume and quality of leads arriving from AI referral sources</p>
</li>
<li><p>Which content is being extracted and cited by which platforms</p>
</li>
<li><p>Brand mention frequency across ChatGPT, Perplexity, Google AIO, and other engines</p>
</li>
</ul>
<p>This is a new measurement discipline, and building it early means having a benchmark before the channel becomes competitive.</p>
<h2>How to Get Started With AEO</h2>
<p>AEO does not require abandoning existing SEO or marketing investments. It requires extending them with a new layer of intent. The following steps represent a practical starting point for any business beginning to build AI visibility.</p>
<h3>Step 1: Audit Your Current AI Visibility</h3>
<p>Before optimizing, understand your baseline. Query the major AI engines directly with questions your ideal clients are likely to ask. See which brands appear. See whether your brand appears. This gives you a clear picture of the gap you are working to close and which platforms to prioritize first.</p>
<h3>Step 2: Restructure Content for Answer Extraction</h3>
<p>Review your existing content and identify pages that address high-value buyer questions. Restructure them to lead with direct answers, use clear heading hierarchies, and include FAQ sections that map to how buyers phrase their queries. Content that already ranks well on Google is often the best candidate for AEO restructuring, since it has existing authority that AI engines can build on.</p>
<h3>Step 3: Implement Comprehensive Schema Markup</h3>
<p>Add structured data to your key pages. At minimum, implement Article, FAQ, and Organization schema. If you have team profiles, add Person schema. Schema is one of the highest-leverage AEO investments because it directly reduces the ambiguity AI systems face when deciding whether to cite your content.</p>
<h3>Step 4: Build Third-Party Authority Signals</h3>
<p>Identify the platforms AI engines draw from most heavily in your industry and build a presence there. For most B2B service providers, this means:</p>
<ul>
<li><p>A complete, keyword-rich LinkedIn company page and individual profiles for key team members</p>
</li>
<li><p>Listings on credible industry directories and review platforms</p>
</li>
<li><p>Contributions to relevant communities and publications that AI engines cite regularly</p>
</li>
<li><p>Consistent NAP (name, address, phone) data across all directories</p>
</li>
</ul>
<h3>Step 5: Track and Iterate</h3>
<p>Set up a monitoring system to track AI citations for your target queries. Measure the volume and quality of leads arriving from AI referral sources. Use this data to identify which content is performing and where gaps remain. AEO is not a one-time project. It is an ongoing system of optimization, similar to how effective SEO has always worked.</p>
<h2>The Bottom Line</h2>
<p>AEO represents a fundamental shift in how businesses earn visibility, not a trend to monitor from a distance. The buyer journey is being compressed by AI-powered answer engines that form shortlists before a prospect ever visits a website. The brands being cited in those answers are winning opportunities they did not have to compete for in the traditional sense. The brands absent from those answers are being filtered out before the conversation even starts.</p>
<p>The good news is that the field is not yet saturated. Most businesses have not started. That creates a genuine first-mover advantage for those willing to invest in building AI visibility now, while the compounding returns are still accessible.</p>
<p><strong>AEO is not about chasing a new algorithm.</strong> It is about building the kind of structured, authoritative, consistently presented brand that AI engines trust enough to recommend. That is a sound long-term investment regardless of how the technology evolves.</p>
<p>For businesses that rely on inbound discovery to generate leads, the question is no longer whether to invest in AEO. It is how quickly they can build the foundation before their competitors do.</p>
<hr />
<p><strong>Not sure where your business stands in AI search?</strong></p>
<p>Most B2B companies have no idea whether ChatGPT, Perplexity, or Gemini are recommending them or their competitors. We audit your AI visibility in 30 minutes and show you exactly where you stand.</p>
<p><a href="https://windgrove.ai">Get Your Free AI Visibility Audit →</a></p>
<h2>Related reading</h2>
<ul>
<li><a href="https://windgrove.ai/blog/aeo-roi">Why Investing in AEO Now Pays Compounding Dividends Later</a></li>
<li><a href="https://windgrove.ai/blog/chatgpt-visibility">Why Your Business Isn't Showing Up in ChatGPT (And How to Fix It)</a></li>
<li><a href="https://windgrove.ai/blog/ai-citation-sources">The Sites That Actually Influence AI Answers in Your Industry</a></li>
</ul>
]]></content:encoded></item><item><title><![CDATA[The Technical AEO Audit: 12 Things Your Website Needs Before AI Tools Will Cite You]]></title><description><![CDATA[Most founders who discover their company isn't showing up in ChatGPT assume it's a content problem. They think they need more blog posts, a PR push, or a stronger social presence. Sometimes that's tru]]></description><link>https://windgrove.hashnode.dev/the-technical-aeo-audit-12-things-your-website-needs-before-ai-tools-will-cite-you</link><guid isPermaLink="true">https://windgrove.hashnode.dev/the-technical-aeo-audit-12-things-your-website-needs-before-ai-tools-will-cite-you</guid><category><![CDATA[citations]]></category><category><![CDATA[aeo]]></category><dc:creator><![CDATA[Mitko Dimitrov]]></dc:creator><pubDate>Fri, 19 Jun 2026 00:00:00 GMT</pubDate><content:encoded><![CDATA[<p>Most founders who discover their company isn't showing up in ChatGPT assume it's a content problem. They think they need more blog posts, a PR push, or a stronger social presence. Sometimes that's true. But more often, the issue is more fundamental: AI crawlers can't properly read your site, don't know who you are as an entity, or can't extract a clean answer from your pages.</p>
<p>Content and citation work can't move the needle on a technically broken foundation. This is the part most AEO conversations skip over. Before you write a single piece of AI-optimised content, your site needs to pass a basic infrastructure test.</p>
<p>Here are the 12 things we check in every audit, what breaking any one of them costs you in citations, and how to check where you stand today.</p>
<h2>TL;DR</h2>
<p>Before any content or citation work can move the needle, your site needs to be technically readable by LLMs. Most sites aren't. The 12 infrastructure items below are what unlock AI citation. Check each one before you spend a dollar on content.</p>
<h2>Why Technical Infrastructure Comes First</h2>
<p>AI engines don't browse your site the way a human does. They crawl it, parse structured signals, cross-reference your entity against third-party sources, and extract answers from your content. If any part of that chain is broken, your content never enters the citation pool, regardless of how good it is.</p>
<p>Think of it this way: you can write the most authoritative, well-structured answer to a target query. But if your robots.txt is blocking the crawler, or your canonical tags are pointing in the wrong direction, or your content has no author attribution, the LLM either can't access it or doesn't trust it enough to cite it.</p>
<p><strong>The uncomfortable truth:</strong> most companies that are invisible in AI-generated answers have a technical problem, not a content problem.</p>
<p>The 12 items below fall into four categories:</p>
<p><strong>Crawlability:</strong> Can AI systems actually access your site and the right pages on it?</p>
<p><strong>Structured data:</strong> Have you told AI engines what your content is, who wrote it, and what organisation it belongs to?</p>
<p><strong>Content readability:</strong> Is your content structured so that an LLM can extract a clean, citable answer?</p>
<p><strong>Entity consistency:</strong> Does your brand present itself identically across every platform an AI engine might cross-reference?</p>
<p>Fix these before anything else. The content work compounds on top of a clean foundation. Without it, you're building on sand.</p>
<h2>The 12-Item Technical AEO Audit</h2>
<h3>1. robots.txt</h3>
<p><strong>What it is:</strong> The file that tells crawlers which pages they can and can't access.</p>
<p><strong>What breaking it costs you:</strong> If your robots.txt is blocking the crawlers that AI platforms use, your entire site is invisible to them. It doesn't matter what's on those pages.</p>
<p><strong>How to audit it:</strong> Go to yourdomain.com/robots.txt. Look for any Disallow rules that might be blocking your key pages or entire directories. Specifically check whether you're blocking any user-agents you don't recognise. Many sites have legacy rules that made sense years ago and now block modern crawlers by accident.</p>
<p><strong>What good looks like:</strong> Your key pages and content directories are fully accessible. You're not running blanket Disallow rules without understanding what they're blocking.</p>
<hr />
<h3>2. <a href="https://windgrove.ai/blog/llmstxt-explained-what-it-is-why-it-matters-and-why-google-now-checks-for-it">llms.txt</a></h3>
<p><strong>What it is:</strong> A newer file, analogous to robots.txt, but built specifically for LLM crawlers. It signals which content on your site is authoritative, citable, and intended for AI consumption.</p>
<p><strong>What breaking it costs you:</strong> Without it, LLMs have no explicit signal about which of your pages to prioritise. They make their own inferences, which aren't always correct.</p>
<p><strong>How to audit it:</strong> Go to yourdomain.com/llms.txt. If you get a 404, you don't have one. Most sites don't.</p>
<p><strong>What good looks like:</strong> The file exists, points to your most important pages and content, and is structured according to the <a href="https://llmstxt.org/">llms.txt specification</a>.</p>
<hr />
<h3>3. XML Sitemap Completeness</h3>
<p><strong>What it is:</strong> The map that tells search engines and AI crawlers which pages exist on your site.</p>
<p><strong>What breaking it costs you:</strong> Pages missing from your sitemap are pages that LLMs may never discover. Even if those pages are technically accessible, they're less likely to be crawled and indexed.</p>
<p><strong>How to audit it:</strong> Go to yourdomain.com/sitemap.xml. Cross-reference the pages listed against the pages you actually want indexed. Look for missing blog posts, landing pages, or service pages. Also check that the sitemap is submitted in <a href="https://search.google.com/search-console/">Google Search Console</a>.</p>
<p><strong>What good looks like:</strong> Every page you want cited is in the sitemap. The sitemap is current and has no broken URLs.</p>
<hr />
<h3>4. Canonical URL Resolution</h3>
<p><strong>What it is:</strong> Canonical tags tell crawlers which version of a URL is the "official" one when duplicates exist.</p>
<p><strong>What breaking it costs you:</strong> www vs non-www conflicts, URL parameters, and canonical mismatches cause LLMs to treat your site as multiple separate entities. This splits your citation authority across versions of the same page and dilutes your overall signal.</p>
<p><strong>How to audit it:</strong> Use a tool like <a href="https://www.screamingfrog.co.uk/seo-spider/">Screaming Frog</a> to crawl your site and flag canonical conflicts. Check that your www and non-www versions redirect to a single canonical. Look for pages where the canonical tag points somewhere unexpected.</p>
<p><strong>What good looks like:</strong> One canonical version per page, consistently enforced. No conflicting signals across URL variants.</p>
<hr />
<h3>5. FAQ Schema</h3>
<p><strong>What it is:</strong> Structured data markup that labels individual question-and-answer pairs on your pages so AI engines can extract and cite them directly.</p>
<p><strong>What breaking it costs you:</strong> Without FAQ schema, an LLM has to infer which part of your page answers a given query. With it, you're handing the engine a pre-packaged answer it can pull verbatim. This is one of the highest-leverage schema types for AI citation.</p>
<p><strong>How to audit it:</strong> Use <a href="https://search.google.com/test/rich-results">Google's Rich Results Test</a> on your key pages. If FAQ schema is present and valid, it will show up. If not, it won't.</p>
<p><strong>What good looks like:</strong> Every page targeting a question-based query has valid FAQ schema. Each Q&amp;A pair is self-contained, specific, and directly answers the question it's paired with.</p>
<hr />
<h3>6. Organisation Schema</h3>
<p><strong>What it is:</strong> Structured data that tells AI engines who you are as a company: your name, URL, logo, contact information, social profiles, and founding details.</p>
<p><strong>What breaking it costs you:</strong> Without it, LLMs are assembling a picture of your organisation from scattered signals. That picture is often incomplete or inconsistent, which reduces citation confidence.</p>
<p><strong>How to audit it:</strong> Use <a href="https://validator.schema.org/">Schema.org's validator</a> or Google's Rich Results Test on your homepage. Check whether Organisation schema is present and whether it includes your name, URL, logo, and sameAs links to your social profiles and directories.</p>
<p><strong>What good looks like:</strong> Valid Organisation schema on your homepage. All key fields populated. sameAs links pointing to your LinkedIn, Crunchbase, Google Business Profile, and any other authoritative directory listings.</p>
<hr />
<h3>7. Article Schema with Author Attribution</h3>
<p><strong>What it is:</strong> Structured data on your blog posts and articles that identifies the piece as an article, names the author, and includes the publication date.</p>
<p><strong>What breaking it costs you:</strong> AI engines look for authorship signals when deciding whether to cite a piece of content. Anonymous content with no author attribution is treated as lower-trust. Content with a named author, credentials, and a publication date is treated as more citable.</p>
<p><strong>How to audit it:</strong> Run your blog posts through Google's Rich Results Test. Check whether Article schema is present. Look for the author field and the datePublished field specifically.</p>
<p><strong>What good looks like:</strong> Every article has Article schema. The author field points to a real person with a bio page. datePublished is accurate and present.</p>
<hr />
<h3>8. H1 Direct Answers</h3>
<p><strong>What it is:</strong> The practice of placing a direct, concise answer to the page's target question in the first 100 words of the content.</p>
<p><strong>What breaking it costs you:</strong> LLMs look near the top of a page for extractable answers. If your content buries the answer 500 words in after a long introduction, the LLM either misses it or passes over your page in favour of one that answers the question immediately.</p>
<p><strong>How to audit it:</strong> Open your key pages and read the first paragraph. Ask yourself: if someone searched for the query this page targets, is the answer clearly stated in the opening lines? Or does the page start with background, context, and scene-setting before getting to the point?</p>
<p><strong>What good looks like:</strong> The first 100 words of every key page contain a direct, specific answer to the target query. No long wind-ups. No "In this article, we'll explore..." openings.</p>
<hr />
<h3>9. Publication Dates on Content</h3>
<p><strong>What it is:</strong> Visible and schema-tagged publication dates on your articles and blog posts.</p>
<p><strong>What breaking it costs you:</strong> AI engines use publication dates to assess content freshness and relevance. Undated content is treated as lower-confidence. For queries where recency matters, undated content often loses to dated content even when the quality is higher.</p>
<p><strong>How to audit it:</strong> Go through your blog and check whether each post displays a publication date visibly on the page. Then verify that the datePublished field is present in the Article schema.</p>
<p><strong>What good looks like:</strong> Every piece of content has a visible, accurate publication date. The same date appears in the Article schema. If you've updated a piece, the dateModified field is also present.</p>
<hr />
<h3>10. Internal Link Architecture</h3>
<p><strong>What it is:</strong> The network of links between your own pages that signals content relationships and helps crawlers navigate your site.</p>
<p><strong>What breaking it costs you:</strong> A site with weak internal linking is a collection of isolated pages from an AI crawler's perspective. LLMs use internal link patterns to understand which pages are most important and how topics relate to each other. Orphaned pages, or pages with no inbound internal links, are often invisible in practice even if they're technically indexed.</p>
<p><strong>How to audit it:</strong> Use Screaming Frog to map your internal link structure. Look for pages with zero or very few inbound internal links. Check whether your most important content pages are linked from multiple relevant locations across the site.</p>
<p><strong>What good looks like:</strong> Key pages are linked from multiple relevant locations. Topic clusters are connected logically. No important pages are sitting as orphans.</p>
<hr />
<h3>11. Google Business Profile Consistency</h3>
<p><strong>What it is:</strong> Your Google Business Profile listing, which AI engines use as a primary signal for local and category-based queries.</p>
<p><strong>What breaking it costs you:</strong> An incomplete, inconsistent, or unclaimed GBP means you're invisible for category-level queries, not just location-specific ones. LLMs pull from GBP data when generating recommendations for "best [service type] in [city]" and similar queries.</p>
<p><strong>How to audit it:</strong> Search for your business name in Google Maps. Check that the listing is claimed, that the business name matches exactly what's on your website, that the description is accurate and complete, and that your hours, phone number, and URL are current.</p>
<p><strong>What good looks like:</strong> Claimed, complete, and consistent. Business name, description, and contact details match your website exactly. Categories are accurate and specific.</p>
<hr />
<h3>12. Entity Alignment Across Directories</h3>
<p><strong>What it is:</strong> The consistency of your brand's name, description, URL, and contact details across every platform an AI engine might cross-reference: LinkedIn, Crunchbase, Apple Maps, Bing Places, and any vertical-specific directories relevant to your industry.</p>
<p><strong>What breaking it costs you:</strong> AI engines cross-reference your brand across the web before deciding how confidently to cite you. If LinkedIn says one thing, Crunchbase says another, and your website says a third, the LLM's confidence in your entity drops. Inconsistency reads as unreliability.</p>
<p><strong>How to audit it:</strong> Search your brand name across the major directories. Check that your business name is identical everywhere (not "Acme Inc." in one place and "Acme" in another). Verify that your URL, description, and founding details are consistent.</p>
<p><strong>What good looks like:</strong> Identical name, URL, and description across every listing. No outdated information sitting on platforms you forgot you signed up for three years ago.</p>
<h2>How to Prioritise When You Find Problems</h2>
<p>Most sites that go through this audit find issues in multiple areas. The question is where to start.</p>
<p>The answer depends on severity, but there's a general order that makes sense for most sites:</p>
<p><strong>Fix crawlability first.</strong> robots.txt, sitemap, and canonical issues are foundational. If crawlers can't access your site or are getting confused about which URL is canonical, nothing else matters. These are also usually the fastest to fix.</p>
<p><strong>Implement structured data next.</strong> Organisation schema, Article schema, and FAQ schema have the highest direct impact on citation probability. They're the difference between an LLM having to guess what your content is about and being told explicitly.</p>
<p><strong>Then address content readability.</strong> H1 direct answers and publication dates are content-level changes. They require going through your existing pages and reformatting. This takes time but compounds quickly once done.</p>
<p><strong>Entity consistency is ongoing.</strong> Directory alignment isn't a one-time fix. As you add new listings and update your business details, consistency needs to be maintained.</p>
<blockquote>
<p><strong>The single highest-impact fix for most sites:</strong> FAQ schema on key pages, combined with H1 direct answers in the opening paragraph. These two changes alone can move citation rates meaningfully for sites that previously had neither.</p>
</blockquote>
<h2>Frequently Asked Questions</h2>
<h3>Can I do a technical AEO audit myself?</h3>
<p>Yes. Everything in this checklist is something a technically capable founder or developer can work through independently. The tools referenced here (Screaming Frog, Google's Rich Results Test, Schema.org's validator, Google Search Console) are all free or low-cost. The audit itself isn't the hard part. The hard part is fixing everything you find, especially the structured data implementation, which requires adding and validating JSON-LD across your entire site. If you have a developer comfortable with schema markup, the technical fixes are manageable. If you don't, the implementation phase is where most teams get stuck.</p>
<h3>How long does it take to implement AEO technical fixes?</h3>
<p>For a site with 50 to 200 pages, a thorough technical AEO implementation typically takes three to four weeks when done properly. Crawlability fixes (robots.txt, sitemap, canonicals) can be done in a day or two. Structured data implementation across all key pages is the most time-consuming part, particularly if your CMS doesn't have native schema support. Entity alignment across directories can be done in parallel and usually takes a week of focused work. Plan for four weeks total if you're doing this alongside other priorities.</p>
<h3>Which technical AEO fix has the biggest impact?</h3>
<p>FAQ schema, consistently. It's the fix that most directly changes how AI engines interact with your content. When you add valid FAQ schema to a page, you're giving LLMs a pre-packaged, extractable answer they can cite verbatim. Without it, they have to infer the answer from unstructured prose. That inference is less reliable and less likely to result in a citation. Pair FAQ schema with H1 direct answers in the opening paragraph and you've addressed the two most common reasons AI engines pass over otherwise good content.</p>
<h2>We Handle All 12 of These in Month 1</h2>
<p>This checklist is the foundation layer of every Windgrove engagement. Before we write a single piece of content or build a single citation, we go through all 12 of these items on your site. We fix what's broken, implement what's missing, and verify everything is working before the content engine starts.</p>
<p>Here is what that looks like in practice:</p>
<p><strong>Week 1:</strong> robots.txt audit and rewrite, sitemap review, canonical conflict identification and resolution, llms.txt configuration.</p>
<p><strong>Week 2:</strong> Organisation schema, Article schema, and FAQ schema implemented across your key pages. Rich Results Test validation on every page touched.</p>
<p><strong>Week 3:</strong> H1 direct answer restructuring on existing content, publication date audit and correction, author attribution added to all articles.</p>
<p><strong>Week 4:</strong> Google Business Profile optimisation, entity alignment across LinkedIn, Crunchbase, Apple Maps, Bing Places, and any vertical-specific directories. Internal link architecture review and gap-filling.</p>
<p>By the end of Month 1, your site is technically ready for AI citation. Month 2 is when the content work begins, and it compounds on a foundation that's actually solid.</p>
<p>If you want to see what this looks like for your specific site, <a href="https://windgrove.ai/">we're happy to take a look</a>. We'll tell you exactly where you stand across all 12 items before any commitment.</p>
]]></content:encoded></item><item><title><![CDATA[AEO vs SEO: Where Should a B2B Tech Company Invest in 2026?]]></title><description><![CDATA[TL;DR
SEO and AEO are not the same game. SEO wins you Google blue links. AEO wins you AI citations. For B2B companies where buyers are using ChatGPT to build shortlists before your sales team even kno]]></description><link>https://windgrove.hashnode.dev/aeo-vs-seo-where-should-a-b2b-tech-company-invest-in-2026</link><guid isPermaLink="true">https://windgrove.hashnode.dev/aeo-vs-seo-where-should-a-b2b-tech-company-invest-in-2026</guid><dc:creator><![CDATA[Mitko Dimitrov]]></dc:creator><pubDate>Thu, 18 Jun 2026 00:00:00 GMT</pubDate><enclosure url="https://cdn.hashnode.com/res/hashnode/image/upload/v1785121015454/e9a308a4-cb41-4891-a61d-e2ce78137ab4.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>TL;DR</h2>
<p>SEO and AEO are not the same game. SEO wins you Google blue links. AEO wins you AI citations. For B2B companies where buyers are using ChatGPT to build shortlists before your sales team even knows a deal exists, AEO is now the higher-<a href="https://windgrove.ai/blog/aeo-roi">ROI</a> investment. But you need both. This post covers what each optimises for, how the signals differ, where they overlap, and how to allocate your budget between them.</p>
<p>Here is the uncomfortable reality of B2B marketing in 2026: your buyers have already ranked you before they contact you.</p>
<p>According to <a href="https://machinerelations.ai/research/b2b-ai-vendor-research-2026">Forrester's 2026 Buyers' Journey Survey of nearly 18,000 global business buyers</a>, 94% of B2B buyers used AI during their most recent purchase process, and AI tools outranked vendor websites, sales reps, and product experts as the most meaningful source of purchase information. A <a href="https://startsomeshift.com/b2b-buying-statistics-2026-ai-research-ratio/">G2 survey of 1,076 B2B decision-makers in March 2026</a> found that 33% of buyers purchased from a vendor they had never previously heard of, discovered entirely through an AI-generated answer.</p>
<p>One in three. Discovered through AI. Not Google. Not a referral. Not a cold email.</p>
<p>If your company is not appearing in those AI-generated answers, you are not losing at the bottom of a search results page. You are losing before the conversation starts.</p>
<p>The question, then, is not whether to care about AI visibility. It is how to allocate your marketing investment between the engine you have been optimising for and the engine your buyers are now using.</p>
<h2>What SEO and AEO Actually Optimise For</h2>
<p>Most people treat these as variations of the same thing. They are not.</p>
<p><strong>SEO</strong> (Search Engine Optimisation) is the practice of making your content rank in Google's blue-link results. The signals that drive rankings are well understood: backlinks from authoritative domains, keyword relevance, page speed, mobile-friendliness, E-E-A-T signals, and internal link architecture. When someone types a query into Google, the algorithm ranks pages and serves a list. Your job is to be near the top of that list.</p>
<p><strong>AEO</strong> (Answer Engine Optimisation) is the practice of making your brand get cited inside AI-generated answers. The signals are different: structured data and schema markup, third-party citation footprint (Reddit, directories, publications), content structured to deliver a direct answer in the first 100 words, entity consistency across the web, and FAQ schema that lets LLMs pull individual Q&amp;A pairs as standalone citations.</p>
<p>When someone asks ChatGPT "what's the best spend management tool for marketing agencies," the model does not serve a ranked list of pages. It synthesises an answer and names brands it considers authoritative. Your job is to be one of those brands.</p>
<h3>The Core Distinction</h3>
<table>
<thead>
<tr>
<th></th>
<th>SEO</th>
<th>AEO</th>
</tr>
</thead>
<tbody><tr>
<td><strong>What it wins</strong></td>
<td>Google blue-link rankings</td>
<td>AI citations and recommendations</td>
</tr>
<tr>
<td><strong>Primary signals</strong></td>
<td>Backlinks, keywords, page authority</td>
<td>Structured data, citation footprint, entity consistency</td>
</tr>
<tr>
<td><strong>How buyers find you</strong></td>
<td>They click your link from a results page</td>
<td>The AI names you in its answer</td>
</tr>
<tr>
<td><strong>Content structure</strong></td>
<td>Optimised for keyword relevance</td>
<td>Optimised for direct-answer extraction</td>
</tr>
<tr>
<td><strong>Measurement</strong></td>
<td>Rankings, organic traffic, clicks</td>
<td>AI visibility score, share of voice across LLMs</td>
</tr>
</tbody></table>
<p>The overlap is real. High-quality content, E-E-A-T signals, and a well-structured site help both. But the execution is different enough that treating them as the same discipline is where most companies go wrong.</p>
<h2>Why B2B Companies Should Prioritise AEO Right Now</h2>
<p>The data on B2B buyer behaviour in 2026 makes the case clearly.</p>
<p><strong>51% of B2B buyers now start their research with AI before Google,</strong> up from just 29% seven months earlier, according to <a href="https://startsomeshift.com/b2b-buying-statistics-2026-ai-research-ratio/">G2's March 2026 Answer Economy Report</a>. That is not a trend line. That is a structural shift that happened faster than most marketing teams adjusted.</p>
<p>And the stakes are high. <a href="https://machinerelations.ai/research/b2b-ai-vendor-research-2026">Forrester research</a> found that B2B buyers use AI tools to compare vendors against each other (55% of buyers) and build internal business cases before engaging any vendor (47% of buyers). These are not peripheral tasks. They are the activities that determine whether you get on the shortlist at all.</p>
<p><a href="https://www.prnewswire.com/news-releases/73-of-b2b-buyers-use-ai-tools-in-purchase-research-multi-source-analysis-finds-302733319.html">A March 2026 analysis of 680 million citations by Averi</a> found that AI search traffic converts at 14.2%, compared to Google organic's 2.8%. That is a 5.1x conversion advantage. Claude users convert at 16.8%, ChatGPT at 14.2%, Perplexity at 12.4%. The conversion data makes the ROI case clear.</p>
<h3>Who Should Prioritise AEO First</h3>
<p>Not every business has the same urgency. Three categories have the highest immediate exposure:</p>
<p><strong>B2B SaaS.</strong> Buyers use ChatGPT and Perplexity to compare features, generate shortlists, and draft RFPs before visiting any vendor site. If your product is not being cited, you are not on the list.</p>
<p><strong>Professional services (consultants, agencies, advisors).</strong> AI citation functions as the new word-of-mouth recommendation. Being named when someone asks "best [service type] for [use case]" is the equivalent of ranking on the first page of Google five years ago.</p>
<p><strong>Fintech.</strong> Financial product buyers are conducting deep research before any vendor contact. The Gartner data cited above found buyers spend 27% of their total buying time on independent research, and just 5 to 6% of that time with any single vendor. That research is increasingly happening inside AI tools.</p>
<p><strong>The uncomfortable truth:</strong> only 22% of marketers currently track AI visibility. Your competitors probably are not optimising for this either. That window will not stay open.</p>
<h2>Where SEO and AEO Overlap (and Why You Still Need Both)</h2>
<p>AEO is not a replacement for SEO. It is an extension of it, applied to a new class of search interface where the engine synthesises answers rather than listing links.</p>
<p>The disciplines share a foundation. Content quality matters for both. E-E-A-T signals — demonstrating genuine expertise, experience, authoritativeness, and trustworthiness — matter for both. A technically broken website hurts both. These are not competing investments; they are layered ones.</p>
<h3>Where They Diverge in Practice</h3>
<p>The divergence shows up in execution. An SEO-optimised article might bury its main argument three paragraphs in after an introductory section designed to satisfy keyword density. An AEO-optimised article puts the direct answer in the first 100 words, because that is where LLMs look first. An SEO strategy pursues backlinks from high-authority domains. An AEO strategy pursues third-party citations: Reddit threads, directory listings, mentions in trade publications, and structured data that tells AI engines exactly who you are and what you do.</p>
<p><strong>The practical implication:</strong> most companies with an existing SEO programme have content that ranks but does not get cited. The articles exist. The domain authority exists. But the structure is wrong for AI extraction, the schema is missing, and the third-party footprint is thin. Fixing that is AEO work — and it builds on, rather than replaces, the SEO foundation.</p>
<blockquote>
<p><strong>The core principle:</strong> SEO earns you a seat on the results page. AEO earns you a name in the answer. For B2B buyers who have already decided before they visit your site, the name in the answer is worth more.</p>
</blockquote>
<h2>A Budget Allocation Framework for 2026</h2>
<p>The right split depends on your current state and your buyer behaviour. Here is a practical starting framework.</p>
<h3>If you have an established SEO programme and are new to AEO</h3>
<p>Your SEO foundation is an asset. Do not abandon it. But shift incremental budget toward AEO rather than doubling down on a channel your buyers are increasingly bypassing.</p>
<p>A reasonable starting allocation: <strong>60% SEO / 40% AEO</strong>, moving to <strong>50/50 within 12 months</strong> as you measure AI-sourced lead volume and adjust based on what is actually driving pipeline.</p>
<h3>If you are starting from scratch</h3>
<p>Build the AEO technical foundation first. Schema markup, structured data, entity consistency, and an llms.txt file are table-stakes infrastructure. None of it requires a developer with months of runway. Four files. A few days of work. Everything else compounds on top.</p>
<p>Then build content in parallel for both channels — but structure every piece for AI extraction first, Google ranking second. The two goals are compatible; the prioritisation just needs to shift.</p>
<h3>If your buyers are clearly using AI to find vendors</h3>
<p>You already know the answer. Go AEO-heavy now. The <a href="https://startsomeshift.com/b2b-buying-statistics-2026-ai-research-ratio/">G2 data</a> is specific: 69% of buyers chose a different vendor than they initially planned, based on AI guidance. If your buyers are in that pool and you are not being cited, you are losing deals to companies that are.</p>
<p><strong>The signal to watch:</strong> ask every new lead "how did you hear about us?" If you start seeing ChatGPT, Perplexity, or Google AI Overviews in the answers, your buyers are already using AI. You just do not have visibility into it yet.</p>
<h2>Frequently Asked Questions</h2>
<h3>Can I do AEO without doing SEO?</h3>
<p>Technically, yes. Practically, the two are hard to fully separate. A website with no domain authority, no indexed content, and no backlinks will struggle to build the citation footprint that AI engines require. That said, the AEO technical layer — schema markup, entity consistency, structured content, third-party directory placements — can be built independently of SEO and will generate AI visibility gains on its own. If you have to choose one to start, and your buyers are B2B, start with AEO. The conversion rate advantage (14.2% vs 2.8% for Google organic) makes the ROI case clear.</p>
<h3>Does AEO hurt my Google rankings?</h3>
<p>No. The changes that improve AI visibility — clearer content structure, faster direct answers, better schema markup, improved entity consistency — are also positive signals for Google. There is no known trade-off. You are not choosing between them; you are choosing how to prioritise your time and budget.</p>
<h3>How do I know if my buyers are using AI to find vendors?</h3>
<p>Three ways to check. First, add a "how did you hear about us?" field to your intake form and look for AI tools in the responses. Second, run the prompts your buyers would run — "best [category] tool for [use case]" — in ChatGPT, Perplexity, and Google AI Overviews, and see if you appear. Third, look at your <a href="https://search.google.com/search-console/">Google Search Console</a> data for traffic from AI Overviews referrals. If you are not tracking AI-sourced traffic at all, you are making budget decisions with incomplete information.</p>
<h3>Should I do AEO in-house or hire an agency?</h3>
<p>The honest answer depends on your team's existing capabilities. The technical foundation — schema, llms.txt, entity audits — can be done in-house if you have someone who understands structured data. The content and citation-building work (Reddit presence, directory placements, third-party publication mentions) is more time-intensive and benefits from specialised execution. Most B2B companies find that the technical layer is manageable in-house, while the ongoing citation footprint work is where agency support pays for itself.</p>
<h2>Where to Start</h2>
<p>The buyers who are going to contact you next month are researching right now. Some of them are asking ChatGPT which vendors to consider. Some are using Perplexity to compare your category. Some are running Deep Research reports that will determine which three companies make their shortlist.</p>
<p>If your brand is not in those answers, it does not matter how well you rank on Google.</p>
<p>The first step is knowing where you actually stand. Not guessing. Not assuming your SEO rankings translate to AI visibility. Actually checking — across ChatGPT, Perplexity, Google AI Overviews, and Gemini — which prompts your buyers are running and whether you appear.</p>
<p><strong>That is exactly what an AEO audit tells you.</strong> We run the prompts your buyers use, map your current AI visibility score, identify the gaps your competitors are filling, and give you a clear picture of what it would take to show up.</p>
<p><a href="https://windgrove.ai/">See where you stand with a free AEO audit.</a></p>
]]></content:encoded></item><item><title><![CDATA[Why Your SaaS Competitors Are Getting Cited by AI Tools (And You're Not)]]></title><description><![CDATA[TL;DR
AI tools don't cite randomly. They pull from five specific signals: structured data, third-party validation, answer-formatted content, entity consistency, and funnel-stage coverage. Most SaaS co]]></description><link>https://windgrove.hashnode.dev/why-your-saas-competitors-are-getting-cited-by-ai-tools-and-youre-not</link><guid isPermaLink="true">https://windgrove.hashnode.dev/why-your-saas-competitors-are-getting-cited-by-ai-tools-and-youre-not</guid><category><![CDATA[citations]]></category><dc:creator><![CDATA[Spencer Duke]]></dc:creator><pubDate>Wed, 17 Jun 2026 00:00:00 GMT</pubDate><enclosure url="https://cdn.hashnode.com/res/hashnode/image/upload/v1785121017559/c76c74cc-6993-4fd2-8d3d-119631861e49.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h2>TL;DR</h2>
<p>AI tools don't cite randomly. They pull from five specific signals: structured data, third-party validation, answer-formatted content, entity consistency, and funnel-stage coverage. Most SaaS companies have none of these in place. Here's what the gap looks like, and how to close it.</p>
<p>You opened ChatGPT. You typed something like "best spend management software for marketing agencies." A competitor showed up. You didn't.</p>
<p>It's not because their product is better. It's not because they have more blog posts. It's because they've built the signals AI engines look for when deciding who to cite — and you haven't.</p>
<p>That distinction matters more than almost anything else in B2B marketing right now. According to <a href="https://www.g2.com/articles/b2b-software-buyer-behavior">G2's 2025 buyer research</a>, 87% of software buyers say AI answer engines changed how they research solutions, and 51% now default to AI chat for software shortlisting. Your buyer is asking ChatGPT or Perplexity for a vendor recommendation before they visit a single website.</p>
<p><strong>If your brand isn't in that answer, you don't exist for that buyer.</strong></p>
<p>The companies getting cited haven't necessarily out-marketed you. They've out-structured you. And the gap is more specific than most founders realise.</p>
<h2>Why SaaS Companies Get Overlooked by AI Engines</h2>
<p>Most SaaS companies have built their digital presence for Google. That means keyword-optimised landing pages, feature-heavy product copy, and blog posts structured to rank for search volume. None of that maps cleanly to how AI engines decide who to cite.</p>
<p>Google ranks pages. AI engines synthesise answers. They pull from a different set of signals entirely: structured data, third-party references, question-answering content formats, and consistent entity information across the web.</p>
<p><strong>The SaaS-specific problem:</strong> product-led content doesn't answer questions.</p>
<p>A page that says "the most powerful spend management platform for growing teams" tells an AI engine nothing useful. A page that says "spend management software helps marketing agencies track client budgets in real time, reducing billing disputes and improving margin visibility" gives the AI something to work with.</p>
<p>The gap isn't content volume. Most SaaS companies have plenty of content. The gap is content structure. And underneath the content problem, there's almost always a technical problem — AI crawlers can't find the site properly, can't parse the schema, and can't establish what the company actually does.</p>
<p>The result: a competitor with a thinner product but a better-structured web presence gets cited. You don't.</p>
<h2>The 5 Signals LLMs Use to Decide Who Gets Cited</h2>
<p>AI engines don't make citation decisions arbitrarily. They weight specific signals when determining which brands are authoritative enough to recommend. Most SaaS companies are missing at least three of these five.</p>
<h3>1. Structured Data (Schema Markup)</h3>
<p>Schema markup tells AI engines what your content is about, who wrote it, what organisation it belongs to, and how to interpret it. Without it, the AI has to guess — and it often guesses wrong or skips you entirely.</p>
<p>The highest-leverage schema types for SaaS companies: FAQ schema (so individual Q&amp;A pairs can be pulled as direct citations), Organisation schema, Article schema with author attribution, and SoftwareApplication schema on product pages.</p>
<p><strong>Before AEO:</strong> No schema. The AI sees a wall of text and can't extract structured meaning.</p>
<p><strong>After AEO:</strong> FAQ schema on every key page. The AI can pull a direct answer to "what does [Product] do?" without any ambiguity.</p>
<h3>2. Answer-Formatted Content</h3>
<p>AI engines look for a direct, extractable answer in the first 100 words of a page. If your article buries the answer 600 words in, the AI moves on to a competitor whose article leads with it.</p>
<p>This is where SaaS content fails most often. Blog posts are written with a narrative arc — context, problem, solution. AI engines don't read narratively. They scan for the answer.</p>
<p><strong>Before AEO:</strong> Article opens with "In today's fast-moving business landscape..." The AI skips it.</p>
<p><strong>After AEO:</strong> Article opens with "Spend management software helps marketing agencies track client budgets across multiple campaigns in real time." The AI cites it.</p>
<h3>3. Third-Party Citation Footprint</h3>
<p>AI engines require external validation before they'll confidently recommend a brand. A company that only exists on its own website will not be cited. Full stop.</p>
<p>For SaaS companies, this means: G2 reviews, Capterra listings, Product Hunt presence, Crunchbase profile, mentions in niche trade publications, Reddit threads in relevant subreddits, and LinkedIn company page consistency. Each of these is a citation source AI engines can cross-reference.</p>
<p><strong>Before AEO:</strong> The company exists on its own domain and nowhere else credible.</p>
<p><strong>After AEO:</strong> Validated across G2, Crunchbase, and three relevant publications. AI engines have external confirmation of who you are.</p>
<h3>4. Entity Consistency</h3>
<p>AI engines cross-reference your brand's presence across the web. If your LinkedIn says one thing, your Crunchbase says another, and your website says a third, the AI's confidence in citing you drops.</p>
<p>This is an invisible problem for most companies. They've updated their positioning three times in two years and never cleaned up the old descriptions across directories and profiles.</p>
<p><strong>Before AEO:</strong> LinkedIn says "AI-powered analytics for e-commerce." Crunchbase says "data platform for retail." Website says "revenue intelligence for DTC brands." The AI is confused. You don't get cited.</p>
<p><strong>After AEO:</strong> Consistent entity description across every platform. The AI knows exactly what you do and trusts it.</p>
<h3>5. Funnel-Stage Coverage</h3>
<p>Most SaaS companies have content at the top of the funnel (awareness) and almost nothing at the bottom (purchase intent). AI engines need content at every stage to recommend you throughout the buyer journey.</p>
<p>Someone asking "what is spend management software?" needs different content than someone asking "how do I get started with [Product]?" If you only have the first and not the second, you'll get cited for awareness queries and disappear when the buyer is ready to act.</p>
<p><strong>Before AEO:</strong> 20 blog posts about industry trends. Zero content explaining how to book a demo or what onboarding looks like.</p>
<p><strong>After AEO:</strong> Full funnel coverage. The AI can recommend you at the research stage, the comparison stage, and the purchase stage.</p>
<h2>What Closing the Gap Actually Looks Like</h2>
<p>The fix isn't a content sprint. It's a sequenced programme that addresses technical infrastructure first, then content, then third-party citation building. In that order, because content published to a broken technical foundation won't get cited no matter how well it's written.</p>
<h3>Month 1: Fix What AI Crawlers Can't Read</h3>
<p>Before writing a single new piece of content, the technical layer has to be in place. This means:</p>
<p><strong>Robots.txt and llms.txt:</strong> Many SaaS sites inadvertently block AI crawlers. A misconfigured robots.txt file can make your entire site invisible to the engines you're trying to appear in. The llms.txt file is newer but increasingly important. It signals directly to LLM crawlers which content is authoritative.</p>
<p><strong>Schema implementation:</strong> FAQ schema, Organisation schema, Article schema with author attribution. This is the single highest-leverage technical change most SaaS sites can make. <a href="https://windgrove.ai/blog/the-aeo-tech-stack-4-things-every-b2b-website-needs-to-get-recommended-by-ai">Our full breakdown of the technical AEO stack</a> covers exactly what needs to be in place.</p>
<p><strong>Canonical URL resolution:</strong> www vs non-www conflicts, duplicate URL parameters, and canonical mismatches cause AI engines to treat one site as multiple separate entities. This dilutes citation authority. Most SaaS sites have at least one of these problems.</p>
<p><strong>Entity consistency audit:</strong> Every profile, directory listing, and external mention gets aligned to the same description. LinkedIn, Crunchbase, Wellfound, Product Hunt — all of them.</p>
<h3>Month 2: Publish Content AI Engines Can Actually Use</h3>
<p>Once the technical foundation is in place, content production starts. Not generic blog posts. Content mapped to specific prompts your buyers are typing into ChatGPT and Perplexity right now.</p>
<p>Each piece follows the same structure:</p>
<p><strong>Opening 100 words:</strong> a direct answer to the target question. No preamble. No "in this article we'll explore." Just the answer.</p>
<p><strong>Body:</strong> supporting evidence, context, and depth that builds credibility for the answer.</p>
<p><strong>FAQ section with schema:</strong> individual Q&amp;A pairs that AI engines can pull as standalone citations. These are often the highest-performing citation units on the page.</p>
<p>For SaaS companies specifically, the content gaps that matter most are comparison pages ("Product A vs Product B"), category definition pages ("what is [category] software and who needs it"), and BOFU content ("how to get started with [Product]"). These are the queries that appear at decision time — and most SaaS companies have zero coverage for them.</p>
<h3>Month 3: Build the Citation Footprint</h3>
<p>Content on your own site is necessary but not sufficient. AI engines require third-party validation. This is where most DIY AEO attempts stall — companies write the content, fix the technical issues, and then wonder why nothing changed. The missing piece is almost always external citation.</p>
<p>For SaaS companies, the priority citation sources are:</p>
<p><strong>G2 and Capterra:</strong> These are heavily <a href="https://windgrove.ai/blog/ai-citation-sources">cited by AI</a> engines for software evaluation queries. A well-maintained G2 profile with recent reviews is one of the fastest paths to appearing in "best [category] software" answers.</p>
<p><strong>Reddit:</strong> ChatGPT, Perplexity, and Google AI all pull significantly from Reddit. A thread on r/SaaS or a relevant vertical subreddit that mentions your product organically can drive AI citations for months.</p>
<p><strong>Niche trade publications:</strong> A mention in a publication your buyers already read carries significant weight with AI engines. Windgrove identifies which publications are currently being cited for your target prompts and pursues placements there first.</p>
<p>The compounding effect of this work is what makes it durable. <a href="https://previsible.io/">According to a Previsible AI Traffic report</a>, visitors arriving from AI-generated answers convert at 4.4x the rate of standard organic search visitors. Once you're being cited, the pipeline impact is measurable. It builds with every new piece of content and every new citation source.</p>
<p>We've run this playbook for <a href="https://windgrove.ai/proof">Opal</a>, a spend management platform for marketing agencies. Before: invisible across all major AI engines for every relevant category query. After: cited consistently in ChatGPT and Perplexity for "spend management software for agencies" and adjacent prompts, with LLM-sourced leads entering the pipeline. The work took 90 days. The results are compounding.</p>
<h2>Frequently Asked Questions</h2>
<h3>How long does it take for a SaaS company to appear in AI answers?</h3>
<p>Most clients begin seeing measurable improvements in AI citation frequency within 60 to 90 days of completing the technical foundation work. The technical changes take effect quickly. Content compounds over time. The programme becomes significantly more effective at the six-month and twelve-month marks as topical authority builds across more of the buyer's query landscape.</p>
<p>The important caveat: if the technical foundation isn't in place first, content alone won't move the needle. The sequence matters.</p>
<h3>Do I need to be on G2 or Capterra to get cited?</h3>
<p>Not strictly, but it helps significantly. G2 and Capterra are among the most heavily cited sources by AI engines for software evaluation queries. A well-maintained profile with recent reviews is one of the fastest paths to appearing in "best [category] software" answers.</p>
<p>That said, third-party citation footprint is what matters — not any single directory. A combination of G2, Crunchbase, relevant subreddit mentions, and one or two trade publication placements will often outperform a G2 profile alone. The goal is giving AI engines multiple external sources to cross-reference.</p>
<h3>What's the difference between being mentioned and being cited with a link?</h3>
<p>Being mentioned means the AI includes your brand name in a response. Being cited means the AI references your content as a source, often with a link to a specific page. Citations carry significantly more authority — they signal that your content was the trusted source for that answer, not just a brand name the AI has heard of.</p>
<p>The distinction matters for pipeline. A mention gets your name in front of a buyer. A citation tells that buyer your content is the authoritative source on the topic. That's a very different level of trust signal.</p>
<h3>What if my competitors have already built a head start?</h3>
<p>The AEO landscape in most B2B SaaS categories is still early. Most of your competitors are still running the same SEO playbook they used in 2022. Even if one or two have started building AI visibility, the category is rarely locked. <a href="https://zeroclicklabs.ai/2026-ai-seo-statistics/">According to ZeroClick Labs' 2026 AI SEO report</a>, roughly five brands capture 80% of AI responses in any given category. Those positions are still being established in most niches. The window is open. It won't be for long.</p>
<h2>The Next Step</h2>
<p>Your competitor <a href="https://windgrove.ai/blog/chatgpt-visibility">showing up in ChatGPT</a> isn't luck. It's a set of specific, fixable signals that you haven't built yet. The good news: most SaaS companies are starting from the same place, and the gap closes faster than most founders expect once the right work is in sequence.</p>
<p>If you want to see exactly where your brand stands across ChatGPT, Perplexity, and Google AI — and what it would take to close the gap — <a href="https://windgrove.ai/">book a discovery call with Windgrove</a>. We'll run a prompt audit, show you where your competitors are getting cited and why, and give you a clear picture of what the fix looks like for your specific category.</p>
<p>No strategy deck. No recommendations you have to implement yourself. We do the work.</p>
<p>You can also <a href="https://windgrove.ai/proof">see the results from our Opal engagement</a> — a before/after look at what 90 days of structured AEO work produced for a B2B SaaS company in a competitive category.</p>
<p>For a deeper look at how this plays out across complex B2B buyer journeys, <a href="https://windgrove.ai/blog/aeo-b2b-saas-pipeline">this piece on AEO for B2B SaaS pipeline</a> covers the full programme structure.</p>
]]></content:encoded></item><item><title><![CDATA[The Sites That Actually Influence AI Answers in Your Industry]]></title><description><![CDATA[Executive Summary
AI answer engines do not discover brands by crawling random websites. They pull from a concentrated set of trusted publishers, directories, review platforms, and community forums tha]]></description><link>https://windgrove.hashnode.dev/ai-citation-sources</link><guid isPermaLink="true">https://windgrove.hashnode.dev/ai-citation-sources</guid><category><![CDATA[aeo]]></category><category><![CDATA[citations]]></category><dc:creator><![CDATA[Spencer Duke]]></dc:creator><pubDate>Fri, 05 Jun 2026 00:00:00 GMT</pubDate><content:encoded><![CDATA[<blockquote>
<p><strong>Executive Summary</strong></p>
<p>AI answer engines do not discover brands by crawling random websites. They pull from a concentrated set of trusted publishers, directories, review platforms, and community forums that have already earned a place in their retrieval corpus.Every industry has its own set of "kingmaker domains": 15 to 25 sites that disproportionately shape which brands get recommended. Getting mentioned on these sites matters far more than optimising your own pages.Across 58.6 million citations analysed between October 2025 and March 2026, the top-cited domains vary sharply by industry: G2 and Reddit dominate SaaS, NerdWallet and Bankrate dominate fintech, Mayo Clinic and NIH dominate medical. The list is short. The stakes are high.Opal, an ad spend management platform for marketing agencies, went from 0% AI visibility to 15.9% and 1,766 brand mentions in 31 days. Not by rewriting their homepage, but by building a citation footprint across the right third-party sources.</p>
</blockquote>
<p>Most AEO advice tells you to optimize your website. Rewrite your headings. Add FAQ schema. Make your content more conversational.</p>
<p>That advice is not wrong. But it is incomplete, and for most brands, it is the wrong place to start.</p>
<p>Here is what the data actually shows: AI answer engines do not build their responses by crawling every website on the internet and ranking the best one. They pull from a pre-selected corpus of trusted sources. A relatively small number of publishers, directories, review sites, and community platforms that the models have learned to trust over time.</p>
<p><strong>The uncomfortable truth:</strong> your website is not in that corpus by default. Getting it there requires more than better content. It requires presence on the sites that are already inside the retrieval set.</p>
<p>This article maps those sites by industry. More importantly, it explains why they exist, how to identify them in any niche, and what it actually means for your AEO strategy.</p>
<h2>How AI Answer Engines Actually Build Their Retrieval Sets</h2>
<p>AI engines do not retrieve information the way a search engine ranks pages. The mechanics are different, and understanding them changes everything about how you approach visibility.</p>
<p>When you ask ChatGPT a question, it does not scan the entire web and return the best result. It decomposes your query into sub-queries, sends those to Bing's index, retrieves and chunks the top-ranking pages, then selects the most relevant passages from that pool. Domain authority carries roughly 40% of the weight in source selection, content quality another 35%, and platform trust signals the remaining 25%, according to citation analysis by ZipTie.dev (2025).</p>
<p>Perplexity operates differently. Every query triggers a live retrieval-augmented generation (RAG) search against a proprietary index of over 200 billion URLs. It scores sources on four factors: semantic clarity, content freshness, structural parse-ability, and entity authority. Perplexity pulls up to 40% more citations from trusted, high-authority websites compared to mid-tier blogs (<a href="https://www.capconvert.com/learn/blog/perplexity-vs-chatgpt-vs-gemini-how-each-ai-engine-sources-and-cites-differently">CapConvert, 2026</a>).</p>
<p>Google AI Overviews pull from Google's own search index and E-E-A-T signals. Gemini trusts structured first-party data. The platforms diverge on which sources they prefer — Only 11% of domains are cited by both ChatGPT and Perplexity, according to a Digital Bloom analysis of 680 million citations.</p>
<h3>What this means in practice</h3>
<p>The practical consequence is straightforward. Each AI engine has already formed a view of which sources are trustworthy in your category. Those sources form the retrieval set. Brands that appear in those sources get cited. Brands that do not, do not.</p>
<p>This is not a ranking problem. It is a presence problem.</p>
<blockquote>
<p><strong>The core principle:</strong> AI systems trust consensus across authoritative sources more than self-published website structure. A well-optimised page on a domain the model has never encountered is worth less than a single mention on a domain it already trusts.</p>
</blockquote>
<p>The brands winning in <a href="https://windgrove.ai/blog/ai-search">AI search</a> right now are not winning because they have better content on their own sites. They are winning because they have built a presence on the sites that are already inside the retrieval set for their category.</p>
<h2>What "Kingmaker Domains" Are</h2>
<p>A kingmaker domain is any site that AI engines consistently pull from when answering queries in a specific industry or category. The term is useful because it captures what these sites actually do: they determine which brands get recommended and which do not.</p>
<p>Kingmaker domains share a common set of characteristics:</p>
<ul>
<li><p><strong>High citation frequency:</strong> across multiple AI platforms, not just one</p>
</li>
<li><p><strong>Editorial independence:</strong> they are not owned by the brands they cover</p>
</li>
<li><p><strong>Category depth:</strong> they have published extensively on a specific topic or vertical for years</p>
</li>
<li><p><strong>Structural quality:</strong> their content is machine-extractable: clear headings, visible dates, author bylines, factual density</p>
</li>
<li><p><strong>Distribution footprint:</strong> their content gets syndicated, referenced, and linked to by other authority sites</p>
</li>
</ul>
<p>The data from <a href="https://higoodie.com/blog/most-cited-domains-in-llms-industry-breakdown/">Goodie AI's analysis of 58.6 million citations</a> (October 2025 to March 2026) illustrates the concentration clearly. Wikipedia leads at 3.4% citation share. YouTube at 2.1%. Reddit at 1.4%. But here is the critical finding: <strong>industry rankings diverge sharply from the overall list.</strong> TechRadar commands 8.86% citation share in B2B SaaS CRM. NerdWallet is the clear leader in retail banking and personal finance. The generic list tells you very little about what actually matters in your specific vertical.</p>
<h3>Why authority concentration has increased</h3>
<p>As AI models have matured, their source selection has become more concentrated, not less. Publishers built for breaking news have lost citation share. Platforms and review sites built for decision-making have gained it.</p>
<p>The reason is structural. AI engines are optimised to answer questions, not to surface the latest news. A site that has spent five years publishing in-depth comparisons of project management software is more useful to an LLM answering "what is the best project management tool for agencies" than a news site that covered the same topic once.</p>
<p>SE Ranking's analysis of 129,000 unique domains found that domains listed on multiple review platforms earned 4.6 to 6.3 citations on average, versus 1.8 for brands absent from those platforms. That is a 3x citation multiplier from directory presence alone.</p>
<p>This is why the kingmaker concept matters. These sites are not just high-traffic destinations. They are the infrastructure through which AI engines validate brand authority.</p>
<h2>The Kingmaker Domains by Industry</h2>
<p>The following breakdown is drawn from citation analysis across ChatGPT, Perplexity, Google AI Overviews, and Gemini. These are not the only sources AI engines pull from in each category. They are the ones that appear with the highest frequency across multiple platforms, meaning a presence on these sites gives you the broadest citation coverage for the least effort.</p>
<h3>B2B SaaS</h3>
<p>The SaaS category is dominated by structured review platforms and developer communities. AI engines treat peer-validated review data as a strong trust signal because it represents aggregated third-party opinion rather than brand-controlled content.</p>
<table>
<thead>
<tr>
<th>Domain</th>
<th>Why It Matters</th>
</tr>
</thead>
<tbody><tr>
<td>G2</td>
<td>Highest citation frequency for SaaS across all major LLMs; structured reviews with verified buyer data</td>
</tr>
<tr>
<td>Capterra</td>
<td>Gartner-owned; high editorial authority; consistently cited in "best software for X" responses</td>
</tr>
<tr>
<td>Reddit</td>
<td>r/entrepreneur, r/SaaS, r/startups pull heavily; real user language matches conversational queries</td>
</tr>
<tr>
<td>GitHub</td>
<td>Critical for dev tools; AI engines treat GitHub presence as an entity validation signal</td>
</tr>
<tr>
<td>Stack Overflow</td>
<td>Developer-facing SaaS; cited for technical questions and tool comparisons</td>
</tr>
<tr>
<td>TechCrunch</td>
<td>Funding and launch coverage; high domain authority; cited for company background and credibility</td>
</tr>
<tr>
<td>Product Hunt</td>
<td>Launch visibility; cited in "new tools" and "alternatives to X" queries</td>
</tr>
</tbody></table>
<p><strong>The SaaS insight:</strong> G2 and Capterra together account for a disproportionate share of AI citations in software categories. A brand with zero reviews on either platform is invisible in the most commercially valuable queries.</p>
<h3>Fintech and Financial Services</h3>
<p>Fintech is tightly regulated and high-trust, which means AI engines default to established financial media and comparison platforms rather than brand-controlled content. Getting cited here requires presence on the sites that have already earned editorial trust.</p>
<table>
<thead>
<tr>
<th>Domain</th>
<th>Why It Matters</th>
</tr>
</thead>
<tbody><tr>
<td>NerdWallet</td>
<td>Dominant citation source for personal finance and payments; AI engines treat it as a primary reference</td>
</tr>
<tr>
<td>Bankrate</td>
<td>Consistently cited for interest rates, financial products, and fee comparisons</td>
</tr>
<tr>
<td>Forbes Advisor</td>
<td>High citation rate for "best X" fintech queries; editorial review format maps well to AI retrieval</td>
</tr>
<tr>
<td>Investopedia</td>
<td>Cited for definitions, comparisons, and educational finance content</td>
</tr>
<tr>
<td>Business Insider</td>
<td>Fintech news and product coverage; cited for company and product credibility</td>
</tr>
</tbody></table>
<p><strong>The fintech insight:</strong> NerdWallet alone accounts for a significant share of citations in the personal finance and payments category. If your product touches consumer finance and you are not in a NerdWallet roundup, you are likely invisible to the most common buyer queries.</p>
<h3>Legal and Professional Services</h3>
<p>Law is a category where AI engines are cautious. They default heavily to established legal directories and institutional sources because the stakes of a wrong recommendation are high.</p>
<table>
<thead>
<tr>
<th>Domain</th>
<th>Why It Matters</th>
</tr>
</thead>
<tbody><tr>
<td>Justia</td>
<td>Free legal information; cited extensively for legal definitions and firm lookups</td>
</tr>
<tr>
<td>FindLaw</td>
<td>One of the most cited legal resources across all LLMs</td>
</tr>
<tr>
<td>Avvo</td>
<td>Attorney directory with peer ratings; cited for "best lawyer in X" queries</td>
</tr>
<tr>
<td>Super Lawyers</td>
<td>Peer-nominated directory; high editorial credibility in legal AI responses</td>
</tr>
<tr>
<td>Martindale-Hubbell</td>
<td>Oldest legal directory; AI engines treat long-standing directories as authority signals</td>
</tr>
</tbody></table>
<p><strong>The legal insight:</strong> Law firm websites almost never get cited directly. The citation chain goes through directories. A firm without profiles on at least three of these platforms is effectively invisible in AI-generated legal recommendations.</p>
<h3>Cybersecurity</h3>
<p>Cybersecurity is a technical category where AI engines weight vendor-neutral analysis and research heavily. Brand-controlled content is treated with more scepticism here than in almost any other vertical.</p>
<table>
<thead>
<tr>
<th>Domain</th>
<th>Why It Matters</th>
</tr>
</thead>
<tbody><tr>
<td>Gartner</td>
<td>Magic Quadrant placements are cited extensively; highest authority signal in enterprise security</td>
</tr>
<tr>
<td>Dark Reading</td>
<td>Editorial news and analysis; cited for threat intelligence and product context</td>
</tr>
<tr>
<td>SC Media</td>
<td>Industry publication; cited for product reviews and security news</td>
</tr>
<tr>
<td>CISA (.gov)</td>
<td>Government authority; cited for compliance and threat advisories</td>
</tr>
<tr>
<td>NIST (.gov)</td>
<td>Framework and standards citations; high trust for technical security queries</td>
</tr>
</tbody></table>
<p><strong>The cybersecurity insight:</strong> Government domains (.gov) carry outsized authority in security queries. A presence in CISA advisories or NIST framework documentation is worth more than most commercial placements.</p>
<h3>Medical and Healthcare</h3>
<p>Healthcare is the most authority-concentrated category in AI citations. AI engines are extremely conservative here, defaulting almost entirely to institutional medical sources. The window for brand-level citation is narrower, but it exists in the comparison and service-finder layer.</p>
<table>
<thead>
<tr>
<th>Domain</th>
<th>Why It Matters</th>
</tr>
</thead>
<tbody><tr>
<td>Mayo Clinic</td>
<td>Highest citation frequency in medical queries across all LLMs</td>
</tr>
<tr>
<td>NIH / PubMed</td>
<td>Government research authority; cited for clinical evidence and treatment options</td>
</tr>
<tr>
<td>WebMD</td>
<td>Consumer-facing health information; cited for symptom and condition queries</td>
</tr>
<tr>
<td>Healthline</td>
<td>Editorially reviewed health content; cited for wellness and treatment questions</td>
</tr>
<tr>
<td>Healthgrades</td>
<td>Provider directory; cited for "best doctor in X" and practice-level queries</td>
</tr>
<tr>
<td>Doximity</td>
<td>Physician verification; cited for professional credentials and speciality queries</td>
</tr>
</tbody></table>
<p><strong>The healthcare insight:</strong> For clinical queries, the citation chain almost never reaches a brand's own website. But for service-finder queries ("best concierge doctor in NYC"), directories like Healthgrades and Doximity are the primary citation layer. That is where brand visibility is won or lost.</p>
<h3>Medtech and Life Sciences</h3>
<table>
<thead>
<tr>
<th>Domain</th>
<th>Why It Matters</th>
</tr>
</thead>
<tbody><tr>
<td>Fierce Biotech</td>
<td>Industry news; cited for product launches and clinical trial coverage</td>
</tr>
<tr>
<td>MedTech Dive</td>
<td>Regulatory and product news; high citation rate for device and diagnostics queries</td>
</tr>
<tr>
<td>Becker's Hospital Review</td>
<td>Healthcare operations and strategy; cited for market context</td>
</tr>
<tr>
<td>FDA (.gov)</td>
<td>Regulatory authority; cited for device approvals and compliance queries</td>
</tr>
</tbody></table>
<h2>How to Identify Kingmaker Domains in Any Niche</h2>
<p>The lists above cover the most common verticals. But every niche has its own citation hierarchy. Here is a repeatable process for mapping it in any category.</p>
<p><strong>Step 1: Run the queries your buyers are running.</strong> Open ChatGPT, Perplexity, and Google AI Overviews. Ask the same question in all three: "what is the best [your category] for [your ICP]." Note every source cited in the response. Do this for 10 to 15 queries. Patterns will emerge quickly.</p>
<p><strong>Step 2: Map the overlap.</strong> The domains that appear across multiple AI platforms for multiple queries are your kingmaker domains. A site cited by only one platform for one query is a secondary signal. A site cited by all three platforms across five different queries is a primary target.</p>
<p><strong>Step 3: Assess your current presence.</strong> For each kingmaker domain you identify, check whether your brand has a listing, a review, a mention, or any form of presence. This gap analysis tells you exactly where to focus.</p>
<p><strong>Step 4: Prioritise by citation frequency and acquisition difficulty.</strong> Some kingmaker domains (G2, Capterra, Crunchbase) are accessible to any brand that creates an account. Others (Forbes Advisor, Gartner) require editorial relationships or research participation. Start with the accessible ones and build toward the harder placements.</p>
<blockquote>
<p><strong>The core principle:</strong> You do not need to be on every site. You need to be on the right 10 to 15 sites for your specific category, consistently and with enough depth that AI engines can extract and cite your brand with confidence.</p>
</blockquote>
<hr />
<h2>What This Looks Like in Practice: The Opal Case Study</h2>
<p>Opal is a charge card and spend management platform built for digital marketing agencies. When Windgrove began working with Opal, the brand had zero AI visibility across ChatGPT, Perplexity, and Google AI Overviews. Their website was well-built. Their product was strong. But they had no presence on the kingmaker domains in their category.</p>
<p>The intervention was not a website rewrite. It was a citation graph build.</p>
<p>Over 31 days, Windgrove secured placements and structured mentions across the third-party sources that AI engines consistently cite for fintech and spend management queries. The results, tracked via Searchable:</p>
<p><strong>AI Visibility Score:</strong> 0% to 15.9% in 31 days</p>
<p><strong>Brand mentions across LLMs:</strong> 1,766 tracked mentions</p>
<p><strong>Primary driver:</strong> Third-party citation footprint, not on-site content changes</p>
<img src="https://cdn.hashnode.com/uploads/posts/6a6513b3e2d2908b72dc5cf1/9fc6cb2d-380d-403b-9bf0-22b6a5025571.png" alt="AEO case study dashboard showing Opal's 30-day results: #1 LLM position, 1,236 AI mentions, 15.9% visibility trend increase." style="display:block;margin:0 auto" /><p>This is what citation graph management looks like in practice. The homepage did not change. The product did not change. What changed was where Opal existed on the web — and which of those places AI engines already trusted.</p>
<p>Read the full <a href="https://windgrove.ai/case-studies/opal">Opal case study</a> or see more results on our <a href="https://windgrove.ai/proof">proof page</a>.</p>
<hr />
<h2>The Strategic Shift: From Website Optimisation to Citation Graph Management</h2>
<p>Most AEO advice is built around a false assumption: that AI engines discover brands the same way Google does, by crawling and ranking individual pages.</p>
<p>They do not.</p>
<p>AI engines form a view of the world from their training data and retrieval corpus. That corpus is dominated by the kingmaker domains in each category. A brand that is absent from those domains is not ranked lower. It is simply not in the conversation.</p>
<p>The strategic implication is significant. The question is not "how do I make my website more AI-friendly?" The question is "which sites does the AI already trust in my category, and how do I get mentioned there?"</p>
<p><strong>The shift in practice looks like this:</strong></p>
<p><strong>Stop:</strong> Publishing more blog posts optimised for Google keywords</p>
<p><strong>Start:</strong> Identifying the 10 to 15 kingmaker domains in your category and building a presence on each</p>
<p><strong>Stop:</strong> Treating your homepage as your primary AI citation asset</p>
<p><strong>Start:</strong> Treating G2, Reddit, NerdWallet, Healthgrades (or their equivalent in your vertical) as your primary citation infrastructure</p>
<p><strong>Stop:</strong> Measuring success by organic traffic</p>
<p><strong>Start:</strong> Measuring AI visibility score and share of voice across tracked prompts</p>
<p>This is not a rejection of on-site content. Technical infrastructure, FAQ schema, and structured content still matter. But they are the floor, not the ceiling. The ceiling is determined by your citation graph.</p>
<hr />
<h2>Want to Know Which Kingmaker Domains Matter for Your Category?</h2>
<p>Identifying the right sites, getting listed on them, and building the citation footprint that moves your AI visibility score is exactly what Windgrove does. Fully managed, with results tracked in Searchable from day one.</p>
<p><a href="https://cal.com/team/windgrove-ai/discovery-call">Book a free discovery call</a> and we will map the kingmaker domains in your category and show you exactly where your brand is missing from the AI retrieval set.</p>
<hr />
<h2>Frequently Asked Questions</h2>
<p><strong>What are kingmaker domains in AI search?</strong></p>
<p>Kingmaker domains are the publishers, directories, review platforms, and community sites that AI engines repeatedly cite when answering industry questions. They matter because AI systems use source trust, repeated mentions, and consensus across independent sites to decide which brands to recommend. A brand present on these sites is far more likely to appear in AI-generated answers than a brand that only exists on its own website.</p>
<p><strong>Why does my website alone not get cited by AI engines?</strong></p>
<p>A website can be well-written and technically sound and still be ignored by AI engines if it lacks third-party validation. AI engines tend to trust consensus from external sources more than self-published pages, especially in commercial categories where reputation, reviews, and editorial coverage influence confidence. Your website signals what you say about yourself. Kingmaker domains signal what others say about you. AI engines weight the latter more heavily.</p>
<p><strong>Which industries rely most on kingmaker domains?</strong></p>
<p>B2B SaaS, fintech, legal, cybersecurity, healthcare, and medtech all show strong concentration around a small set of recurring sources. In each category, a handful of domains account for a disproportionate share of AI citations, particularly in comparison and recommendation queries. Healthcare is the most concentrated, with Mayo Clinic, NIH, and WebMD dominating clinical queries. SaaS is dominated by G2, Capterra, and Reddit.</p>
<p><strong>How do I find the kingmaker domains in my niche?</strong></p>
<p>Run the same buyer questions in ChatGPT, Perplexity, and Google AI Overviews, then record the domains cited most often across multiple queries. Aim for 10 to 15 different queries that reflect how your buyers actually search. The domains that appear across multiple platforms and multiple queries are your kingmakers. The overlap between platforms is the highest-priority target list.</p>
<p><strong>What is the fastest way to get mentioned in AI answers?</strong></p>
<p>The fastest path is not more on-site content. It is getting listed, reviewed, or mentioned on the domains that already sit inside the retrieval sets for your category. For SaaS, that means G2 and Capterra. For fintech, NerdWallet and Forbes Advisor. For healthcare, Healthgrades and Doximity. Once those placements are in place, supporting them with structured, citation-friendly content on your own site compounds the effect over time.</p>
<h2>Related reading</h2>
<ul>
<li><a href="https://windgrove.ai/blog/what-is-aeo">What Is AEO (Answer Engine Optimization) and How Does It Work?</a></li>
<li><a href="https://windgrove.ai/blog/ai-search">How AI Search Engines Recommend B2B Companies (And Why Most Are Invisible)</a></li>
<li><a href="https://windgrove.ai/blog/metrics">How We Measure Success: The Metrics Behind Your AI Visibility Program</a></li>
</ul>
]]></content:encoded></item><item><title><![CDATA[How to Check If ChatGPT Can Read Your Website (And What to Do If It Can't)]]></title><description><![CDATA[Executive Summary: ChatGPT and other AI engines cannot reliably cite your website if your technical infrastructure blocks crawlers, lacks structured data, or buries answers too deep in your content.Yo]]></description><link>https://windgrove.hashnode.dev/chatgpt-crawl-check</link><guid isPermaLink="true">https://windgrove.hashnode.dev/chatgpt-crawl-check</guid><category><![CDATA[aeo]]></category><category><![CDATA[chatgpt]]></category><dc:creator><![CDATA[Mitko Dimitrov]]></dc:creator><pubDate>Thu, 04 Jun 2026 00:00:00 GMT</pubDate><content:encoded><![CDATA[<blockquote>
<p><strong>Executive Summary:</strong> ChatGPT and other AI engines cannot reliably cite your website if your technical infrastructure blocks crawlers, lacks structured data, or buries answers too deep in your content.You can run a basic AI readability check in under 30 minutes using five free tools: your robots.txt file, Google Search Console, an <a href="https://windgrove.ai/blog/llmstxt-explained-what-it-is-why-it-matters-and-why-google-now-checks-for-it">llms.txt</a> check, a schema validator, and a direct ChatGPT prompt test.The most common reasons websites fail AI readability checks are a misconfigured robots.txt, missing FAQ schema, no sitemap submitted to Search Console, and content that never directly answers the query it targets.Windgrove offers a free AI visibility audit at <a href="https://windgrove.ai/audit">windgrove.ai/audit</a> that diagnoses exactly where your site is failing and what it would take to fix it.</p>
</blockquote>
<p>Most business owners who discover their company isn't showing up in ChatGPT assume the problem is content. They think they need more blog posts, better copy, or a PR push.</p>
<p>Sometimes that is true. But often, the issue is more basic.</p>
<p>AI crawlers cannot find your site. Or they can find it but cannot read it properly. Or they can read it but have no way to know who you are, what you do, or why you should be cited as an authority.</p>
<p><strong>The uncomfortable truth:</strong> a website that ranks on Google can still be completely invisible to ChatGPT. The two engines pull from different signals, weight different structures, and require different foundations. Your Google rankings tell you nothing about your AI visibility.</p>
<p>This guide walks you through exactly how to check whether ChatGPT can read your website, what the most common failure points look like, and what to do when you find them.</p>
<h2>Why "Readable by Google" Does Not Mean "Readable by ChatGPT"</h2>
<p>Google's crawler, Googlebot, is built to index pages for a ranked list of blue links. It follows links, reads HTML, and scores pages based on signals like backlinks, keyword density, and page authority. The output is a ranked list. The reader decides which link to click.</p>
<p>ChatGPT works differently. It synthesizes an answer. It draws from its training data, its live browsing capability (in the GPT-4o and GPT-4 Turbo models), and increasingly from real-time retrieval via tools like Bing. When a buyer asks "what's the best spend management tool for agencies," ChatGPT does not return ten links. It returns a recommendation, often with specific product names, reasons, and comparisons.</p>
<blockquote>
<p><strong>The core principle:</strong> ChatGPT is not ranking your page. It is deciding whether to name your brand in an answer. That is a fundamentally different question, and it requires a fundamentally different kind of readability.</p>
</blockquote>
<p>For your website to be cited by ChatGPT, three things need to be true:</p>
<ul>
<li><p><strong>AI crawlers can access your pages.</strong> Your robots.txt and llms.txt files must not block the crawlers that AI platforms use to retrieve content.</p>
</li>
<li><p><strong>Your content is structured for extraction.</strong> LLMs look for direct, clearly framed answers near the top of a page. Content buried in dense paragraphs does not get pulled.</p>
</li>
<li><p><strong>Your site exists beyond your own domain.</strong> AI engines require third-party validation. A brand that only lives on its own website will not be confidently cited.</p>
</li>
</ul>
<p>Most websites fail on at least one of these. Many fail on all three.</p>
<h2>Step 1: Check Your robots.txt File</h2>
<p>Your robots.txt file tells crawlers which parts of your site they are allowed to access. It was designed for Googlebot. The problem is that AI platforms use their own crawlers, and a misconfigured robots.txt can block them without you realising it.</p>
<p><strong>How to check it:</strong></p>
<ol>
<li><p>Go to <code>yourdomain.com/robots.txt</code> in your browser.</p>
</li>
<li><p>Look for any <code>Disallow:</code> rules that apply to <code>/</code> (the entire site) or to specific high-value pages like your blog, product pages, or homepage.</p>
</li>
<li><p>Check whether any <code>User-agent:</code> rules specifically block crawlers like <code>GPTBot</code> (OpenAI's crawler), <code>PerplexityBot</code>, or <code>Google-Extended</code>.</p>
</li>
</ol>
<p><strong>What a problem looks like:</strong></p>
<pre><code class="language-plaintext">User-agent: GPTBot
Disallow: /
</code></pre>
<p>That single block makes your entire website invisible to ChatGPT's live browsing feature. Many sites have this in place without the owner knowing, often because a developer added it during a site build and never removed it.</p>
<p><strong>What a clean file looks like:</strong></p>
<pre><code class="language-plaintext">User-agent: *
Allow: /

User-agent: GPTBot
Allow: /

User-agent: PerplexityBot
Allow: /
</code></pre>
<p>If you find blocking rules for AI crawlers, remove them. This is typically a one-line edit in your CMS or hosting configuration and requires no developer for most platforms.</p>
<h2>Step 2: Check Whether Your Pages Are Actually Indexed</h2>
<p>A crawler that can access your site still needs to find your pages. If your sitemap is missing or broken, AI engines may only ever see a fraction of your content.</p>
<p><strong>How to check it:</strong></p>
<ol>
<li><p>Open <a href="https://search.google.com/search-console">Google Search Console</a> and navigate to the <strong>Indexing</strong> section.</p>
</li>
<li><p>Check the <strong>Pages</strong> report. Look at how many pages are indexed versus how many are not.</p>
</li>
<li><p>Go to <strong>Sitemaps</strong> and verify that a sitemap has been submitted and is returning a success status.</p>
</li>
</ol>
<p><strong>Common problems to look for:</strong></p>
<table>
<thead>
<tr>
<th>Issue</th>
<th>What It Means</th>
</tr>
</thead>
<tbody><tr>
<td>No sitemap submitted</td>
<td>AI crawlers have no structured map of your content</td>
</tr>
<tr>
<td>Sitemap submitted but showing errors</td>
<td>Pages are being excluded from the index</td>
</tr>
<tr>
<td>Large gap between total pages and indexed pages</td>
<td>Significant content is invisible to all crawlers</td>
</tr>
<tr>
<td>Key product or service pages marked "Crawled, not indexed"</td>
<td>These pages will not be cited by AI engines</td>
</tr>
</tbody></table>
<p><strong>The Opal example:</strong> When Windgrove began working with <a href="https://windgrove.ai/case-studies/opal">Opal</a>, a spend management platform for digital marketing agencies, the site had only 4 indexed pages, no blog, and no sitemap submitted to Google Search Console. The result was 0% AI visibility across ChatGPT, Perplexity, and Google AI Overviews. Competitors were capturing every relevant query. Opal was not in the conversation at all.</p>
<p>Submitting a clean XML sitemap was one of the first fixes. It is also one of the most impactful.</p>
<h2>Step 3: Check for llms.txt</h2>
<p>llms.txt is a relatively new file, but it is becoming an important signal for AI readability. Think of it as robots.txt, but written specifically for large language models. It tells LLM crawlers which pages on your site are authoritative, which content is safe to cite, and how to navigate your site's structure.</p>
<p><strong>How to check it:</strong></p>
<p>Go to <code>yourdomain.com/llms.txt</code> in your browser.</p>
<ul>
<li><p>If you get a 404 error, the file does not exist. This is not a catastrophic failure, but it is a missed opportunity. AI crawlers that support llms.txt will have less guidance about which of your pages to prioritise.</p>
</li>
<li><p>If the file exists, check that it includes your most important pages: homepage, product or service pages, key blog content, and any pages you want AI engines to cite.</p>
</li>
</ul>
<p><strong>Why it matters:</strong></p>
<p>AI crawlers that support llms.txt use it to prioritise what to read and what to skip. A well-configured llms.txt file means your highest-value pages get more crawl attention. A missing or poorly structured file means crawlers make those decisions on their own, and they may not choose the pages you want cited.</p>
<p>Setting up llms.txt does not require a developer. It is a plain text file you can create and upload directly to your site's root directory.</p>
<h2>Step 4: Validate Your Structured Data (Schema Markup)</h2>
<p>Structured data is one of the highest-leverage signals for AI citation. Schema.org markup tells AI engines exactly what a piece of content is: who wrote it, what organisation it belongs to, what question it answers, and how to parse the information on the page.</p>
<p>Without it, AI engines are guessing. With it, they have a clear, machine-readable map of your content.</p>
<p><strong>How to check it:</strong></p>
<p>Use <a href="https://search.google.com/test/rich-results">Google's Rich Results Test</a> or the <a href="https://validator.schema.org/">Schema Markup Validator</a> to check your pages.</p>
<ol>
<li><p>Paste in your homepage URL.</p>
</li>
<li><p>Check which schema types are detected.</p>
</li>
<li><p>Repeat for your most important product or service pages.</p>
</li>
</ol>
<p><strong>What to look for:</strong></p>
<ul>
<li><p><strong>Organisation schema</strong> — Does the tool recognise your brand name, logo, and description as a defined entity?</p>
</li>
<li><p><strong>FAQ schema</strong> — Are any FAQ sections on your site marked up so LLMs can extract individual Q&amp;A pairs as standalone citations?</p>
</li>
<li><p><strong>Article schema</strong> — Do your blog posts include author attribution, publication date, and a clear subject?</p>
</li>
</ul>
<p><strong>What missing schema actually costs you:</strong></p>
<p>Without FAQ schema, an LLM reading your page has to guess which part of your content answers a given query. With FAQ schema, each question and answer is explicitly labelled. The LLM does not have to guess. It extracts the answer directly and cites your page.</p>
<p>That difference is the gap between appearing in an AI-generated answer and being skipped entirely.</p>
<blockquote>
<p><strong>Key insight:</strong> FAQ schema is the single most impactful schema type for AI citation. If your site has none, that is the first thing to fix.</p>
</blockquote>
<h2>Step 5: Ask ChatGPT Directly</h2>
<p>The most direct test is also the most revealing. Open ChatGPT (using GPT-4o with browsing enabled) and ask it the questions your buyers are actually asking.</p>
<p><strong>Prompts to try:</strong></p>
<ul>
<li><p>"What is [your brand name]?"</p>
</li>
<li><p>"What does [your brand name] do?"</p>
</li>
<li><p>"What are the best [your product category] for [your target customer]?"</p>
</li>
<li><p>"Who are the top [your service type] in [your city or industry]?"</p>
</li>
</ul>
<p><strong>What you are looking for:</strong></p>
<ul>
<li><p>Does your brand appear in the answer at all?</p>
</li>
<li><p>If it does appear, is the information accurate and current?</p>
</li>
<li><p>If it does not appear, which competitors are being recommended instead?</p>
</li>
</ul>
<h3>Interpreting the Results</h3>
<p><strong>Your brand appears with accurate information.</strong> ChatGPT can access and synthesise your content. The next question is whether you are appearing at the right funnel stages, for the right queries, and with the right positioning.</p>
<p><strong>Your brand appears but the information is outdated or wrong.</strong> ChatGPT is pulling from training data rather than live pages. This usually means your structured data is weak, your content is not clearly labelled, or your pages are not being crawled frequently enough.</p>
<p><strong>Your brand does not appear at all.</strong> This is the most common result for B2B companies that have not done any AI optimisation. It means one or more of the issues in steps 1 through 4 are blocking your visibility, or your content is not structured in a way that LLMs can extract and cite.</p>
<p><strong>Your brand appears only when you search your name directly.</strong> Branded visibility is the floor, not the goal. The real opportunity is appearing when buyers search your category without knowing your name. If you only show up for your own brand name, you are invisible to the buyers who matter most.</p>
<h2>What Fixing These Issues Actually Looks Like</h2>
<p>Running the five checks above will tell you where your site stands. What happens next depends on what you find.</p>
<p>For most B2B websites, the audit surfaces a combination of technical blockers and content structure problems. The technical issues are fixable quickly. The content structure work takes longer but compounds over time.</p>
<h3>The Opal Result: 0% to 15.9% AI Visibility in 31 Days</h3>
<p>Opal is a charge card and spend management platform for digital marketing agencies. When Windgrove audited the site in late March 2026, the picture was stark: 4 indexed pages, no blog, no sitemap, no structured data, and 0% AI visibility across every tracked prompt.</p>
<p>Competitors were capturing every relevant query in ChatGPT, Perplexity, and Google AI Overviews. Opal was not in the conversation.</p>
<p>The fix was sequenced deliberately. Technical infrastructure first: sitemap submission, robots.txt overhaul, llms.txt build, meta rewrites, heading hierarchy restructure, and indexing issue resolution. Content second: 8 AEO-optimised articles, all core pages rewritten, and bottom-of-funnel pages targeting buyers ready to convert.</p>
<p><strong>The outcome after 31 days:</strong></p>
<table>
<thead>
<tr>
<th>Metric</th>
<th>Before</th>
<th>After</th>
</tr>
</thead>
<tbody><tr>
<td>AI Visibility Score</td>
<td>0%</td>
<td>15.9%</td>
</tr>
<tr>
<td>AI Brand Mentions</td>
<td>0</td>
<td>1,766</td>
</tr>
<tr>
<td>Site Health Score</td>
<td>66.2</td>
<td>80.7</td>
</tr>
<tr>
<td>AEO Articles Live</td>
<td>0</td>
<td>8</td>
</tr>
<tr>
<td>Avg LLM Position</td>
<td>—</td>
<td>#1</td>
</tr>
</tbody></table>
<p>Within seven days of launching the first content, Opal ranked #2 for "ad pay" and #2 for "ad spend cards." Both are bottom-of-funnel terms. Buyers searching those terms are already in the market, evaluating options, and ready to move.</p>
<p>Most sites take three to six months to see meaningful movement on terms like those. Opal was there in a week.</p>
<p>You can read the <a href="https://windgrove.ai/case-studies/opal">full Opal case study here</a>.</p>
<h2>The Checklist: Your 5-Point AI Readability Audit</h2>
<p>Use this as a quick reference before running your own checks. Each item maps to one of the five steps above.</p>
<table>
<thead>
<tr>
<th>Check</th>
<th>Where to Look</th>
<th>Pass Condition</th>
</tr>
</thead>
<tbody><tr>
<td>robots.txt</td>
<td>yourdomain.com/robots.txt</td>
<td>No <code>Disallow</code> rules for GPTBot, PerplexityBot, or Google-Extended</td>
</tr>
<tr>
<td>Sitemap and indexing</td>
<td>Google Search Console &gt; Indexing</td>
<td>Sitemap submitted, key pages indexed, no major crawl errors</td>
</tr>
<tr>
<td>llms.txt</td>
<td>yourdomain.com/llms.txt</td>
<td>File exists and lists key pages</td>
</tr>
<tr>
<td>Structured data</td>
<td>Google Rich Results Test</td>
<td>Organisation schema detected; FAQ schema on relevant pages</td>
</tr>
<tr>
<td>ChatGPT visibility</td>
<td>ChatGPT with browsing (GPT-4o)</td>
<td>Brand appears in category-level queries, not just branded searches</td>
</tr>
</tbody></table>
<p><strong>If you pass all five:</strong> your technical foundation is solid. The next priority is content structure and third-party citation footprint.</p>
<p><strong>If you fail two or more:</strong> your site has structural blockers that are actively suppressing AI visibility. These need to be fixed before content or citation work will have any meaningful impact.</p>
<p><strong>If you are unsure what you found:</strong> that is what the audit is for.</p>
<h2>Get a Free AI Visibility Audit</h2>
<p>The five checks above will tell you whether your site has obvious blockers. What they will not tell you is the full picture: which prompts your buyers are using, which competitors are winning those prompts, and what specific changes would move you from invisible to cited.</p>
<p>That is what Windgrove's free AI visibility audit covers.</p>
<p><strong>The audit looks at:</strong></p>
<ul>
<li><p>Your full technical infrastructure (robots.txt, sitemap, llms.txt, schema, canonical issues)</p>
</li>
<li><p>Your current AI visibility score across ChatGPT, Perplexity, and Google AI Overviews</p>
</li>
<li><p>The specific prompts your target buyers are using and where your competitors are appearing</p>
</li>
<li><p>The highest-leverage fixes based on your current site, your category, and your competitive landscape</p>
</li>
</ul>
<p>There is no obligation and no generic deck. You will get a clear picture of where your site stands and what it would take to improve it.</p>
<p><a href="https://windgrove.ai/audit">Run your free audit at windgrove.ai/audit</a></p>
<p>If you would prefer to talk through what you find, you can also <a href="https://cal.com/team/windgrove-ai/discovery-call">book a free strategy call</a> with the Windgrove team. Most calls run 30 minutes. You will leave with a prioritised list of actions, not a sales pitch.</p>
<p>Your product deserves to be found. The question is whether the engine buyers are using can actually read your site.</p>
<h2>Frequently Asked Questions</h2>
<h3>Does ChatGPT crawl websites in real time?</h3>
<p>ChatGPT's base model (without browsing) relies on training data with a knowledge cutoff date. GPT-4o and GPT-4 Turbo with browsing enabled can retrieve live web content via Bing. For your site to be cited by ChatGPT in browsing mode, your pages must be accessible to Bing's crawler (Bingbot) and not blocked by your robots.txt. For training data inclusion, your content needs to have been publicly accessible, well-structured, and authoritative enough to be included in OpenAI's training datasets.</p>
<h3>If my site ranks on Google, why wouldn't ChatGPT know about it?</h3>
<p>Google rankings are based on backlinks, keyword relevance, and page authority signals. ChatGPT citations are based on whether your content is structured for extraction, whether your brand appears in third-party sources LLMs trust, and whether your pages are accessible to the crawlers AI platforms use. A site can rank highly on Google and still be invisible to ChatGPT if the underlying content structure does not meet AI extraction requirements.</p>
<h3>What is GPTBot and should I allow it?</h3>
<p>GPTBot is OpenAI's web crawler. It is used to retrieve content for ChatGPT's browsing feature and potentially for future training data. Allowing GPTBot in your robots.txt means ChatGPT can access and read your pages in real time. Blocking it means ChatGPT's browsing mode will not retrieve your content when answering queries related to your category. Most businesses should allow GPTBot unless they have a specific reason not to.</p>
<h3>What is llms.txt and do I need it?</h3>
<p>llms.txt is a plain text file placed at your domain root (yourdomain.com/llms.txt) that provides structured guidance to LLM crawlers about your site's content. It is analogous to robots.txt but designed specifically for AI systems rather than traditional search crawlers. It is not yet universally required, but AI platforms that support it use it to prioritise which pages to read and cite. Setting it up takes under an hour and has no downside.</p>
<h3>How long does it take to see results after fixing AI readability issues?</h3>
<p>Technical fixes (robots.txt, sitemap, schema) can take effect within days to a few weeks, depending on how quickly crawlers re-index your site. Content changes typically take longer to compound. Based on Windgrove's work with <a href="https://windgrove.ai/case-studies/opal">Opal</a>, a site that went from 0% to 15.9% AI visibility in 31 days, the timeline for meaningful results after a full technical and content overhaul is typically 30 to 90 days.</p>
<h3>Can I fix these issues myself or do I need an agency?</h3>
<p>The five checks in this guide are DIY-friendly. Fixing a robots.txt file, submitting a sitemap, and adding basic schema markup are tasks most non-developers can handle with the right instructions. Where it gets more complex is content restructuring for AI citation, building a third-party citation footprint, and tracking visibility across multiple AI engines over time. That is where a specialised AEO agency adds the most value.</p>
<h3>What is an AI visibility score?</h3>
<p>An AI visibility score measures the percentage of tracked prompts where your brand is mentioned across major LLMs (ChatGPT, Perplexity, Google AI Overviews, Gemini). A score of 0% means your brand does not appear in any of the queries your target buyers are using. A score of 15% means you appear in 15% of those queries. Windgrove tracks this using <a href="https://searchable.ai/">Searchable</a>, which provides prompt-level data including mention rate, average position, and sentiment across platforms.</p>
<h2>Related reading</h2>
<ul>
<li><a href="https://windgrove.ai/blog/chatgpt-visibility">Why Your Business Isn't Showing Up in ChatGPT (And How to Fix It)</a></li>
<li><a href="https://windgrove.ai/blog/what-is-aeo">What Is AEO (Answer Engine Optimization) and How Does It Work?</a></li>
<li><a href="https://windgrove.ai/blog/aeo-roi">Why Investing in AEO Now Pays Compounding Dividends Later</a></li>
</ul>
]]></content:encoded></item><item><title><![CDATA[HowTo vs Article Schema: Which One Gets You Cited by AI Engines?]]></title><description><![CDATA[Executive Summary  
HowTo schema and Article schema send fundamentally different signals to AI engines. Applying the wrong one to your content is one of the most common and costly structured data mist]]></description><link>https://windgrove.hashnode.dev/howto-vs-article-schema-which-one-gets-you-cited-by-ai-engines</link><guid isPermaLink="true">https://windgrove.hashnode.dev/howto-vs-article-schema-which-one-gets-you-cited-by-ai-engines</guid><category><![CDATA[citations]]></category><dc:creator><![CDATA[Spencer Duke]]></dc:creator><pubDate>Wed, 03 Jun 2026 00:00:00 GMT</pubDate><enclosure url="https://cdn.hashnode.com/res/hashnode/image/upload/v1785121025595/13a01fc3-8d5d-44ed-a086-40facdd9ea88.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<blockquote>
<p><strong>Executive Summary</strong>  </p>
<p>HowTo schema and Article schema send fundamentally different signals to AI engines. Applying the wrong one to your content is one of the most common and costly structured data mistakes.  </p>
<p>HowTo schema is built for procedural, step-by-step content. Article schema is built for editorial, informational content. The distinction is not cosmetic. It changes how AI engines extract, interpret, and cite your pages.  </p>
<p>Pages with schema markup are <a href="https://www.thehoth.com/blog/structured-data-for-ai-search/">3x more likely to earn AI citations</a> than pages without it. But generic or misapplied schema can actively underperform having no schema at all.  </p>
<p>Windgrove's AEO audits include a full structured data review across every core page, identifying missing schema, misapplied types, and the specific fixes that move your AI visibility score.</p>
</blockquote>
<p>Most content teams treat schema markup as a box to tick. Add it once, validate it in Google's Rich Results Test, and move on.</p>
<p>That approach misses the point entirely.</p>
<p>Schema markup is not just a technical formality. It is the signal that tells AI engines what kind of content they are reading, how to extract it, and whether to cite it. When you apply the wrong schema type to a piece of content, you are not just failing to benefit from structured data. You are potentially confusing the very systems you are trying to get cited by.</p>
<p><strong>The uncomfortable truth:</strong> most sites that have schema markup have the wrong schema markup. Not because they implemented it incorrectly in a technical sense, but because they applied it without understanding what each type is actually communicating to an AI engine.</p>
<p>This article breaks down the functional difference between HowTo and Article schema, explains when each applies, and shows what getting it right actually looks like in practice.</p>
<h2>Why Schema Markup Matters for AI Visibility</h2>
<p>Before getting into the specific types, it is worth understanding what schema markup actually does inside the AI citation pipeline.</p>
<p>AI engines like ChatGPT, Perplexity, and Google AI Overviews do not read your website the way a human does. They process structured signals. When your content has schema markup, you are not just decorating your HTML. You are giving the AI engine a machine-readable declaration of what your content covers, who created it, and how it is structured.</p>
<blockquote>
<p><strong>The mechanism:</strong> Schema feeds Google's Knowledge Graph and Bing's entity index. AI engines draw on those enriched indexes when generating answers. Your JSON-LD does not get parsed in real time by an LLM. It gets absorbed upstream, during indexing, and shapes how confidently an AI engine can cite your page.</p>
</blockquote>
<p>This is why the type of schema you apply matters. Different schema types send different signals about what kind of content a page contains. An AI engine retrieving a procedural answer to a "how do I" query is looking for different structural cues than one retrieving a definition or an editorial analysis.</p>
<p><strong>Schema builds clarity, not authority.</strong> That distinction is critical. Structured data alone does not make you trustworthy. It makes you legible. A credible page with correct schema gets cited more consistently. A weak page with schema gets parsed more efficiently and ignored just as quickly.</p>
<p>The largest independent study on schema and citation rates, conducted by Growth Marshal across 730 citations, found that attribute-rich schema earns a 61.7% citation rate. Generic, minimally populated schema actually underperforms having no schema at all, at 41.6% versus 59.8% for pages with no schema. The lesson is not "add schema." It is "add the right schema, fully populated."</p>
<h2>What Article Schema Actually Signals</h2>
<p>Article schema is the structured data type for editorial, informational, and long-form content. It tells AI engines: this page contains a piece of writing with a topic, an author, and a publication date. It is designed for content that explains, analyses, or informs.</p>
<p>For AEO specifically, the Article schema does two things that matter:</p>
<p><strong>It communicates content freshness.</strong> The <code>datePublished</code> and <code>dateModified</code> properties tell AI engines when the content was created and last updated. AI systems prioritize fresh, current information. A page without these signals is treated as undated, which reduces citation confidence, particularly for time-sensitive topics.</p>
<p><strong>It anchors authorship to a verifiable entity.</strong> Pairing Article schema with Person schema (via the <code>author</code> property) connects your content to a named individual with credentials. This feeds directly into E-E-A-T evaluation, the framework AI engines use to assess experience, expertise, authoritativeness, and trustworthiness.</p>
<h3>When to Use Article Schema</h3>
<p>Article schema applies to any page where the primary purpose is to inform, educate, or analyse. Common use cases include:</p>
<p>Blog posts and editorial content</p>
<p>Industry analysis and thought leadership pieces</p>
<p>News articles and announcements</p>
<p>Guides that explain a concept rather than walk through steps</p>
<p>Comparison of content and research summaries</p>
<h3>Required Properties to Populate</h3>
<p>Minimal Article schema is not enough. Populate every relevant property:</p>
<table>
<thead>
<tr>
<th>Property</th>
<th>What It Signals</th>
</tr>
</thead>
<tbody><tr>
<td><code>headline</code></td>
<td>The topic of the content</td>
</tr>
<tr>
<td><code>author</code></td>
<td>Who wrote it (link to a Person entity)</td>
</tr>
<tr>
<td><code>datePublished</code></td>
<td>When it was first published</td>
</tr>
<tr>
<td><code>dateModified</code></td>
<td>When it was last updated</td>
</tr>
<tr>
<td><code>publisher</code></td>
<td>The organization behind the content</td>
</tr>
<tr>
<td><code>description</code></td>
<td>A concise summary of the content</td>
</tr>
<tr>
<td><code>image</code></td>
<td>A representative image (required for Top Stories eligibility)</td>
</tr>
</tbody></table>
<p>Leaving these fields empty is not neutral. It is a signal that the content is either incomplete or unverifiable. Both outcomes reduce citation probability.</p>
<h2>What HowTo Schema Actually Signals</h2>
<p>HowTo schema is built for procedural content. It tells AI engines: this page contains a task that can be completed by following a defined sequence of steps. It is designed for content that guides a user through a process from start to finish.</p>
<p>The distinction from Article schema is not subtle. Article schema says "here is information." HowTo schema says "here is a process." AI engines treat these as fundamentally different content types and retrieve them for different kinds of queries.</p>
<blockquote>
<p><strong>The core principle:</strong> "How do I" queries trigger AI Overviews 73% of the time. HowTo schema is the structured data signal that tells AI engines your page is the answer to exactly those queries.</p>
</blockquote>
<p>When a user asks ChatGPT or Perplexity "how do I set up <a href="https://windgrove.ai/blog/llmstxt-explained-what-it-is-why-it-matters-and-why-google-now-checks-for-it">llms.txt</a> for my website," the AI engine is looking for ordered, extractable steps. A page with HowTo schema makes that extraction reliable. A page without it forces the AI to infer the structure from the prose, which introduces error and reduces citation confidence.</p>
<h3>When to Use HowTo Schema</h3>
<p>HowTo schema applies when the content describes a task with a defined beginning, middle, and end. Use it for:</p>
<p>Step-by-step tutorials and technical guides</p>
<p>Setup and configuration walkthroughs</p>
<p>Process documentation (onboarding flows, installation guides)</p>
<p>DIY instructions and how-to explainers</p>
<p>Any content where the primary value is in following the sequence</p>
<h3>Required Properties to Populate</h3>
<table>
<thead>
<tr>
<th>Property</th>
<th>What It Signals</th>
</tr>
</thead>
<tbody><tr>
<td><code>name</code></td>
<td>The name of the task being explained</td>
</tr>
<tr>
<td><code>step</code></td>
<td>An array of <code>HowToStep</code> objects (each with a name and text)</td>
</tr>
<tr>
<td><code>tool</code></td>
<td>Tools or software required to complete the task</td>
</tr>
<tr>
<td><code>supply</code></td>
<td>Materials or resources needed</td>
</tr>
<tr>
<td><code>totalTime</code></td>
<td>Estimated time to complete</td>
</tr>
<tr>
<td><code>estimatedCost</code></td>
<td>Cost, if applicable</td>
</tr>
</tbody></table>
<p>Each <code>HowToStep</code> should have its own <code>name</code> and <code>text</code> property. The more granular and complete the step structure, the more reliably AI engines can extract and present individual steps as standalone answers.</p>
<h3>The Common Mistake</h3>
<p>The most frequent error is applying Article schema to content that is actually procedural. A "how to implement schema markup" guide is not an article. It is a process. Marking it up as Article schema tells the AI engine it is reading editorial content, not a guide. The AI may still parse and cite the page, but it is working harder to do so, and the citation accuracy drops.</p>
<h2>HowTo vs Article Schema: Side-by-Side Comparison</h2>
<p>Here is a direct comparison of the two schema types across every dimension that matters for AEO:</p>
<table>
<thead>
<tr>
<th>Dimension</th>
<th>Article Schema</th>
<th>HowTo Schema</th>
</tr>
</thead>
<tbody><tr>
<td><strong>Content type</strong></td>
<td>Editorial, informational, analytical</td>
<td>Procedural, step-by-step, instructional</td>
</tr>
<tr>
<td><strong>Primary query match</strong></td>
<td>"What is X" / "Why does X happen"</td>
<td>"How do I X" / "How to X"</td>
</tr>
<tr>
<td><strong>Key properties</strong></td>
<td><code>headline</code>, <code>author</code>, <code>datePublished</code>, <code>dateModified</code></td>
<td><code>name</code>, <code>step</code>, <code>HowToStep</code>, <code>tool</code>, <code>totalTime</code></td>
</tr>
<tr>
<td><strong>E-E-A-T signal</strong></td>
<td>Strong (via author + publisher linkage)</td>
<td>Moderate (task-focused, less author-centric)</td>
</tr>
<tr>
<td><strong>AI extraction pattern</strong></td>
<td>Pulls topic, summary, author credentials</td>
<td>Pulls individual steps as ordered sequences</td>
</tr>
<tr>
<td><strong>Voice search fit</strong></td>
<td>Low to moderate</td>
<td>High (step-by-step answers suit voice delivery)</td>
</tr>
<tr>
<td><strong>Google rich result</strong></td>
<td>Top Stories, article carousels</td>
<td>HowTo rich results (scaled back in 2023, still signals AI)</td>
</tr>
<tr>
<td><strong>Combine with</strong></td>
<td>FAQPage, Person, Organization</td>
<td>FAQPage, VideoObject</td>
</tr>
</tbody></table>
<h3>The Grey Area: Content That Could Be Either</h3>
<p>Some content genuinely sits between the two types. A "guide to AEO for B2B SaaS" could be structured as either an editorial overview (Article) or a step-by-step implementation plan (HowTo). The right call depends on the dominant intent of the content.</p>
<p>Ask yourself: is the reader primarily trying to understand something, or trying to do something?</p>
<p>If they are trying to <strong>understand</strong>: use Article schema.</p>
<p>If they are trying to <strong>do</strong>: use HowTo schema.</p>
<p>When the content is genuinely both, Article schema is the safer default. You can always supplement it with FAQPage schema to capture the question-and-answer extraction layer. What you should not do is apply HowTo schema to content that does not have a clear sequential structure. AI engines will attempt to extract steps that do not exist, which produces inaccurate citations.</p>
<h3>You Can Use Both on the Same Page</h3>
<p>Schema types are not mutually exclusive. A page can carry Article schema at the top level and FAQPage schema for a Q&amp;A section at the bottom. A tutorial page can carry HowTo schema for the step-by-step section and Article schema for the introductory editorial content. The key is that each schema block accurately reflects the content it is marking up. Mismatches between what the schema declares and what the page actually contains reduce citation reliability across the board.</p>
<h2>What This Looks Like in Practice: The Opal Case Study</h2>
<p>Schema selection decisions do not happen in isolation. They are part of a broader technical AEO infrastructure that either enables or suppresses your content's ability to be cited.</p>
<p>When Windgrove began working with <a href="https://windgrove.ai/case-studies/opal">Opal</a>, a charge card and spend management platform for digital marketing agencies, the site had four indexed pages, no blog, and zero AI visibility. Opal was not appearing in ChatGPT, Perplexity, or Google AI Overviews for any of the queries its buyers were using.</p>
<p>The first workstream was entirely technical. Before a single piece of content was written, the team audited every blocker between Opal's pages and AI crawlers. That included a full structured data review: identifying which pages had no schema, which had schema applied incorrectly, and which needed specific types to match the content structure.</p>
<p>The results were measurable and fast.</p>
<blockquote>
<p><strong>In 31 days, Opal went from 0% AI visibility to 15.9%, accumulating 1,766 brand mentions across LLMs and the web, with zero ad spend.</strong></p>
</blockquote>
<p>Within one week of launching the Ad Pay page, Opal ranked #2 for "ad pay" and #2 for "ad spend cards." These are bottom-of-funnel, high-intent terms. The buyers searching them are not researching. They are ready to evaluate and purchase.</p>
<p><strong>The structural data work was not the only factor.</strong> The content was also written specifically for AI citation, with direct answers in the opening 100 words, proper heading hierarchy, and FAQPage schema layered alongside Article schema on editorial content. But the schema implementation was the foundation. Without it, the content would have landed on a site that AI crawlers could not reliably parse.</p>
<p>That is the real cost of getting schema wrong. It is not just a missed optimization. It is content that works hard to get written and published, then sits invisible because the technical layer was not in place to support it.</p>
<h2>The Broader Schema Stack: Where HowTo and Article Fit In</h2>
<p>HowTo and Article schema are two pieces of a larger structured data strategy. Neither operates in isolation. Understanding where they sit within the full schema stack helps you make better decisions across your entire content library.</p>
<h3>The Priority Schema Types for AEO</h3>
<table>
<thead>
<tr>
<th>Schema Type</th>
<th>Primary Purpose</th>
<th>AEO Value</th>
</tr>
</thead>
<tbody><tr>
<td><strong>FAQPage</strong></td>
<td>Marks up Q&amp;A pairs</td>
<td>Highest citation rate; mirrors conversational AI query structure</td>
</tr>
<tr>
<td><strong>HowTo</strong></td>
<td>Marks up procedural, step-by-step content</td>
<td>High; triggers on "how do I" queries (73% of AI Overviews)</td>
</tr>
<tr>
<td><strong>Article / BlogPosting</strong></td>
<td>Marks up editorial content</td>
<td>High; communicates freshness and authorship for E-E-A-T</td>
</tr>
<tr>
<td><strong>Organisation</strong></td>
<td>Declares brand identity</td>
<td>Critical for entity recognition and Knowledge Graph anchoring</td>
</tr>
<tr>
<td><strong>Person</strong></td>
<td>Declares author credentials</td>
<td>Feeds E-E-A-T; makes authorship machine-readable</td>
</tr>
<tr>
<td><strong>Product</strong></td>
<td>Marks up commercial offerings</td>
<td>High for product-led businesses; enables pros/cons extraction</td>
</tr>
</tbody></table>
<h3>How They Connect</h3>
<p>The most effective AEO schema implementations do not use these types in isolation. They connect them.</p>
<p>An editorial article should carry:</p>
<p>Article schema (with <code>headline</code>, <code>author</code>, <code>datePublished</code>, <code>dateModified</code>)</p>
<p>Person schema linked via the <code>author</code> property</p>
<p>Organization schema linked via the <code>publisher</code> property</p>
<p>FAQPage schema if there is a Q&amp;A section</p>
<p>A procedural guide should carry:</p>
<p>HowTo schema (with fully populated <code>step</code> arrays)</p>
<p>FAQPage schema for any supporting Q&amp;A content</p>
<p>Article schema if there is substantial introductory editorial content</p>
<p>This layering is what <a href="https://schema.org/">Schema.org</a> was designed for. The <code>@id</code> property allows you to connect entities across schema blocks, so your author entity is the same verifiable Person across every page on your site. That consistency matters. AI engines cross-reference entity signals. Inconsistent or disconnected schema reduces the confidence with which they cite you.</p>
<p><strong>The bottom line:</strong> schema type selection is not a one-time decision. It is a content-by-content judgement that should be built into your publishing workflow. Every piece of content you publish should have a defined schema type before it goes live, not after.</p>
<h2>Find Out Where Your Schema Is Failing</h2>
<p>Most sites have a schema problem they do not know about. Missing types on key pages. Article schema applied to procedural content. HowTo schema with empty step arrays. Publisher and author fields left blank. Each of these is a signal to AI engines that your content is either unverifiable or structurally ambiguous.</p>
<p>The good news is that schema errors are fixable. Quickly. And the citation impact is measurable.</p>
<p><strong>Start with a free AI visibility audit from Windgrove.</strong> The <a href="https://windgrove.ai/audit">Windgrove audit</a> covers your full structured data stack: which pages have schema, which types are applied, which properties are missing, and where the mismatches between schema declarations and visible content are suppressing your citation rates. You will get a clear picture of exactly what is broken and what to fix first.</p>
<p>If you want to go deeper, <a href="https://cal.com/team/windgrove-ai/discovery-call">book a free consultation</a>. Windgrove's AEO engagements start with a complete technical foundation review before any content is written or published. Schema is always Month 1 work, because content built on a weak technical foundation does not compound. It just sits there.</p>
<p>Schema markup is not the whole story. But it is the infrastructure layer that everything else depends on. Get it right, and your content has a real shot at being cited. Get it wrong, and you are publishing into a void.</p>
<p>The AI engines are not going anywhere. The question is whether your content is structured for them.</p>
<h2>Frequently Asked Questions</h2>
<h3>What is the difference between HowTo and Article schema?</h3>
<p>HowTo schema is for procedural, step-by-step content where the reader is trying to complete a task. Article schema is for editorial, informational, or analytical content where the reader is trying to understand something. The distinction changes how AI engines extract and cite your page.</p>
<h3>When should I use HowTo schema?</h3>
<p>Use HowTo schema when your content walks a reader through a defined sequence of steps from start to finish. Tutorials, setup guides, configuration walkthroughs, and installation instructions are the clearest use cases. If your content does not have real, ordered steps, do not force the schema.</p>
<h3>When should I use Article schema?</h3>
<p>Use Article schema for blog posts, guides, analysis, commentary, and thought leadership content. It works best when the page is meant to explain or inform rather than instruct. Populate the <code>author</code>, <code>datePublished</code>, and <code>dateModified</code> fields fully — these are the properties AI engines use to assess freshness and credibility.</p>
<h3>Can I use both HowTo and Article schema on the same page?</h3>
<p>Yes. A page can carry Article schema for the main editorial content and HowTo schema for a genuine step-by-step section within it. Each schema block must accurately reflect the content it marks up. Mismatches between what the schema declares and what the page actually contains reduce citation reliability.</p>
<h3>Does applying the wrong schema type hurt my AI visibility?</h3>
<p>It can. The wrong schema type tells AI engines your content is something it is not. That creates a mismatch between the query, the schema signal, and the actual content — which reduces citation confidence. Minimally populated or misapplied schema can actually underperform having no schema at all, according to <a href="https://www.thehoth.com/blog/structured-data-for-ai-search/">research across 730 AI citations</a>. Getting the type right, and populating every relevant property, is what moves the needle.</p>
]]></content:encoded></item><item><title><![CDATA[How to Structure a Product Page for ChatGPT Recommendations]]></title><description><![CDATA[Executive SummaryChatGPT does not rank pages — it synthesizes answers. A product page built for Google keyword rankings will not automatically earn AI recommendations; it requires a different structur]]></description><link>https://windgrove.hashnode.dev/how-to-structure-a-product-page-for-chatgpt-recommendations</link><guid isPermaLink="true">https://windgrove.hashnode.dev/how-to-structure-a-product-page-for-chatgpt-recommendations</guid><category><![CDATA[aeo]]></category><dc:creator><![CDATA[Spencer Duke]]></dc:creator><pubDate>Tue, 02 Jun 2026 00:00:00 GMT</pubDate><content:encoded><![CDATA[<blockquote>
<p><strong>Executive Summary</strong>ChatGPT does not rank pages — it synthesizes answers. A product page built for Google keyword rankings will not automatically earn AI recommendations; it requires a different structure entirely.Three layers determine whether ChatGPT recommends your product: technical accessibility (can the AI read your page?), content structure (does the page answer the right questions in the right format?), and third-party authority (do external sources confirm your credibility?).Pages with properly implemented FAQ schema see an 89% boost in citation probability, and content updated within 30 days receives 3.2x more <a href="https://windgrove.ai/blog/ai-citation-sources">AI citations</a> than stale pages (Moz, 2025).Windgrove helps B2B and SaaS companies implement all three layers — from schema markup and content restructuring to citation footprint building — so their product pages earn consistent recommendations across ChatGPT, Perplexity, and Google AI Overviews.</p>
</blockquote>
<p>Most product pages were built to rank. Clean URL, target keyword in the H1, meta description filled in. That is the Google playbook, and it still matters. But it is no longer enough.</p>
<p>When a buyer asks ChatGPT "What's the best project management tool for a 10-person agency?" they are not getting a list of blue links. They are getting a synthesized answer with named vendors, reasons why, and often a direct recommendation. That answer is drawn from what the AI has learned to trust — and trust, in this context, is built through structure, not just content.</p>
<p>Your product page either speaks the language ChatGPT understands, or it does not show up at all.</p>
<p>This guide covers exactly how to structure a product page so it earns a place in that answer.</p>
<h2>Why Product Pages Fail the ChatGPT Test</h2>
<p>ChatGPT does not crawl your website in real time the way Googlebot does. It pulls from what it was trained on, what its retrieval layer surfaces, and what third-party sources confirm. A product page that buries key information in JavaScript, uses vague feature-first copy, and has no external validation is effectively invisible to that process.</p>
<p>The failure modes are predictable:</p>
<ul>
<li><strong>JavaScript-rendered content</strong> that AI crawlers cannot read because the page requires script execution to load product details</li>
<li><strong>Feature-first descriptions</strong> that name capabilities without connecting them to buyer problems or use cases</li>
<li>**No FAQ schema,**categorize so individual questions cannot be extracted as standalone citation units</li>
<li><strong>Stale content</strong> with no publication date, no author attribution, and no recent updates</li>
<li><strong>Zero external footprint</strong> — the product exists only on its own website, with no third-party directories, reviews, or mentions to confirm its authority</li>
</ul>
<p><strong>The uncomfortable truth:</strong> a competitor with a thinner product but a better-structured page will get recommended before you do. ChatGPT rewards clarity and structure, not just quality.</p>
<p>The fix is not a full redesign. It is a systematic rewrite and technical layer applied to pages you already have.</p>
<h2>Layer 1: Technical Accessibility — Can ChatGPT Even Read Your Page?</h2>
<p>Before content structure or authority signals matter, the AI has to be able to read your page. This is the binary gate. If GPTBot is blocked or your product content only renders via JavaScript, nothing else you do will move the needle.</p>
<h3>Unblock AI Crawlers</h3>
<p>Check your <code>robots.txt</code> file. Many sites unintentionally block OpenAI's crawler (<code>GPTBot</code>) through blanket disallow rules or legacy configurations. Verify that the following crawlers are explicitly permitted:</p>
<ul>
<li><code>GPTBot</code> (OpenAI / ChatGPT)</li>
<li><code>ClaudeBot</code> (Anthropic)</li>
<li><code>PerplexityBot</code></li>
<li><code>Googlebot</code> (for AI Overviews)</li>
</ul>
<p>If your site uses a <code>llms.txt</code> file — a newer convention analogous to <code>robots.txt</code> but designed specifically for LLM access — ensure it signals which pages are authoritative and citable.</p>
<h3>Fix JavaScript Rendering</h3>
<p>View the page source of your product pages. If the product name, description, pricing, and key attributes do not appear in the raw HTML, your page has a JavaScript rendering problem. AI crawlers typically do not execute JavaScript. What is not in the HTML is not being read.</p>
<p>The fix: ensure all critical product content is server-side rendered and present in the initial HTML response.</p>
<h3>Use Descriptive URL Structures</h3>
<p>URL structure is a signal. <code>/products/crm-for-small-agencies</code> communicates context to AI systems. <code>/p?id=4421</code> communicates nothing. Descriptive URLs improve entity clarity and help AI systems categorize what your page is about before they even parse the content.</p>
<h3>Keep Your Sitemap Current</h3>
<p>Submit an up-to-date XML sitemap that includes every product and solution page. Missing pages in your sitemap are pages AI systems may never discover. This is a quick audit item with disproportionate impact.</p>
<blockquote>
<p><strong>Key takeaway:</strong> Technical accessibility is pass/fail. A product page that cannot be read cannot be recommended, regardless of how well the content is written.</p>
</blockquote>
<h2>Layer 2: Schema Markup — Give ChatGPT a Machine-Readable Brief</h2>
<p>Schema markup is the single highest-leverage technical intervention for AI visibility. It tells AI systems exactly what your page is about, who it is for, what it costs, and how others have rated it — in a format they can parse with confidence.</p>
<p><a href="https://moz.com/">According to a Moz study of 15,000 articles (2025)</a>, properly structured content receives 25-35% more AI citations than unstructured equivalents. Schema is a core driver of that gap.</p>
<h3>The Essential Schema Stack for Product Pages</h3>
<p>For B2B and SaaS companies, treat your product pages, pricing pages, feature comparison pages, and use-case landing pages as product pages for schema purposes. Each should carry the following:</p>
<table>
<thead>
<tr>
<th>Schema Type</th>
<th>What It Communicates</th>
</tr>
</thead>
<tbody><tr>
<td><strong>Product schema</strong></td>
<td>Name, description, brand, SKU/GTIN, image, category</td>
</tr>
<tr>
<td><strong>Offer schema</strong></td>
<td>Price, currency, availability, price validity date</td>
</tr>
<tr>
<td><strong>AggregateRating schema</strong></td>
<td>Average rating, total review count</td>
</tr>
<tr>
<td><strong>FAQPage schema</strong></td>
<td>Individual Q&amp;A pairs extractable as standalone citations</td>
</tr>
<tr>
<td><strong>Article schema</strong></td>
<td>Author attribution, publication date, last modified date</td>
</tr>
<tr>
<td><strong>Organization schema</strong></td>
<td>Brand entity, logo, contact details, social profiles</td>
</tr>
</tbody></table>
<h3>Why FAQPage Schema Matters Most</h3>
<p>FAQPage schema deserves special attention. AI engines pull individual Q&amp;A pairs from FAQ schema as direct citation units. A question and its answer can be surfaced in a ChatGPT response without the rest of your page ever being mentioned. That means each FAQ entry is its own citation opportunity.</p>
<p><strong>The format that works:</strong> each FAQ answer should be 40-60 words, self-contained, and written as a direct answer — not a teaser that requires the reader to visit the page for the full response.</p>
<h3>Implementation Format</h3>
<p>Use JSON-LD, not microdata. It is cleaner, easier to maintain, and preferred by AI systems. Place the JSON-LD block in the <code>&lt;head&gt;</code> of each page so it is present in the initial HTML response.</p>
<p>Validate your schema using <a href="https://search.google.com/test/rich-results">Google's Rich Results Test</a> before publishing. Missing required fields are the most common implementation failure.</p>
<h2>Layer 3: Content Structure — Write for the Answer, Not the Page</h2>
<p>This is where most product pages fall apart. The copy was written to persuade a human visitor. ChatGPT needs something different: a page that answers specific questions directly, early, and in a format it can extract with confidence.</p>
<h3>Lead With the Use Case, Not the Feature</h3>
<p>The most common product page mistake is feature-first copy. "AI-powered pipeline automation" tells a buyer what the product does mechanically. It does not tell ChatGPT who it is for or what problem it solves.</p>
<p>Reframe every description around the use case:</p>
<ul>
<li><strong>Instead of:</strong> "AI-powered pipeline automation with real-time dashboards."</li>
<li><strong>Write:</strong> "[Product] helps sales teams of 10-50 people close more deals without manual data entry — by automating pipeline updates and surfacing at-risk opportunities in real time."</li>
</ul>
<p>The second version answers the implicit buyer question. ChatGPT can extract it, contextualize it, and cite it in a recommendation.</p>
<h3>Front-Load the Direct Answer</h3>
<p>Every product page has an implicit question it is trying to answer. State that answer in the first 100 words. AI systems look for the direct answer near the top of the page. Content that buries the answer 500 words in does not get cited as confidently as content that states it immediately.</p>
<p><strong>The inverted pyramid applies:</strong> conclusion first, supporting detail second.</p>
<h3>Add Explicit Use-Case Headers</h3>
<p>Use H2 and H3 headings that mirror how buyers phrase their questions:</p>
<ul>
<li>"Best for freelancers"</li>
<li>"How [Product] handles client billing for agencies"</li>
<li>"Ideal for teams under 20 people"</li>
<li>"[Product] vs. manual tracking: what changes"</li>
</ul>
<p>These section headers create extraction points. When ChatGPT is asked "what's the best tool for freelance invoicing," it looks for pages with headings that directly match that intent.</p>
<h3>Build a Robust FAQ Section</h3>
<p>Every product page needs a standalone FAQ section — not just for human visitors, but as a structured citation layer for AI engines. <a href="https://www.onely.com/blog/how-to-get-your-product-discovered-by-chatgpt/">According to research cited by Onely (2026)</a>, FAQs boost citation probability by 89%.</p>
<p>The questions to include:</p>
<ul>
<li>What does [Product] do?</li>
<li>Who is [Product] designed for?</li>
<li>How does [Product] compare to [main alternative]?</li>
<li>What does [Product] cost?</li>
<li>How do I get started with [Product]?</li>
</ul>
<p>Each answer: 40-60 words, self-contained, no references to other sections of the page.</p>
<h3>Keep It Fresh</h3>
<p><strong>Content updated within 30 days receives 3.2x more AI citations than older pages</strong> (Moz, 2025). Add a visible publication date and "last updated" timestamp to every product page. Update pricing, feature lists, and FAQ answers on a regular cadence. Stale pages signal low authority to AI retrieval systems.</p>
<h2>Layer 4: Comparison Pages — Win the Evaluation Stage</h2>
<p>When a buyer asks ChatGPT "what's better, [Product A] or [Product B]," the AI draws from comparison content. If your site does not have that content, you are handing the evaluation stage to whoever does.</p>
<p>Comparison pages are one of the most underutilized assets in B2B product marketing. They serve two functions simultaneously: they rank for "[Your Product] vs [Competitor]" queries in traditional search, and they become the source ChatGPT cites when a buyer is in the final evaluation stage.</p>
<h3>What a Comparison Page Needs</h3>
<p>A comparison page built for AI citation is different from a typical "why we're better" page. It needs to be structured, balanced enough to be credible, and specific enough to be extractable.</p>
<p>Include at minimum:</p>
<ul>
<li>A clear summary of what each product does and who it is for</li>
<li>A comparison table with 5-7 dimensions (pricing, key features, integrations, support, ideal team size)</li>
<li>A section explicitly named "Who should choose [Your Product]" and "Who should choose [Competitor]"</li>
<li>Pricing for both products (updated regularly)</li>
<li>A FAQ section addressing common evaluation questions</li>
</ul>
<p><strong>The credibility principle:</strong> a comparison page that acknowledges where a competitor is stronger is more likely to be cited by ChatGPT than one that presents your product as superior on every dimension. AI systems are trained on human-generated content, and humans trust balanced assessments.</p>
<p>Build one comparison page per major competitor. Prioritize the comparisons buyers are actually making — check what queries are already driving traffic to your site, and build comparison pages for the competitors appearing in those queries.</p>
<h2>Layer 5: Third-Party Authority — The Signal ChatGPT Actually Trusts</h2>
<p>ChatGPT does not confidently recommend companies that only exist on their own websites. Third-party validation is required. This is not a soft preference — it is a structural requirement of how AI citation works.</p>
<p>AI systems cross-reference a brand's presence across the web. A company with strong on-site content but no external footprint will be cited less confidently, and often not at all, for competitive category queries.</p>
<h3>The Authority Signals That Matter</h3>
<p><strong>Review platforms:</strong> For SaaS companies, <a href="https://www.g2.com/">G2</a> and <a href="https://www.capterra.com/">Capterra</a> are among the most heavily cited sources in AI-generated software recommendations. Aim for at least 50 reviews with an average rating above 4.0. Fewer reviews signal too small a sample to trust; lower ratings actively harm recommendation eligibility.</p>
<p><strong>Directory listings:</strong> Crunchbase, LinkedIn company page, and vertical-specific directories create entity consistency signals. When AI systems see the same brand described consistently across multiple sources, citation confidence increases. Inconsistency — different descriptions, different founding dates, different team sizes — reduces it.</p>
<p><strong>Third-party publications:</strong> Mentions in industry publications, roundups, and editorial content are citation sources ChatGPT draws from directly. A single well-placed mention in a relevant publication can drive AI citations for months.</p>
<p><strong>Reddit and forum presence:</strong> Reddit is one of the most heavily cited sources across all major LLMs because it represents real human opinion at scale. A thread where your product is recommended organically, or where you answer a question helpfully, becomes a persistent citation source.</p>
<h3>Entity Consistency Is Non-Negotiable</h3>
<p>AI systems build a model of your brand from every source they can find. If your LinkedIn says one thing, your Crunchbase says another, and your website says a third, the AI's confidence in citing you drops. Audit your brand's presence across all major platforms and align the descriptions, founding information, and positioning before you invest in new content.</p>
<blockquote>
<p><strong>The core principle:</strong> your product page is the destination. Third-party sources are the signals that tell ChatGPT the destination is worth recommending.</p>
</blockquote>
<h2>The Product Page Audit: Where to Start</h2>
<p>With five layers to address, the question is sequencing. Not everything can be done at once, and some fixes unlock others. Here is the order of operations:</p>
<h3>Priority Sequence</h3>
<ol>
<li><strong>Technical accessibility first.</strong> Confirm GPTBot is not blocked in your <code>robots.txt</code>. Check that product content renders in raw HTML. Fix JavaScript rendering issues before doing anything else. This is the binary gate — nothing downstream matters until it is clear.</li>
<li><strong>Schema markup second.</strong> Implement Product, Offer, AggregateRating, and FAQPage schema on your highest-value product pages. Start with the pages that represent your most competitive queries. Validate with <a href="https://search.google.com/test/rich-results">Google's Rich Results Test</a> before moving on.</li>
<li><strong>Content restructuring third.</strong> Rewrite product descriptions to lead with use case and buyer context. Add publication dates and author attribution. Build or expand FAQ sections on every product page. This work compounds — every page you rewrite becomes a persistent citation asset.</li>
<li><strong>Comparison pages fourth.</strong> Build one comparison page per major competitor. Prioritize the comparisons buyers are already making. These pages address the evaluation stage directly and capture buyers at the point of decision.</li>
<li><strong>Third-party authority ongoing.</strong> Directory listings, review platform profiles, and publication mentions are not one-time tasks. They require consistent attention. Set a quarterly audit to check entity consistency across all external platforms.</li>
</ol>
<h3>Quick Diagnostic: Test Your Own Visibility</h3>
<p>Before starting, run this test:</p>
<ul>
<li>Open ChatGPT and ask: "What are the best [your product category] tools for [your target buyer]?"</li>
<li>Note whether your product appears, and if so, how it is characterized</li>
<li>Ask: "Tell me about [Your Product]" — what does ChatGPT know? What is missing or wrong?</li>
</ul>
<p>The answers tell you where the gaps are. If your product does not appear at all, start with technical accessibility and schema. If it appears but is described incorrectly, the content restructuring and entity consistency work is the priority.</p>
<h2>Conclusion</h2>
<p>A product page optimized for ChatGPT recommendations is not a fundamentally different page. It is the same page, rebuilt with the right technical foundation, the right content structure, and the right external signals.</p>
<p>The shift is in how you think about the reader. For Google, the reader is a human clicking a link. For ChatGPT, the reader is an AI synthesizing an answer. Both deserve a page that is clear, structured, and credible — but the specific requirements are different enough that a page built only for one will underperform with the other.</p>
<p><strong>The window is still open.</strong> Most companies in most categories have not made these changes. The brands that restructure their product pages now will build a compounding advantage that becomes harder to displace as AI-driven discovery grows.</p>
<p>Start with the diagnostic. Find out where you stand. Then work through the five layers in sequence.</p>
<p>If you would rather have this done for you, <a href="https://windgrove.ai/">Windgrove</a> handles the full execution — technical infrastructure, content restructuring, schema implementation, and citation footprint building — for B2B and SaaS companies that want to be recommended by AI, not just ranked by Google.</p>
<h2>Frequently Asked Questions</h2>
<p><strong>What does it mean to structure a product page for ChatGPT recommendations?</strong> Structuring a product page for ChatGPT means making it technically accessible to AI crawlers, implementing schema markup so the AI can parse product details in machine-readable format, writing content that leads with use-case answers rather than features, and building external authority signals that confirm your brand's credibility.</p>
<p><strong>Why is my product not appearing in ChatGPT recommendations?</strong> The most common reasons are: GPTBot is blocked in your <code>robots.txt</code>, product content only renders via JavaScript (which AI crawlers cannot execute), your page lacks schema markup, or your brand has no third-party validation footprint. Run the diagnostic described in this article to identify which layer is the root cause.</p>
<p><strong>What schema markup should I add to a product page for AI visibility?</strong> The core schema stack for product pages includes Product schema (name, description, brand, image), Offer schema (price, currency, availability), AggregateRating schema (average rating, review count), FAQPage schema (individual Q&amp;A pairs), and Article schema (author attribution, publication date). Implement in JSON-LD format and validate using Google's Rich Results Test.</p>
<p><strong>How do FAQ sections help with ChatGPT citations?</strong> FAQ sections with FAQPage schema allow AI engines to extract individual Q&amp;A pairs as standalone citation units. Each question and its answer can be surfaced in a ChatGPT response independently of the rest of the page. Research indicates FAQs boost citation probability by 89% compared to pages without them.</p>
<p><strong>How often should I update my product pages for AI visibility?</strong> Content updated within 30 days receives 3.2x more AI citations than older pages (Moz, 2025). Add a visible "last updated" timestamp and refresh pricing, feature lists, and FAQ answers on a regular cadence — at minimum quarterly, ideally monthly for your highest-priority product pages.</p>
<p><strong>Do comparison pages help with ChatGPT recommendations?</strong> Yes. When buyers ask ChatGPT to compare two products, the AI draws from comparison content on the web. A structured comparison page with a feature table, pricing for both products, and a "who should choose each" section gives ChatGPT exactly what it needs to cite your page in evaluation-stage queries.</p>
<p><strong>What third-party sources does ChatGPT trust for product recommendations?</strong> ChatGPT draws from G2 and Capterra for software reviews, industry publications for editorial mentions, Reddit and forums for community validation, and directories like Crunchbase and LinkedIn for entity confirmation. A brand that appears consistently across multiple trusted sources is cited with greater confidence than one that exists only on its own website.</p>
]]></content:encoded></item><item><title><![CDATA[How Windgrove Works: The Client Process Behind Every AI Visibility Engagement]]></title><description><![CDATA[Executive SummaryMost AEO agencies hand you a strategy deck and a list of recommendations. Windgrove does the work, writing content, fixing technical infrastructure, managing Reddit accounts, securing]]></description><link>https://windgrove.hashnode.dev/how-windgrove-works-the-client-process-behind-every-ai-visibility-engagement</link><guid isPermaLink="true">https://windgrove.hashnode.dev/how-windgrove-works-the-client-process-behind-every-ai-visibility-engagement</guid><dc:creator><![CDATA[Mitko Dimitrov]]></dc:creator><pubDate>Mon, 01 Jun 2026 00:00:00 GMT</pubDate><enclosure url="https://cdn.hashnode.com/res/hashnode/image/upload/v1785121029772/9b1e8a13-0d73-4780-9b6e-c0f114fa34df.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<blockquote>
<p><strong>Executive Summary</strong>Most AEO agencies hand you a strategy deck and a list of recommendations. Windgrove does the work, writing content, fixing technical infrastructure, managing Reddit accounts, securing directory placements, and reporting on results using <a href="https://searchable.com/">Searchable</a> data.The engagement model is built around a single principle: low touch for the client, fully managed by Windgrove. The typical client time commitment is 30 minutes per month.Every engagement follows a structured 90-day programme, Foundation, Authority, Visibility, with a clear reporting cadence: weekly Slack updates, bi-weekly progress reports, monthly strategy calls, and quarterly business reviews.Windgrove backs every 90-day engagement with a results guarantee: if your AI visibility score has not meaningfully increased and LLM-sourced leads are not entering your funnel, Windgrove works for free until they are.</p>
</blockquote>
<p>Most agencies will tell you what to do. Very few will actually do it.</p>
<p>That distinction matters more in AEO than in almost any other discipline. <a href="https://windgrove.ai/blog/what-is-aeo-answer-engine-optimization-and-how-does-it-work">Answer Engine Optimization</a>will requires simultaneous execution across technical infrastructure, content production, third-party citation building, and funnel attribution. It is not a strategy you hand off to an internal team after a kick-off call. It is a continuous, compounding programme — and the results depend entirely on the consistency of execution.</p>
<p>This article explains exactly how Windgrove runs client engagements: what the onboarding looks like, what the reporting cadence is, what we need from you, and what you can expect at the end of 90 days.</p>
<p><strong>The short version:</strong> you show up to one call, get a Slack update every Friday, and review results once a month. We handle everything else.</p>
<h2>It Starts With One Hour</h2>
<p>The first thing that happens when a new client joins Windgrove is an onboarding call. One hour. That is the only time-intensive ask we make of you in the entire engagement.</p>
<p>On that call, we map everything we need to start executing immediately:</p>
<ul>
<li><strong>Your brand</strong> — positioning, voice, what makes you different, what you want to be known for</li>
<li><strong>Your competitors</strong> — who is already winning AI citations in your category, and what they are doing to earn them</li>
<li><strong>Your target prompts</strong> — the specific questions your buyers are asking ChatGPT, Perplexity, and Google AI Overviews right now</li>
<li><strong>Your priorities</strong> — which funnel stages matter most, which markets you are focused on, and what a meaningful result looks like for your business</li>
</ul>
<p>By the end of that call, we will have everything we need. There is no back-and-forth after that. No follow-up questionnaires. No waiting for internal approvals before we can start.</p>
<p>We are moving.</p>
<h3>What We Need From You</h3>
<p>Before the onboarding call, we ask for three things:</p>
<ol>
<li><strong>Access to your CMS</strong> (Webflow, WordPress, or equivalent) for technical changes</li>
<li><strong>Access to GA4 and Google Search Console</strong> so we can audit what is already working and set up proper attribution tracking</li>
<li><strong>The ability to loop in a co-founder or key team member</strong> for content and podcast assets when relevant</li>
</ol>
<p>That is typically all that is needed. Because this is a custom engagement, there may occasionally be one or two additional requests specific to your setup. But the model is built to minimise your involvement — not to require it.</p>
<p><strong>The channels we use:</strong> Slack, Google Meet, and Gmail. No new tools to learn. No project management platforms to navigate.</p>
<h2>The Reporting Cadence: What You Hear From Us and When</h2>
<p>One of the most common frustrations with agency relationships is not knowing what is happening. You send a brief, someone disappears for two weeks, and then you get a report that raises more questions than it answers.</p>
<p>Windgrove is structured to prevent that entirely. The cadence is predictable, the reporting is specific, and you never have to chase us down for a status update.</p>
<table>
<thead>
<tr>
<th>Frequency</th>
<th>Format</th>
<th>What It Covers</th>
</tr>
</thead>
<tbody><tr>
<td>Weekly (every Friday)</td>
<td>Slack message</td>
<td>What was executed that week, what is live, what is coming next</td>
</tr>
<tr>
<td>Bi-weekly</td>
<td>Written report</td>
<td>Visibility movement, trajectory, Searchable prompt-level data, content published</td>
</tr>
<tr>
<td>Monthly</td>
<td>30-minute Google Meet</td>
<td>Full Searchable data review, funnel attribution, next 30-day prioritisation</td>
</tr>
<tr>
<td>Quarterly</td>
<td>Deeper session</td>
<td>ROI review, compounding results, market expansion strategy</td>
</tr>
</tbody></table>
<h3>The Friday Slack Update</h3>
<p>Every Friday, you receive a detailed Slack message covering exactly what happened that week. Not a vague summary. Specific: which articles went live, which technical fixes were deployed, which directory placements were secured, and what is scheduled for the following week.</p>
<p>This keeps you informed without requiring your involvement. You can read it in two minutes and move on with your day.</p>
<h3>The Bi-Weekly Progress Report</h3>
<p>Every two weeks, we share a more detailed written report that shows visibility movement across your tracked prompts. This includes Searchable data, prompt-level changes, and a breakdown of what content went live. You can see exactly which prompts are moving, which are holding, and where we are focusing next.</p>
<h3>The Monthly Strategy Session</h3>
<p>The monthly call is 30 minutes. We walk through your full <a href="https://searchable.com/">Searchable</a> data together: what moved, what is compounding, and what we are prioritising in the next sprint.</p>
<p>We also review your HubSpot and Typeform data to track LLM-sourced leads moving through your pipeline — from waitlist join to consultation booked to introductory call. The goal of this call is not to discuss visibility numbers in isolation. It is to connect what we are doing to booked consultations and closed pipeline.</p>
<h3>The Quarterly Business Review</h3>
<p>Once per quarter, we run a deeper session covering compounding results, full ROI analysis, and your evolving AI search strategy. For clients expanding into new markets, this is where we align on what needs to be built before those markets go live — not reactively after.</p>
<p><strong>The total time commitment on your end: approximately 30 minutes per month.</strong> Everything else is handled by Windgrove.</p>
<h2>The 90-Day Programme: Foundation, Authority, Visibility</h2>
<p>Windgrove engagements begin with a 90-day commitment. The structure exists for a specific reason: AI visibility changes take time to compound. The technical foundation has to be in place before content can be cited. Content has to accumulate before citation patterns shift meaningfully. Trying to compress that timeline produces noise, not results.</p>
<p>The 90 days are divided into three distinct phases.</p>
<h3>Month 1: Foundation</h3>
<p>No content production starts until the site can be properly read and cited by AI engines. This is non-negotiable.</p>
<p>Month 1 is entirely technical. Windgrove audits the site, identifies every blocker between your pages and AI crawlers, and resolves them before a single article is published. The work includes:</p>
<ul>
<li>Schema.org structured data implementation (FAQ schema, Organisation schema, Article schema with author attribution)</li>
<li>robots.txt and llms.txt configuration</li>
<li>XML sitemap optimisation and submission</li>
<li>Canonical URL resolution</li>
<li>Reformatting of existing articles to front-load direct answers</li>
<li>Google Business Profile optimisation</li>
<li>Entity consistency audit across LinkedIn, Crunchbase, Apple Maps, Bing Places, and relevant vertical directories</li>
<li>Baseline Searchable report and funnel attribution setup</li>
</ul>
<p><strong>The deliverable at the end of Month 1:</strong> a site that AI crawlers can fully read, index, and cite — with a documented baseline showing exactly where visibility stands before content begins.</p>
<h3>Month 2: Authority</h3>
<p>With the foundation solid, Month 2 shifts to content production and citation footprint expansion.</p>
<p>Windgrove publishes four new AEO-optimised articles per week during this phase. Each article is mapped to a specific tracked prompt. The structure is deliberate: the opening 100 words contain a direct answer to the target question, the body provides supporting depth, and a FAQ section with schema markup allows LLMs to pull individual Q&amp;A pairs as standalone citations.</p>
<p>Alongside content production, Month 2 activates the Reddit presence — including account warmup and engagement in relevant subreddits — and secures the first wave of directory placements. Existing content is fully rewritten to match AEO citation mechanics.</p>
<p>To understand what this looks like in practice, the <a href="https://windgrove.ai/case-studies/opal">Opal case study</a> is the clearest example. Opal went from zero AI visibility to a 15.9% AI Visibility Score and 1,766 brand mentions in 31 days. The first eight AEO articles went live during this phase. Within seven days of publishing the Ad Pay page, Opal ranked #2 for "ad pay" and #2 for "ad spend cards" — both bottom-funnel terms with direct commercial intent.</p>
<h3>Month 3: Visibility</h3>
<p>Month 3 is about dominating the conversations your target customers are already having.</p>
<p>The work expands to include YouTube setup and video syndication, comparison pages targeting "[your brand] vs [competitor]" queries, podcast guest pipeline activation, and emerging topic capture. For clients expanding into new markets, location pages and local entity profiles are built before those markets go live.</p>
<p><strong>Target outcomes at the end of 90 days:</strong></p>
<ul>
<li>5 to 10 percentage point increase in AI Visibility Score</li>
<li>3 to 5 percentage point increase in Share of Voice above 10%</li>
<li>10 or more LLM-attributed leads entering the funnel</li>
</ul>
<p>After 90 days, engagements move to month-to-month continuation at the same rate, with scope adjusted based on what the data shows.</p>
<h2>What the Results Look Like in Practice</h2>
<p>The Opal engagement is the most complete picture of what this process produces.</p>
<p>When Windgrove began working with Opal in late March 2026, the product was strong but AI search visibility was non-existent. The site had four indexed pages, no blog, weak metadata, and no sitemap in Google Search Console. Anyone searching for Opal's category in ChatGPT, Perplexity, or Google AI Overviews was finding only competitors.</p>
<p>Thirty-one days later:</p>
<ul>
<li><strong>AI Visibility Score:</strong> 0% to 15.9%</li>
<li><strong>Brand mentions across LLMs:</strong> 0 to 1,766</li>
<li><strong>Site health score:</strong> 66.2 to 80.7 (top 10% of benchmarked sites)</li>
<li><strong>AEO articles live:</strong> 0 to 8</li>
</ul>
<p>Opal is now surfacing as a named recommendation inside Perplexity AI responses. When a buyer searches "what cards allow you to pay ad invoices with a credit card," Opal Ad Pay appears in the answer — not as an ad, not as a sponsored result, but as the recommended solution with a direct link to the product page.</p>
<p><img src="https://cdn.hashnode.com/uploads/posts/6a6513b3e2d2908b72dc5cf1/37da5a47-ef99-42b7-b195-09466470740c.png" alt="AEO case study dashboard showing LLM ranking improvement to #1 position with AI visibility trend chart and performance metrics." /></p>
<p>That is what the Foundation and Authority phases produce when executed in sequence. The 15.9% AI Visibility Score at 31 days is a starting point. As content compounds and citation frequency increases, that number climbs. You can read the <a href="https://windgrove.ai/case-studies/opal">full Opal case study here</a>.</p>
<p>For a deeper look at how Windgrove measures progress across every engagement, the <a href="https://windgrove.ai/blog/how-we-measure-success-the-metrics-behind-your-ai-visibility-program">metrics breakdown</a> covers the full Searchable data model — including how AI Visibility Score, Share of Voice, and funnel-stage visibility are tracked and reported.</p>
<h2>The Guarantee</h2>
<p>Most agencies ask you to trust the process. Windgrove backs it with a guarantee.</p>
<blockquote>
<p>If at the end of 90 days your AI Visibility Score has not meaningfully increased and LLM-sourced leads are not entering your funnel, Windgrove works for free until they are. No awkward conversations. No renegotiation. We eat the cost until we earn it.</p>
</blockquote>
<p>This guarantee exists because the model is built around outcomes, not activity. A Friday Slack update that lists tasks completed is not a result. A bi-weekly report that shows visibility movement is not a result. The result is a buyer finding your product in ChatGPT when they were not looking for you by name — and taking the next step.</p>
<p><strong>Why the 90-day window?</strong> Because that is how long it takes for the compounding mechanics to become visible in the data. Month 1 fixes the foundation. Month 2 builds the content and citation footprint. Month 3 is where the visibility shifts show up clearly in Searchable — and where LLM-sourced leads begin entering the funnel in measurable numbers.</p>
<p>The guarantee is not a hedge. It is a reflection of how confident Windgrove is in the process when it is executed in full.</p>
<p>For a broader perspective on why the timing of AEO investment matters — and what happens to brands that wait — the article on <a href="https://windgrove.ai/blog/aeo-benefits-why-timing-matters">why investing in AEO now pays compounding dividends</a> is worth reading before you make a decision.</p>
<h2>Where to Start</h2>
<p>If your product is strong but buyers are not finding you in ChatGPT, Perplexity, or Google AI Overviews, the gap is usually smaller than it looks. And it almost always starts with the same two things: a technical foundation that AI crawlers cannot read, and content that was built for Google rather than for citation.</p>
<p>The fastest way to understand where you stand is a free AI visibility audit. Windgrove will review your current visibility across tracked prompts, identify the highest-leverage gaps, and show you what a 30/60/90-day AEO programme looks like for your specific situation.</p>
<p><a href="https://windgrove.ai/audit"><strong>Run your free AI visibility audit at windgrove.ai/audit</strong></a></p>
<p>If you would rather talk through the process first, you can <a href="https://cal.com/team/windgrove-ai/discovery-call">book a free consultation</a> directly. No obligation. No generic deck. Just a clear-eyed look at where your AI visibility stands and what it would take to move it.</p>
<p>The brands that are winning AI citations today are not necessarily the ones with the best products. They are the ones that built the right infrastructure first. That window is still open — but it is narrowing every quarter as more companies recognise what is at stake.</p>
<p>If you want to understand what questions to ask before committing to any AEO agency, the article on <a href="https://windgrove.ai/blog/questions-to-ask-an-aeo-agency">10 questions to ask an AEO agency before you sign</a> is a useful place to start.</p>
<h2>Frequently Asked Questions</h2>
<h3>How much time does working with Windgrove actually require?</h3>
<p>The typical client commitment is approximately 30 minutes per month — the monthly strategy call. Beyond that, the onboarding call (one hour, once) and occasional requests for access or approvals specific to your setup. Windgrove handles all execution: content production, technical fixes, Reddit management, directory placements, reporting, and attribution tracking.</p>
<h3>What does the onboarding call cover and what happens immediately after?</h3>
<p>The onboarding call is one hour. Windgrove maps your brand positioning, key competitors, target prompts, and priorities. By the end of the call, there is enough information to begin executing immediately. No follow-up questionnaires, no waiting for internal sign-off. Month 1 technical work begins within days of the call.</p>
<h3>How does Windgrove measure whether the engagement is working?</h3>
<p>Windgrove uses <a href="https://searchable.com/">Searchable</a> as the primary measurement platform. It tracks AI Visibility Score (the percentage of tracked prompts where your brand is mentioned across LLMs), Share of Voice against named competitors, prompt-level data including position and sentiment, and funnel-stage visibility segmented by TOFU, MOFU, and BOFU. Secondary metrics include LLM-sourced leads entering the funnel, tracked via "how did you hear about us" attribution and CRM pipeline data.</p>
<h3>What happens if results do not materialise within 90 days?</h3>
<p>Windgrove works for free until they do. If at the end of 90 days the AI Visibility Score has not meaningfully increased and LLM-sourced leads are not entering the funnel, the engagement continues at no charge until those outcomes are achieved. No renegotiation, no awkward conversations.</p>
<h3>Is Windgrove a fit if we already have an SEO agency?</h3>
<p>Yes. Windgrove is not an SEO agency and does not compete with traditional search optimisation work. <a href="https://windgrove.ai/blog/what-is-aeo-answer-engine-optimization-and-how-does-it-work">AEO and SEO operate on different signals</a> — content that ranks on Google does not automatically get cited by AI engines, and vice versa. Windgrove's work runs in parallel with existing SEO programmes and typically complements rather than conflicts with them.</p>
]]></content:encoded></item><item><title><![CDATA[Is AEO Actually Working, or Is It Just SEO with a New Name?]]></title><description><![CDATA[Executive SummaryAI answer engines (ChatGPT, Perplexity, Google AI Overviews) now handle over 35% of all search queries, a 120% increase from 2023 levels, and they pull from entirely different signals]]></description><link>https://windgrove.hashnode.dev/is-aeo-actually-working-or-is-it-just-seo-with-a-new-name</link><guid isPermaLink="true">https://windgrove.hashnode.dev/is-aeo-actually-working-or-is-it-just-seo-with-a-new-name</guid><dc:creator><![CDATA[Mitko Dimitrov]]></dc:creator><pubDate>Fri, 29 May 2026 00:00:00 GMT</pubDate><enclosure url="https://cdn.hashnode.com/res/hashnode/image/upload/v1785121032130/dd9db492-ca8c-4a17-a36f-92d7fc705b11.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<blockquote>
<p><strong>Executive Summary</strong>AI answer engines (ChatGPT, Perplexity, Google AI Overviews) now handle over 35% of all search queries, a 120% increase from 2023 levels, and they pull from entirely different signals than traditional search.AEO is not SEO with a new coat of paint. It requires a distinct technical foundation, different content structure, and a third-party citation footprint that Google rankings do not provide.The data is clear: brands implementing AEO alongside SEO report up to 43% increases in brand visibility and 37% more qualified leads compared to SEO-only approaches.Windgrove has delivered measurable AI visibility results for clients in fintech, professional services, and SaaS, taking one client from 0% AI visibility to a #1 average LLM position in 31 days.</p>
</blockquote>
<hr />
<p>Every few years, a new acronym lands in the marketing world and the skeptics ask the same question: is this real, or is it just the last thing repackaged?</p>
<p>AEO is getting that question right now. <a href="https://windgrove.ai/blog/what-is-aeo">Answer Engine Optimization</a>. Optimize for AI. Get cited by ChatGPT. Show up in Perplexity. It sounds, to a reasonable person, like SEO with a rebrand and a fresh coat of hype.</p>
<p>The skepticism is understandable. But it is wrong.</p>
<p>AEO is not a renaming exercise. The mechanics are genuinely different, the signals are different, and the results, when the work is done properly, are measurable and distinct from anything SEO alone produces. This article makes that case with data, explains exactly where the two disciplines diverge, and shows what AEO looks like when it works in practice.</p>
<p><strong>The core question:</strong> Does optimizing for AI answer engines produce outcomes that SEO cannot? The answer is yes. Here is why.</p>
<h2>The Fundamental Difference: Ranking vs. Synthesizing</h2>
<p>Google ranks pages. AI engines synthesize answers.</p>
<p>That single sentence captures the entire distinction. When someone types "best spend management tool for agencies" into Google, they receive a list of links and decide where to click. When they ask the same question to ChatGPT or Perplexity, they receive a personalized, conversational recommendation, often with specific vendor names, reasons why, and comparisons, all generated from the AI's internal understanding of the market.</p>
<p>The buyer never visits ten websites. They read one answer and often act on it.</p>
<blockquote>
<p><strong>The uncomfortable truth:</strong> AI engines do not rank your page. They synthesise an answer and name the brands they consider authoritative. If your brand is not in that answer, you are invisible to a buyer who has already made up their mind before they visit a single website.</p>
</blockquote>
<p>This is not a subtle distinction. It changes the entire logic of how you earn visibility.</p>
<h3>What Google rewards vs. what AI engines reward</h3>
<table>
<thead>
<tr>
<th>Signal</th>
<th>Traditional SEO</th>
<th>AEO</th>
</tr>
</thead>
<tbody><tr>
<td>Primary ranking factor</td>
<td>Backlinks + keyword relevance</td>
<td>Entity authority + citation footprint</td>
</tr>
<tr>
<td>Content structure</td>
<td>Keyword density, headings, meta</td>
<td>Direct answer in first 40-60 words</td>
</tr>
<tr>
<td>Technical requirements</td>
<td>Crawlability, page speed, Core Web Vitals</td>
<td>Schema markup, llms.txt, canonical clarity</td>
</tr>
<tr>
<td>Third-party validation</td>
<td>Backlinks from authority domains</td>
<td>Mentions on Reddit, directories, publications</td>
</tr>
<tr>
<td>Measurement</td>
<td>Rankings, organic traffic</td>
<td>AI visibility score, share of voice, LLM citations</td>
</tr>
</tbody></table>
<p>A page can rank on page one of Google and be completely invisible inside an AI-generated answer. The reverse is also true: a well-structured page with strong entity signals and a solid citation footprint can appear in AI responses without ranking anywhere near the top of traditional search.</p>
<p><a href="https://blog.hubspot.com/marketing/answer-engine-optimization-trends">According to HubSpot's 2026 Consumer Trends Report</a>, the average prompt length in ChatGPT is 23 words, compared to 3.37 words in traditional search. That is not a user typing a keyword. That is a user having a conversation. And the AI engine pulling from your content needs to understand your brand as an entity, not just a collection of keyword-optimized pages.</p>
<p>That is the structural difference. And it is why SEO work, done without AEO in mind, does not transfer.</p>
<h2>The Data Behind AEO's Results</h2>
<p>The skeptic's next move is to ask for proof. Fair enough. Here is what the research shows.</p>
<p><a href="https://www.wavesandalgorithms.com/intelligence-reports/answer-engine-optimization">A meta-analysis of 27 AEO case studies</a> published between 2024 and 2025 found weighted average improvements across the following metrics:</p>
<ul>
<li><strong>AI citation frequency:</strong> +213%</li>
<li><strong>Voice search visibility:</strong> +167%</li>
<li><strong>Conversion rate from AI referrals:</strong> +28%</li>
<li><strong>Brand mention sentiment:</strong> +42%</li>
</ul>
<p>These are not small movements. A 213% increase in how often an AI engine cites your brand is the difference between being invisible and being the default recommendation.</p>
<p>HubSpot's own internal data is equally direct. Their marketing team reported <strong>3x better lead conversion from AEO-sourced traffic</strong> compared to other acquisition channels. The reason is structural: AI engines personalize answers based on what they know about the user's context, so the traffic that arrives from an AI citation is pre-qualified in a way that a Google click rarely is.</p>
<p><strong>The zero-click reality:</strong> According to <a href="https://medium.com/@snargunde/answer-engine-optimization-aeo-the-future-of-search-in-the-ai-era-2025-insights-4f48f8cb7ae0">BrightEdge's March 2025 research</a>, over 60% of search queries now generate AI-powered answers that reduce direct website clicks. Google's AI Overviews alone answer more than 50% of queries without requiring a click. A Bain and Company analysis puts the economic stakes plainly: zero-click search behaviour could impact $750 billion in revenue by 2028.</p>
<p>That is not a trend to monitor. That is a shift to act on.</p>
<h3>Why the "it's just SEO" argument falls apart</h3>
<p>The strongest version of the skeptic's argument goes like this: AEO is just good SEO. Write quality content, earn authority, get cited. Nothing new.</p>
<p>There is some truth here. Good content remains foundational. But the argument collapses on three specific points:</p>
<ol>
<li><strong>Technical infrastructure is different.</strong> SEO does not require an llms.txt file, FAQ schema on every applicable page, or canonical URL resolution tuned for LLM crawlers. These are AEO-specific interventions. A technically sound SEO site can still be completely unreadable by AI engines.</li>
<li><strong>Citation footprint is different.</strong> AI engines require third-party validation from sources they trust: Reddit threads, directory listings, publications. A company with strong backlinks but no Reddit presence and no directory listings will not be confidently cited by an LLM, regardless of its domain authority.</li>
<li><strong>Content structure is different.</strong> Google rewards comprehensive, well-linked content. AI engines reward content that delivers a direct answer in the first 40-60 words. Content structured for Google often buries the answer. That content does not get cited. Research confirms that answer engines demonstrate a measurable 2.7x preference for content that directly addresses user questions in the opening text.</li>
</ol>
<p>Three distinct requirements. None of them is SEO.</p>
<h2>What AEO Actually Looks Like in Practice: The Opal Case</h2>
<p>Theory is useful. Numbers from a real engagement are more useful.</p>
<p>In late March 2026, Opal, a charge card and spend management platform built for digital marketing agencies, had a strong product and zero AI search visibility. Their site had four indexed pages, no blog, no sitemap submitted to Google Search Console, and a 0% AI visibility score across ChatGPT, Perplexity, and Google AI Overviews. Buyers searching for their category were finding only competitors.</p>
<p><a href="https://windgrove.ai/case-studies/opal">The full Opal case study is documented here</a>. The summary is this: in 31 days, the numbers moved from zero to the following.</p>
<p><img src="https://cdn.hashnode.com/uploads/posts/6a6513b3e2d2908b72dc5cf1/8f64b8ed-0fc6-41bc-bbbc-911d0770ec34.jpg" alt="Windgrove Opal case study results showing 0 to 15.9% AI visibility score and 1766 brand mentions in 31 days" /></p>
<ul>
<li><strong>AI Visibility Score:</strong> 0% to 15.9%</li>
<li><strong>AI Brand Mentions:</strong> 0 to 1,766</li>
<li><strong>Average LLM Position:</strong> #1</li>
<li><strong>Site Health Score:</strong> 66.2 to 80.7 (top 10% of benchmarked sites)</li>
<li><strong>AEO Articles Published:</strong> 8</li>
</ul>
<h3>How the results were produced</h3>
<p>The work followed a deliberate sequence. Technical foundation first. Content second.</p>
<p>Before a single article was written, every blocker between Opal's pages and AI crawlers was cleared. The XML sitemap was submitted. The robots.txt was rewritten. An llms.txt file was built from scratch. Meta titles and descriptions were rewritten across every core page. Heading hierarchy and internal linking were restructured. Indexing issues on key products and landing pages were resolved.</p>
<p>Only once the foundation was solid did content production begin. Every article was mapped to a specific buyer query, filtered through three criteria: search volume, buyer intent, and competitive opportunity. The result was eight published articles and a set of bottom-of-funnel pages, all structured so AI engines could read, understand, and cite them.</p>
<p>Within one week of launching the Ad Pay page, Opal ranked #2 for "ad pay" and #2 for "ad spend cards." These are not informational terms. Businesses searching for those phrases are in the market and ready to move.</p>
<p><strong>The compounding effect:</strong> The 15.9% AI visibility score reached in 31 days is not a ceiling. It is a starting point. Each article published from a clean technical foundation adds a new surface for AI engines to cite. The 1,766 brand mentions accumulated in month one were earned, not paid. Zero ad spend. Pure citation.</p>
<p>This is what AEO produces that SEO alone cannot. Not just traffic. Authoritative presence inside the conversations buyers are already having.</p>
<p>You can see more results like this on the <a href="https://windgrove.ai/proof">Windgrove proof page</a>.</p>
<h2>The Five Levers That Separate AEO from SEO</h2>
<p>If AEO were just SEO, you could apply the same playbook and expect the same results. You cannot. The levers are different.</p>
<p>Here are the five interventions that move AI visibility scores and have no direct equivalent in traditional SEO:</p>
<h3>1. Answer-first content structure</h3>
<p>AI engines look for a direct answer in the first 40 to 60 words of a page. If your content spends the opening paragraphs building context before delivering the answer, it will not be cited. The structure that wins in SEO, comprehensive, well-linked, thorough, is often the structure that loses in AEO.</p>
<p>Every piece of AEO content needs to pass a simple test: if an AI engine read only the first 100 words of this page, would it have a complete, citable answer to the target query? If the answer is no, the content needs to be restructured.</p>
<h3>2. Schema markup at scale</h3>
<p>Websites with comprehensive schema markup implementation are <a href="https://www.wavesandalgorithms.com/intelligence-reports/answer-engine-optimization">78% more likely to be referenced by AI systems</a> compared to those without structured data. FAQ schema is particularly high-leverage: it allows AI engines to pull individual question-and-answer pairs as standalone citations, independent of the full page.</p>
<p>Most SEO implementations include basic schema. AEO requires FAQ schema on every applicable page, Article schema with author attribution and publication date, and Organization schema that connects all content back to a verified entity.</p>
<h3>3. Entity consistency across the web</h3>
<p>AI engines cross-reference a brand's presence across multiple sources before deciding whether to cite it confidently. If your LinkedIn company description says one thing, your Crunchbase listing says another, and your website says a third, the LLM's confidence in recommending you drops.</p>
<p>Entity consistency, aligning how your brand is described across every directory, profile, and publication, is an AEO-specific discipline. It has no meaningful SEO equivalent.</p>
<h3>4. Third-party citation footprint</h3>
<p>A company that exists only on its own website will not be confidently recommended by an AI engine. LLMs require external validation: Reddit threads, directory listings, press mentions, and reviews. These sources function as the AI equivalent of backlinks, but they operate differently. A single well-placed Reddit thread, properly structured and upvoted, can generate LLM citations for months.</p>
<h3>5. Funnel-stage content coverage</h3>
<p>Most companies have some top-of-funnel content and almost nothing at the bottom. Someone asking ChatGPT "how do I sign up for [your product]" or "what is the onboarding process for [your service]" will get no answer if you have not written content that tells AI engines what the next step is.</p>
<p>BOFU visibility is where AEO converts. It is also where the vast majority of companies have a 0% AI visibility score.</p>
<hr />
<p>These five levers compound. A clean technical foundation amplifies every piece of content published on top of it. Content published on a broken foundation, without schema, without entity consistency, without a citation footprint, sits there and does not move the needle.</p>
<p>You can read more about the full AEO approach on the <a href="https://windgrove.ai/blog">Windgrove blog</a>.</p>
<h2>The Honest Answer: <a href="https://windgrove.ai/blog/aeo-vs-seo-where-should-a-b2b-tech-company-invest-in-2026">AEO and SEO</a> Are Not Enemies</h2>
<p>Here is where the nuance matters.</p>
<p>AEO is not a replacement for SEO. Organic search still drives 25 to 42% of website traffic across major industries, according to <a href="https://almcorp.com/blog/aeo-geo-benchmarks-2025-conductor-analysis-complete-guide/">Conductor's AEO/GEO Benchmarks Report</a>, which analyzed 13,770 domains and 3.3 billion sessions. Traditional search is not dead. It is being compressed.</p>
<p>The question is not "should I do AEO or SEO?" The question is, "Am I building a presence that works across both engines?"</p>
<p>Strong SEO performance and strong AEO performance are correlated. A site with clean technical infrastructure, authoritative content, and a credible external presence will perform better in both channels. The work is not redundant. But it is also not identical.</p>
<blockquote>
<p><strong>The practical reality:</strong> Most companies have invested years in SEO and have almost nothing purpose-built for AI citation. That asymmetry is the gap. The companies closing it now are capturing category authority in AI search before their competitors realise the opportunity exists.</p>
</blockquote>
<h3>Where to start if you are behind</h3>
<p>If you are unsure where your AI visibility stands right now, the starting point is an audit. Not a content audit. An AI visibility audit: which prompts your buyers are using, which AI engines are answering them, which brands are being cited, and whether you are one of them.</p>
<p>That audit tells you two things. First, how large the gap is. Second, which interventions will close it fastest.</p>
<p>For most companies, the gap is larger than expected. A brand that ranks on page one of Google for its core category terms often has a 0% AI visibility score on the same terms. The work done for Google does not transfer automatically. It has to be rebuilt for a different engine.</p>
<p>The good news: the window is still open. AI visibility in most B2B categories is not yet dominant. The brands investing in AEO now are establishing positions that will be significantly harder to displace in 12 months than they are today.</p>
<h2>Find Out Where You Stand</h2>
<p>The first step is knowing your current AI visibility score. Not an estimate. Actual data: which prompts your buyers are using, which platforms are answering them, and whether your brand appears in those answers.</p>
<p>Windgrove offers a <a href="https://windgrove.ai/audit">free AI visibility audit</a> that maps your current position across ChatGPT, Perplexity, Google AI Overviews, and Gemini. You will see exactly where you are visible, where you are invisible, and which gaps represent the highest-leverage opportunities.</p>
<p>If you want to go further, you can <a href="https://cal.com/team/windgrove-ai/discovery-call">book a free consultation</a> with the Windgrove team. No generic deck. No obligation. A direct conversation about your specific situation and what a 30, 60, or 90-day AEO plan would look like for your business.</p>
<p>The question of whether AEO works has an answer. The more useful question is whether your brand is showing up in the answers your buyers are already getting.</p>
<hr />
<h2>Frequently Asked Questions</h2>
<p><strong>Is AEO the same as SEO?</strong> No. AEO and SEO share some foundations, including the importance of quality content and technical site health, but they target different engines with different signals. SEO optimizes for Google's ranking algorithm, which rewards backlinks, keyword relevance, and page authority. AEO optimizes for AI answer engines like ChatGPT, Perplexity, and Google AI Overviews, which require a direct answer in the first 40 to 60 words of content, comprehensive schema markup, entity consistency across the web, and a third-party citation footprint from sources like Reddit and industry directories. A page can rank on page one of Google and be completely invisible in AI-generated answers.</p>
<p><strong>How long does AEO take to show results?</strong> Initial results, including improvements in featured snippet acquisition and early AI citations, can appear within two to four weeks of implementing structural changes. Meaningful movement in AI visibility scores typically occurs within one to three months. The Opal case study showed a 0% to 15.9% AI visibility score increase and 1,766 brand mentions in 31 days, achieved by fixing the technical foundation first and then publishing AEO-structured content.</p>
<p><strong>Do I need to stop doing SEO to invest in AEO?</strong> No. AEO and SEO are complementary, not competing. Organic search still drives 25 to 42% of website traffic across major industries, and strong SEO performance often correlates with stronger AEO performance. The goal is a presence that works across both engines. Most companies are significantly under-invested in AEO relative to SEO, so the priority is closing that gap, not abandoning traditional search.</p>
<p><strong>What metrics should I track for AEO?</strong> The primary metrics for AEO are AI visibility score (the percentage of tracked prompts where your brand is mentioned), share of voice across AI platforms, average LLM position, and funnel-stage visibility across TOFU, MOFU, and BOFU queries. Secondary metrics include brand mentions in AI-generated responses, LLM-attributed leads entering your funnel, and conversion rates from AI-sourced traffic.</p>
<p><strong>How do I know if my current content is AEO-ready?</strong> The simplest test: read the first 100 words of any page on your site. If an AI engine read only those 100 words, would it have a complete, citable answer to the query that page targets? If the answer is buried further down, the page is not AEO-ready. Other indicators include missing FAQ schema, no author attribution, canonical URL conflicts, and the absence of an llms.txt file. A <a href="https://windgrove.ai/audit">free AI visibility audit</a> will surface these gaps systematically.</p>
]]></content:encoded></item><item><title><![CDATA[llms.txt Explained: What It Is, Why It Matters, and Why Google Now Checks for It]]></title><description><![CDATA[Executive Summaryllms.txt is a plain-text file that lives at the root of your website and tells AI crawlers what your site is about, which pages matter, and how to read your content efficiently.As of ]]></description><link>https://windgrove.hashnode.dev/llmstxt-explained-what-it-is-why-it-matters-and-why-google-now-checks-for-it</link><guid isPermaLink="true">https://windgrove.hashnode.dev/llmstxt-explained-what-it-is-why-it-matters-and-why-google-now-checks-for-it</guid><dc:creator><![CDATA[Spencer Duke]]></dc:creator><pubDate>Wed, 27 May 2026 00:00:00 GMT</pubDate><content:encoded><![CDATA[<blockquote>
<p><strong>Executive Summary</strong>llms.txt is a plain-text file that lives at the root of your website and tells AI crawlers what your site is about, which pages matter, and how to read your content efficiently.As of May 2026, Google's Chrome Lighthouse now includes a dedicated "Agentic Browsing" audit that checks whether your site has an llms.txt file, signalling that AI readiness has moved from optional to measurable.AI search traffic converts at 14.2% compared to Google's 2.8%, making it 5x more valuable per visitor. Yet most business websites are still invisible to AI engines because their technical foundation was never built for them (Exposure Ninja, 2026).Windgrove helped Opal go from 0% AI visibility to 15.9% and 1,766 brand mentions across LLMs in just 31 days. llms.txt configuration was one of the first technical steps in that process.</p>
</blockquote>
<hr />
<p>Most business owners have never heard of llms.txt. That is not surprising. It is a relatively new file format, proposed in 2024 and still gaining mainstream awareness. But the gap between "haven't heard of it" and "need to act on it" is closing fast.</p>
<p>In May 2026, Google added a check for llms.txt to Chrome's Lighthouse auditing tool — the same tool developers and marketers use to measure site performance and technical health. The new "Agentic Browsing" category evaluates how well your site is structured for machine interaction, and the presence of an llms.txt file is one of the explicit signals it looks for.</p>
<p>That is a meaningful shift. Google has publicly said llms.txt is not required for traditional search rankings. But Chrome is now flagging its absence as a readiness gap for <a href="https://windgrove.ai/blog/how-to-make-your-website-ai-agent-friendly-what-cloudflares-research-actually-means-for-you">AI agents</a>. Those two things can both be true at once, and understanding the distinction is exactly what this article is about.</p>
<p><strong>The core question:</strong> If your buyers are increasingly using ChatGPT, Perplexity, and Google AI Overviews to find products and services like yours, does your site give those systems what they need to find, read, and recommend you?</p>
<p>For most sites, the honest answer is no. llms.txt is one part of fixing that.</p>
<h2>What Is llms.txt?</h2>
<p>llms.txt is a plain-text file placed at the root of a website — typically accessible at <code>yourdomain.com/llms.txt</code> — that provides AI language models with a structured, machine-readable summary of your site's content and purpose.</p>
<p>Think of it as a table of contents written specifically for AI systems. A sitemap tells search engine crawlers which pages exist. llms.txt goes further. It tells AI agents what your site is actually about, which pages are most important, and how to interpret your content without crawling every page from scratch.</p>
<h3>The Analogy That Makes It Click</h3>
<blockquote>
<p>robots.txt tells crawlers what they can and cannot access. llms.txt tells AI agents what is worth their attention and how to understand it.</p>
</blockquote>
<p>The format was first proposed in 2024 by Jeremy Howard of Answer.AI as a community standard, not a Google-mandated protocol. It is written in simple Markdown and typically contains:</p>
<ul>
<li>A brief description of the organisation and what it does</li>
<li>Links to the most important pages on the site (product pages, documentation, key articles)</li>
<li>Optional context about content categories, intended audiences, and how the site is structured</li>
</ul>
<p>Here is a simplified example of what an llms.txt file looks like in practice:</p>
<pre><code># Acme Software

&gt; Acme builds project management tools for remote engineering teams.

## Core Pages
- [Product Overview](https://acme.com/product): What Acme does and who it is for
- [Pricing](https://acme.com/pricing): Plans and pricing for teams of all sizes
- [Documentation](https://acme.com/docs): Full technical documentation

## Blog
- [How we reduced onboarding time by 40%](https://acme.com/blog/onboarding)
- [Remote team productivity guide](https://acme.com/blog/remote-teams)
</code></pre>
<p>That is it. No code. No developer required. A text file with a clear structure.</p>
<h3>What llms.txt Is Not</h3>
<p>It is not a replacement for robots.txt, which controls crawler access permissions. It does not override your sitemap. And it does not directly determine whether you rank in Google's traditional blue-link results.</p>
<p>What it does is reduce the cognitive load on AI agents that are trying to understand your site quickly. Google's own Lighthouse documentation states it plainly: <strong>"Without llms.txt, agents may spend more time crawling the site to understand its high-level structure and primary content."</strong></p>
<p>Less time crawling means faster comprehension. Faster comprehension means a higher likelihood of accurate, confident citation.</p>
<h2>Why Google Now Checks for It</h2>
<p>On 20 May 2026, <a href="https://searchengineland.com/google-llms-txt-chrome-lighthouse-478246">Search Engine Land reported</a> that Google added llms.txt detection to Chrome's Lighthouse auditing tool under a new category called "Agentic Browsing." This is the same Lighthouse tool that measures Core Web Vitals, accessibility, and SEO performance — the scores that developers and marketers already track closely.</p>
<p>The new category does not produce a traditional 0-100 score. It surfaces a pass/fail ratio across a set of agentic readiness signals. The checks include:</p>
<ul>
<li>WebMCP integration</li>
<li>Accessibility tree integrity</li>
<li>Layout stability through Cumulative Layout Shift (CLS)</li>
<li><strong>Presence of an llms.txt file</strong></li>
</ul>
<p>The framing matters. Google is not checking for llms.txt because it affects traditional search rankings. It is checking because Chrome is increasingly used by AI agents, and agents benefit from sites that are easier for machines to navigate.</p>
<h3>The Nuance Worth Understanding</h3>
<p>Google's John Mueller addressed this directly in May 2026, responding to a question from SEO expert Lily Ray about the apparent contradiction between Google's own use of llms.txt and its official guidance that the file is not needed for search:</p>
<blockquote>
<p>"The short answer is that it's not done for search. There's more to websites than just SEO. It's worth separating 'discovery' (finding the website or pages with a global search engine) vs 'functionality' (once someone has found the page, helping them to best do the task they want to do)."</p>
</blockquote>
<p>This is a useful distinction. llms.txt does not help Google <em>find</em> your site. It helps AI agents <em>use</em> your site more effectively once they arrive. For businesses whose buyers increasingly interact with AI tools, that functional layer is exactly where visibility is won or lost.</p>
<p><strong>Lighthouse Flags AI Readiness Gaps:</strong> Google's Lighthouse now flags the absence of llms.txt as a gap in your site's machine readiness. Whether or not it affects your ranking today, it signals where the web is heading. Sites that are easy for AI agents to read will be cited more confidently than sites that are not.</p>
<p>The same Lighthouse category also emphasises accessibility tree integrity and layout stability — signals that AI agents rely on as their "primary data model" when navigating pages. llms.txt is one layer of a broader machine-readability picture.</p>
<h2>Why This Matters for Your Business Right Now</h2>
<p>The timing of Google's Lighthouse update is not coincidental. It reflects a broader shift in how buyers find and evaluate products and services.</p>
<p><strong>51% of B2B software buyers now start their research with an AI chatbot</strong> more often than with Google (G2, April 2026). That number has moved from a prediction to a reality. If your site is not structured for AI comprehension, you are invisible when it matters most. That is when your buyer is forming their shortlist.</p>
<p>The traffic quality argument makes this even more urgent. <a href="https://exposureninja.com/blog/ai-search-statistics/">According to Exposure Ninja's 2026 AI Search Statistics report</a>, AI search traffic converts at <strong>14.2%</strong> compared to Google's 2.8%. That means a visitor arriving from an AI-generated answer is roughly <strong>five times more likely to convert</strong> than one arriving from a traditional search result.</p>
<h3>The Scale of What AI Agents Are Already Doing</h3>
<p>The <a href="https://www.humansecurity.com/learn/resources/2026-state-of-ai-traffic-cyberthreat-benchmarks/">2026 State of AI Traffic report from Human Security</a> documented a fundamental shift in how automated systems interact with the web:</p>
<ul>
<li>Monthly volumes of AI-driven traffic <strong>grew 187% from January to December 2025</strong>, nearly tripling over the calendar year.</li>
<li>Traffic from AI agents and agentic browsers grew <strong>7,851% year over year</strong></li>
<li>OpenAI's bots alone accounted for approximately <strong>69% of all observed AI bot traffic</strong> in 2025</li>
</ul>
<p>These are not future projections. This is what already happened in 2025. AI agents are reading, navigating, and increasingly transacting on the web at a scale most businesses have not yet accounted for in their technical infrastructure.</p>
<h3>The Conversion Gap Is Already Costing You</h3>
<p>Here is the scenario that plays out every day for businesses without proper AI infrastructure:</p>
<ol>
<li>A buyer asks ChatGPT or Perplexity: "What's the best [your category] tool for [your ICP]?"</li>
<li>The AI generates a confident answer naming three or four competitors.</li>
<li>Your brand is not mentioned. Not because your product is inferior, but because the AI could not efficiently read and understand your site.</li>
<li>The buyer never visits your site. The deal goes elsewhere.</li>
</ol>
<p>llms.txt alone does not solve this. But it is part of the technical foundation that makes your site legible to the systems your buyers are now relying on. And it is one of the fastest, lowest-effort technical fixes available.</p>
<h2>What Goes Into a Well-Configured llms.txt File</h2>
<p>There is no single mandatory format for llms.txt. The community standard proposed by Jeremy Howard of Answer.AI in 2024 provides a flexible Markdown-based structure, and most implementations follow a similar pattern.</p>
<p>A well-configured llms.txt file typically includes four components:</p>
<table>
<thead>
<tr>
<th>Component</th>
<th>Purpose</th>
<th>Example</th>
</tr>
</thead>
<tbody><tr>
<td><strong>H1 heading</strong></td>
<td>Identifies the site or organisation</td>
<td><code># Acme Software</code></td>
</tr>
<tr>
<td><strong>Blockquote description</strong></td>
<td>One-sentence summary of what the site does</td>
<td><code>&gt; Project management tools for remote engineering teams</code></td>
</tr>
<tr>
<td><strong>Section links</strong></td>
<td>Curated list of key pages with short descriptions</td>
<td>Core pages, documentation, blog articles</td>
</tr>
<tr>
<td><strong>Optional context</strong></td>
<td>Notes on content type, audience, or site structure</td>
<td><code>## Notes: Content is written for engineering managers</code></td>
</tr>
</tbody></table>
<h3>What to Include and What to Skip</h3>
<p>The goal is not to list every page on your site. It is to give AI agents a fast, accurate orientation. Think of it as the briefing document you would hand someone before they read your entire website.</p>
<p><strong>Include:</strong></p>
<ul>
<li>Your homepage and core product or service pages</li>
<li>Pricing page (if public)</li>
<li>Key blog articles or resources that represent your expertise</li>
<li>Contact or booking pages (so AI agents can recommend the right next step)</li>
<li>Any comparison or "vs" pages you have published</li>
</ul>
<p><strong>Skip:</strong></p>
<ul>
<li>Privacy policy, terms of service, and legal pages</li>
<li>Admin or login pages</li>
<li>Thin or duplicate content</li>
<li>Pages under active construction</li>
</ul>
<h3>The llms-full.txt Variant</h3>
<p>Some sites also publish an <code>llms-full.txt</code> file, which contains the full text content of key pages rather than just links. This is particularly useful for documentation-heavy sites where AI agents frequently need to extract detailed technical information. For most business websites, the standard llms.txt is sufficient.</p>
<p><strong>One important note:</strong> llms.txt is a guide, not a command. AI agents can choose to follow it or ignore it. But providing the file removes friction from the process. An agent that has a clear map of your site will almost always produce a more accurate, more confident summary of what you do. That accuracy is what drives citation.</p>
<h2>llms.txt in Practice: The Opal Case Study</h2>
<p>Understanding what llms.txt does in theory is one thing. Seeing what happens when it is part of a complete AI visibility overhaul is another.</p>
<p>In late March 2026, Windgrove began working with <a href="https://opalspend.com/">Opal</a>, a charge card and spend management platform built for digital marketing agencies. The situation was stark: a strong product with zero AI search visibility. Opal had only four indexed pages, no blog, weak metadata, and no sitemap submitted to Google Search Console. Anyone searching their category in ChatGPT, Perplexity, or Google AI Overviews was finding only competitors.</p>
<h3>The Technical Foundation First</h3>
<p>Before publishing a single piece of content, Windgrove cleared every blocker between Opal's pages and AI crawlers. The work included:</p>
<ul>
<li>Submitting an XML sitemap to Google Search Console</li>
<li>Overhauling and redeploying robots.txt</li>
<li><strong>Building and configuring llms.txt for LLM indexing and citation</strong></li>
<li>Rewriting meta titles and descriptions across every core page</li>
<li>Restructuring heading hierarchy and internal linking</li>
<li>Resolving indexing issues on key product and landing pages</li>
</ul>
<p>llms.txt was not an afterthought. It was part of the first wave of technical work, alongside the sitemap and robots.txt. The goal was to make Opal's site fully legible to AI systems before any content was published on top of it.</p>
<h3>The Results After 31 Days</h3>
<p>The <a href="https://windgrove.ai/case-studies/opal">full Opal case study</a> documents what happened next:</p>
<table>
<thead>
<tr>
<th>Metric</th>
<th>Before</th>
<th>After (31 days)</th>
</tr>
</thead>
<tbody><tr>
<td>AI Visibility Score</td>
<td>0%</td>
<td>15.9%</td>
</tr>
<tr>
<td>Brand Mentions (LLMs + web)</td>
<td>0</td>
<td>1,766</td>
</tr>
<tr>
<td>Site Health Score</td>
<td>66.2</td>
<td>80.7 (top 10% of benchmarked sites)</td>
</tr>
<tr>
<td>AEO Articles Live</td>
<td>0</td>
<td>8</td>
</tr>
<tr>
<td>Average LLM Position</td>
<td>—</td>
<td>#1</td>
</tr>
</tbody></table>
<p>Within one week of launching the Ad Pay page, Opal ranked #2 for "ad pay" and #2 for "ad spend cards" — bottom-funnel terms where buyers are already in-market and evaluating options.</p>
<p><strong>The key insight from Opal's results:</strong> Most sites take three to six months to see meaningful movement on bottom-funnel terms. Opal was there in seven days. That is what a clean technical foundation does before content is built on top of it.</p>
<p>Opal is now being surfaced inside Perplexity AI responses as a named recommendation. Not as an ad, not as a sponsored result. As the answer. The 15.9% AI visibility score reached in 31 days is a starting point, not a ceiling. Content compounds on a strong technical foundation. It stagnates on a broken one.</p>
<p>You can see Windgrove's broader track record of results at <a href="https://windgrove.ai/proof">windgrove.ai/proof</a>.</p>
<h2>How to Get Started With llms.txt</h2>
<p>Creating an llms.txt file is not technically complex. The challenge is doing it strategically — knowing which pages to include, how to describe your organisation accurately, and how to structure the file so it actually improves AI comprehension rather than just existing as a checkbox.</p>
<h3>The Basic Steps</h3>
<ol>
<li><strong>Audit what you have.</strong> Before writing the file, map your site's most important pages. Product pages, service pages, key blog articles, your about page, and your contact or booking page are the typical starting points.</li>
<li><strong>Write a clear one-sentence description.</strong> This is the most important line in the file. It tells AI agents what your organisation does, who it serves, and what makes it distinct. Vague descriptions produce vague citations.</li>
<li><strong>Create the file in Markdown.</strong> Use a plain text editor. The format is simple: an H1 with your brand name, a blockquote with your description, and organised sections of links with short annotations.</li>
<li><strong>Place it at your domain root.</strong> The file must be accessible at <code>yourdomain.com/llms.txt</code>. It should not require authentication to access.</li>
<li><strong>Keep it updated.</strong> As your site evolves — new products, new content, new pages — your llms.txt should reflect those changes.</li>
</ol>
<h3>What llms.txt Cannot Do Alone</h3>
<p>A common mistake is treating llms.txt as a standalone fix. It is not. It is one component of a broader technical and content infrastructure that AI engines use to evaluate whether your site is worth citing.</p>
<p>The full picture includes:</p>
<ul>
<li><strong>XML sitemap:</strong> ensures AI crawlers can discover all your pages</li>
<li><strong>robots.txt:</strong> controls which pages crawlers can access</li>
<li><strong>Schema.org structured data:</strong> tells AI engines exactly what each piece of content is, who wrote it, and what it is about</li>
<li><strong>Canonical URL resolution:</strong> prevents AI engines from treating duplicate pages as separate entities</li>
<li><strong>Author attribution on content:</strong> a key trust signal for citation</li>
<li><strong>AEO-optimised content structure:</strong> articles that front-load direct answers so AI engines can extract and cite them accurately</li>
</ul>
<p>llms.txt improves the efficiency of AI comprehension. The rest of the stack determines what there is to comprehend. Both matter. You can read more about how these layers work together on the <a href="https://windgrove.ai/blog">Windgrove blog</a>.</p>
<h2>Frequently Asked Questions</h2>
<h3>Does llms.txt affect my Google search rankings?</h3>
<p>No. Google has confirmed that llms.txt does not influence traditional blue-link search rankings. Google's John Mueller stated in May 2026 that llms.txt is "not done for search" and is instead about functionality — helping AI agents use your site more effectively once they arrive. The two systems are separate. llms.txt is an AI readability signal, not an SEO ranking factor.</p>
<h3>How long does it take to create an llms.txt file?</h3>
<p>For a straightforward business website, creating a basic llms.txt file takes between 30 minutes and two hours. The file itself is plain text written in Markdown — no coding required. The more time-consuming part is deciding which pages to include and writing a precise one-sentence description of your organisation. A vague description produces vague AI citations, so that line is worth getting right.</p>
<h3>Will AI engines like ChatGPT and Perplexity automatically read my llms.txt file?</h3>
<p>llms.txt is a guide, not a command. AI agents and crawlers can choose to follow it or ignore it. That said, major AI platforms including OpenAI's crawler (OAI-SearchBot) and Perplexity's crawler are increasingly designed to look for and use llms.txt files when present. Providing the file removes friction and improves the accuracy of how AI systems describe and cite your business.</p>
<h3>Is llms.txt enough to make my site visible in AI search results?</h3>
<p>No. llms.txt improves AI comprehension efficiency, but it is one layer of a broader technical and content stack. A complete AI-ready site also needs a properly configured XML sitemap, clean robots.txt, Schema.org structured data, resolved canonical URLs, author attribution on content, and AEO-optimised articles that front-load direct answers. llms.txt without the rest of the stack is like having a clean front door on a building with no address.</p>
<h3>How do I know if my site has the right AI visibility infrastructure in place?</h3>
<p>The fastest way is a structured audit. Windgrove offers a <a href="https://windgrove.ai/audit">free AI visibility audit</a> that covers technical infrastructure gaps. That includes llms.txt presence, schema implementation, canonical issues, and content structure. You get a clear picture of exactly what is preventing AI engines from reading and citing your site. If you want to understand the full scope of what is possible, you can also <a href="https://windgrove.ai/audit">book a free consultation</a> with the Windgrove team.</p>
]]></content:encoded></item><item><title><![CDATA[How Windgrove AI Billing Works, And Why You Won't Encounter Unexpected Charges]]></title><description><![CDATA[How Windgrove Billing Works at a Glance
Before you read further, here is the short version:

Windgrove does not use hidden charges. Every fee is agreed upon before work begins.
You will always know wh]]></description><link>https://windgrove.hashnode.dev/billing</link><guid isPermaLink="true">https://windgrove.hashnode.dev/billing</guid><category><![CDATA[service]]></category><dc:creator><![CDATA[Mitko Dimitrov]]></dc:creator><pubDate>Fri, 22 May 2026 00:00:00 GMT</pubDate><content:encoded><![CDATA[<h2>How Windgrove Billing Works at a Glance</h2>
<p><strong>Before you read further, here is the short version:</strong></p>
<ul>
<li>Windgrove does not use hidden charges. Every fee is agreed upon before work begins.</li>
<li>You will always know what you are paying for. Scope and pricing are defined upfront, in plain language.</li>
<li>In limited cases, third-party websites or platforms charge separately for listings or placements. When that happens, Windgrove tells you in advance, explains exactly what the cost covers, and waits for your approval before spending a dollar.</li>
<li>Nothing moves forward without your sign-off. If you decline an optional cost, it does not proceed.</li>
</ul>
<p>That is the operating principle behind everything Windgrove does. The rest of this guide explains how it works in practice and what you can expect at every stage of the relationship.</p>
<h2>Why Billing Clarity Matters More in AI Services</h2>
<p>Buying AI services in 2026 is not like buying a fixed-price subscription. <a href="https://zylo.com/blog/ai-cost/">According to Zylo's 2026 SaaS Management Index</a>, annual SaaS spend rose 8% while application counts stayed flat, meaning buyers are paying more without necessarily getting more. Nearly one-third of AI vendors now use hybrid pricing models that combine flat fees with usage-based or value-based charges. The result is billing that shifts mid-contract, often without clear warning.</p>
<p>That makes pricing transparency a real differentiator, not just a nice-to-have. When buyers evaluate an AI agency, they are not only asking "will this work?" They are asking "will I be in control of what I spend?"</p>
<p>The table below shows how two approaches compare:</p>
<table>
<thead>
<tr>
<th>What some AI services do</th>
<th>What Windgrove does</th>
</tr>
</thead>
<tbody><tr>
<td>Define scope loosely, adjust pricing later</td>
<td>Define scope and fees before work begins</td>
</tr>
<tr>
<td>Bundle optional features into base billing</td>
<td>Separate core fees from optional third-party costs</td>
</tr>
<tr>
<td>Notify clients after charges are incurred</td>
<td>Communicate all additional costs before acting</td>
</tr>
<tr>
<td>Require clients to opt out of extra spend</td>
<td>Require explicit client approval before any extra spend</td>
</tr>
<tr>
<td>Use vague line items on invoices</td>
<td>Invoice only for agreed, clearly described work</td>
</tr>
</tbody></table>
<p><strong>The real issue:</strong> as the <a href="https://www.filevine.com/guides/ai-trust-index-survey-report/">2026 AI Trust Index</a> found, trust in AI services is conditional. It rises when clients can verify what happened, see a clear audit trail, and know that a human is accountable for every decision. Windgrove builds that accountability into billing from day one.</p>
<h2>How Windgrove Billing Works in Practice</h2>
<p>The billing process is straightforward. Here is what it looks like from the moment you engage with Windgrove to the moment an invoice arrives.</p>
<ol>
<li><strong>Scope and pricing are agreed before work begins.</strong> Before Windgrove starts any campaign or project, you receive a clear breakdown of what is included, what it costs, and what the deliverables are. There are no open-ended agreements that allow fees to expand without notice.</li>
<li><strong>Work is delivered against the agreed scope.</strong> Windgrove's team executes on what was defined. If something changes during the project, that conversation happens before any additional work is done, not after.</li>
<li><strong>Invoices reflect only what was agreed.</strong> When a bill arrives, every line item maps to something you approved. There are no surprise additions, no retroactive charges, and no fees for work you did not know was happening.</li>
<li><strong>Optional opportunities are presented separately.</strong> If Windgrove identifies an additional tactic or placement that could benefit your visibility, it is presented as a separate recommendation with its own cost and rationale. You decide whether to proceed.</li>
<li><strong>You can ask billing questions at any point.</strong> Windgrove welcomes questions before, during, or after a campaign. If anything on an invoice is unclear, the team will explain it. Clarity is not a courtesy here; it is how the relationship is designed to work.</li>
</ol>
<blockquote>
<p><strong>The principle:</strong> <a href="https://improvado.io/blog/automated-client-reporting">Transparent reporting and built-in human oversight</a> are what separate agencies that clients trust long-term from those that create friction at every renewal. Windgrove's billing process is built around that standard.</p>
</blockquote>
<h2>When Additional Costs Can Happen, and How Approval Works</h2>
<p>Windgrove's own fees are fixed and agreed upfront. There is, however, one category of cost that can arise outside of that: third-party website listings and placements.</p>
<h3>What this means in practice</h3>
<p>Part of building AI visibility is getting your business cited and listed on external platforms, directories, and authoritative websites. Some of those platforms charge a fee for inclusion. Those fees are not Windgrove's revenue; they go directly to the third-party platform. But they do require additional budget beyond your core Windgrove engagement.</p>
<blockquote>
<p><strong>Important:</strong> Windgrove never incurs a third-party cost on your behalf without telling you first. If a listing or placement opportunity involves an external charge, here is exactly what happens:</p>
</blockquote>
<ul>
<li>Windgrove identifies the opportunity and assesses whether it is worth pursuing for your visibility goals.</li>
<li>You receive a clear explanation of what the platform is, what the fee covers, and why it is being recommended.</li>
<li>The cost is presented to you before any commitment is made.</li>
<li>You decide. If you approve, Windgrove proceeds. If you decline, the opportunity is set aside and no charge is incurred.</li>
</ul>
<p>This process applies every single time. There is no situation in which Windgrove spends additional budget on your behalf without documented approval.</p>
<p>As <a href="https://www.prnewswire.com/in/news-releases/avalara-predicts-2026-will-reshape-global-business-through-ai-transparency-and-compliance-agility-302613936.html">Avalara's 2026 compliance outlook</a> noted, accountability and auditable decision-making are fast becoming the baseline expectation in AI-era business relationships. Windgrove treats that standard not as a regulatory checkbox, but as the foundation of how client relationships should work.</p>
<h2>What You Will Not Experience with Windgrove</h2>
<p>Some billing concerns come up regularly when businesses evaluate AI agencies. Here is a direct answer to each one.</p>
<table>
<thead>
<tr>
<th>The concern</th>
<th>The reality at Windgrove</th>
</tr>
</thead>
<tbody><tr>
<td>"Will fees appear that I didn't agree to?"</td>
<td>No. Every charge maps to something you approved before work began.</td>
</tr>
<tr>
<td>"Could a third-party cost be added without my knowledge?"</td>
<td>No. Any external platform fee is disclosed and approved before it is incurred.</td>
</tr>
<tr>
<td>"Will I get a vague invoice I can't make sense of?"</td>
<td>No. Line items describe real, agreed deliverables in plain language.</td>
</tr>
<tr>
<td>"What if I say no to an optional cost?"</td>
<td>Nothing happens. The spend does not proceed.</td>
</tr>
<tr>
<td>"Can the scope quietly expand mid-project?"</td>
<td>No. Scope changes are discussed before they are acted on.</td>
</tr>
</tbody></table>
<p>Budget control stays with you. That is not a policy Windgrove follows reluctantly; it is the way the business was built. Clients who feel in control of their spend are the ones who stay, grow, and refer others. Transparency is not a trust signal. It is the operating model.</p>
<h2>Questions to Ask Before Approving Any Agency Invoice</h2>
<p>Whether you are working with Windgrove or evaluating any AI agency, these questions will tell you quickly whether billing is being handled properly.</p>
<h3>Before you sign</h3>
<ul>
<li>Is the scope of work written down and specific?</li>
<li>Are the fees fixed, or can they change based on usage or platform decisions?</li>
<li>Is there a clear distinction between the agency's fees and any third-party costs?</li>
</ul>
<h3>Before you approve additional spend</h3>
<ul>
<li>Has the agency explained what the extra cost is for and why it is being recommended?</li>
<li>Do you have the option to decline without affecting the core engagement?</li>
<li>Is there a written record of your approval before the spend is made?</li>
</ul>
<h3>Before you sign off on an invoice</h3>
<ul>
<li>Does every line item correspond to something you agreed to?</li>
<li>Can the agency explain any item you do not recognise?</li>
<li>Are there any charges you were not notified of in advance?</li>
</ul>
<p>A trustworthy agency will not hesitate on any of these. Windgrove welcomes them. If you have billing questions before approving a campaign, <a href="https://windgrove.ai/">reach out to the team directly</a> before anything is confirmed.</p>
<h2>FAQ: Windgrove AI Billing and Unexpected Charges</h2>
<p><strong>Does Windgrove have hidden charges?</strong> No. Every fee is agreed upon before work begins. There are no charges added after the fact without your knowledge or approval.</p>
<p><strong>Can my invoice change without my approval?</strong> No. If anything outside the original scope is needed, Windgrove discusses it with you first. Work does not proceed, and charges are not incurred, until you have confirmed.</p>
<p><strong>What happens if a third-party listing costs extra?</strong> Windgrove will explain the opportunity, the cost, and the reason for recommending it. You decide whether to proceed. If you say no, the listing is not pursued and no charge is incurred.</p>
<p><strong>Can I ask billing questions before approving a campaign?</strong> Yes. Windgrove encourages this. Contact the team at <a href="mailto:contact@windgrove.ai">contact@windgrove.ai</a> with any questions about scope, fees, or third-party costs before giving your approval.</p>
<p><strong>How does Windgrove keep billing clear from the start?</strong> By defining scope and pricing before work begins, separating core fees from optional third-party costs, and requiring explicit client approval before any additional spend is made. The process is designed so you are never surprised.</p>
<h2>Related reading</h2>
<ul>
<li><a href="https://windgrove.ai/blog/aeo-roi">Why Investing in AEO Now Pays Compounding Dividends Later</a></li>
<li><a href="https://windgrove.ai/blog/metrics">How We Measure Success: The Metrics Behind Your AI Visibility Program</a></li>
</ul>
]]></content:encoded></item><item><title><![CDATA[How to Track LLM Crawls With Log Files, Looker Studio, and Other Low-Cost Methods: Without Fooling Yourself]]></title><description><![CDATA[Executive Summary

**TL;DR — four things to know before you read on:**Log files are the only reliable source of truth. GA4 and most analytics platforms miss the majority of AI crawler activity. Server]]></description><link>https://windgrove.hashnode.dev/how-to-track-llm-crawls-with-log-files-looker-studio-and-other-low-cost-methods-without-fooling-your</link><guid isPermaLink="true">https://windgrove.hashnode.dev/how-to-track-llm-crawls-with-log-files-looker-studio-and-other-low-cost-methods-without-fooling-your</guid><category><![CDATA[tracking]]></category><category><![CDATA[aeo]]></category><dc:creator><![CDATA[Mitko Dimitrov]]></dc:creator><pubDate>Thu, 21 May 2026 00:00:00 GMT</pubDate><content:encoded><![CDATA[<h2>Executive Summary</h2>
<blockquote>
<p>**TL;DR — four things to know before you read on:**<strong>Log files are the only reliable source of truth.</strong> GA4 and most analytics platforms miss the majority of AI crawler activity. Server logs or CDN logs capture every request, including bots that never execute JavaScript.<strong>The cheapest functional stack is Cloudflare or host logs, plus BigQuery or Google Sheets, plus Looker Studio.</strong> Each layer is either free or near-free at typical marketing-team volumes.<strong>Not all AI traffic means the same thing.</strong> Training crawlers, retrieval crawlers, and user-triggered fetchers behave differently, have different implications for your content strategy, and should never be lumped into a single "LLM traffic" row.<strong>The goal is a trustworthy baseline, not a bot hit counter.</strong> Raw crawler numbers mean almost nothing on their own. What you want is a clean, consistent dataset you can compare against referral traffic, citation mentions, and actual business outcomes over time.</p>
</blockquote>
<p>This guide walks through what you can measure accurately, compares low-cost tool options, and shows a practical workflow from logs to dashboard — including where the data breaks down and how to avoid building a report that looks authoritative but tells you nothing useful.</p>
<h2>Why Tracking LLM Crawls Is Suddenly Worth Doing</h2>
<p>AI crawlers are no longer a footnote in your server logs. According to <a href="https://blog.cloudflare.com/radar-2025-year-in-review/">Cloudflare's 2025 Year in Review</a>, AI bots averaged 4.2% of all HTML page requests across Cloudflare's network in 2025 — while user-action crawling, the category most closely tied to live AI queries, grew more than 15x during the same period. By Q1 2026, AI crawlers had become the second-largest bot category after search engines, representing <a href="https://technologychecker.io/blog/web-traffic-statistics">22% of all bot traffic</a> on Cloudflare's global network.</p>
<blockquote>
<p><strong>Key context:</strong> Googlebot still accounted for 4.5% of HTML requests in 2025 — slightly more than all other AI bots combined. That matters. If you see a spike in bot traffic and assume it is ChatGPT citing your content, it is more likely to be Googlebot, a training crawler, or an undeclared scraper. Context prevents expensive misreads.</p>
</blockquote>
<p>The category that deserves the most attention is <strong>user-triggered fetching</strong>: bots like ChatGPT-User and Perplexity-User that fire when a real person asks an AI a question and the model retrieves your page in real time. According to <a href="https://blog.cloudflare.com/ai-crawler-traffic-by-purpose-and-industry/">Cloudflare's crawler purpose analysis</a>, ChatGPT-User alone represented nearly three-quarters of user-action crawl traffic in the studied period, and its volume showed clear daily cycles tied to human usage patterns.</p>
<p>That is the signal worth tracking. The rest — training crawls, undeclared scrapers — tells you about data collection activity, not about whether AI engines are actively surfacing your brand to buyers.</p>
<h2>What You Can Measure Accurately — and What You Cannot</h2>
<p>Before you build anything, understand what the data can and cannot prove. Most measurement failures in this space come from treating a log file as a citation report. It is not.</p>
<table>
<thead>
<tr>
<th><strong>You CAN measure from logs</strong></th>
<th><strong>You CANNOT reliably infer from logs alone</strong></th>
</tr>
</thead>
<tbody><tr>
<td>Which declared user agents hit your site</td>
<td>Whether the AI actually cited or recommended you</td>
</tr>
<tr>
<td>Which URLs were crawled, and how often</td>
<td>Whether a human buyer saw the AI's answer</td>
</tr>
<tr>
<td>Response codes (200, 404, 5XX) per bot</td>
<td>Whether a crawler visit led to a referral session</td>
</tr>
<tr>
<td>Crawl frequency and timing patterns</td>
<td>The commercial intent behind any given crawl</td>
</tr>
<tr>
<td>Bytes transferred per request</td>
<td>Whether a training crawl will improve future AI answers</td>
</tr>
<tr>
<td>Bot categories (training, retrieval, user-action)</td>
<td>Whether an undeclared bot is from a specific platform</td>
</tr>
</tbody></table>
<p><strong>User-agent strings can be spoofed.</strong> Any server can send a request claiming to be GPTBot. The confidence level of your data improves significantly when you cross-reference user agents against published IP ranges. <a href="https://developers.openai.com/api/docs/bots">OpenAI publishes separate IP JSON files</a> for GPTBot, OAI-SearchBot, and ChatGPT-User. Cloudflare's verified bot classification does this cross-referencing automatically, which is one reason it is the recommended starting point for teams without a dedicated data engineer.</p>
<p><strong>AI crawl traffic and AI referral traffic are two different datasets.</strong> Crawl data lives in your server logs. Referral data, when it appears at all, shows up in GA4 or your analytics platform as a session from a source like <code>perplexity.ai</code> or <code>chat.openai.com</code>. The two are related but do not map to each other cleanly. Track them separately, report them separately, and resist the urge to merge them into a single "AI visibility" metric.</p>
<p>The most honest use of log-based AI crawler data is as a <strong>baseline</strong>: a consistent, repeatable measure of which bots are active on your site, which pages they prioritize, and how that pattern shifts over time.</p>
<h2>The Cheapest Ways to Track LLM Traffic, Ranked by Cost and Effort</h2>
<p>There is no single right answer. The best stack depends on whether you already use Cloudflare, how much log volume your site generates, and how comfortable your team is with SQL. Here is an honest comparison.</p>
<table>
<thead>
<tr>
<th><strong>Tool / Method</strong></th>
<th><strong>Setup Effort</strong></th>
<th><strong>Monthly Cost</strong></th>
<th><strong>Strengths</strong></th>
<th><strong>Weaknesses</strong></th>
</tr>
</thead>
<tbody><tr>
<td><strong>Cloudflare AI Crawl Control</strong></td>
<td>Low (1–2 hours)</td>
<td>Free on most plans</td>
<td>Built-in bot classification, verified IPs, one-click controls, no parsing required</td>
<td>Only available if you use Cloudflare as your CDN/DNS</td>
</tr>
<tr>
<td><strong>Host / server logs (Apache, Nginx)</strong></td>
<td>Medium (parsing required)</td>
<td>Free</td>
<td>Complete request-level data, no third-party dependency</td>
<td>Raw files are unwieldy; need grep, awk, or a log parser to extract signal</td>
</tr>
<tr>
<td><strong>BigQuery + Looker Studio</strong></td>
<td>Medium (3–6 hours setup)</td>
<td>Free tier covers most small sites; ~$5/month at scale</td>
<td>Scalable, SQL-queryable, connects natively to Looker Studio for dashboards</td>
<td>Requires some SQL comfort and a log ingestion step</td>
</tr>
<tr>
<td><strong>Google Sheets + Looker Studio</strong></td>
<td>Low–Medium</td>
<td>Free</td>
<td>Good for prototyping; no SQL needed</td>
<td>Breaks above ~50,000 rows; manual refresh unless scripted</td>
</tr>
<tr>
<td><strong>Cloudflare Logpush to BigQuery</strong></td>
<td>Medium–High</td>
<td>Free for logs; BigQuery storage costs apply</td>
<td>Automated, structured, verified bot data delivered directly to your warehouse</td>
<td>Requires Cloudflare Business plan or above for Logpush</td>
</tr>
</tbody></table>
<h3>The recommended default for non-engineers</h3>
<p>If your site already runs behind Cloudflare, start with <strong>AI Crawl Control</strong> (found under Security &gt; Bots in your dashboard). It gives you bot traffic breakdowns by category, request frequency, and crawl purpose without any log parsing. For reporting, connect a Google Sheets export or a Cloudflare Analytics API pull to <a href="https://lookerstudio.google.com/">Looker Studio</a> and you have a functional dashboard in an afternoon.</p>
<p>If you are not on Cloudflare, the practical path is: pull your host logs, run a grep filter for known AI user agents, push the cleaned output to BigQuery using a scheduled upload, and connect BigQuery to Looker Studio as a data source. <a href="https://docs.cloud.google.com/bigquery/docs/visualize-looker-studio">Google's own documentation</a> covers the BigQuery-to-Looker Studio connection step by step.</p>
<p><strong>Looker Studio is a reporting layer, not a processing layer.</strong> Do not push raw log files directly into it and expect clean results. Pre-aggregate the data in BigQuery or Sheets first. Dashboards built on pre-calculated tables load faster and are far easier to maintain.</p>
<h2>A Practical Low-Cost Workflow: Logs to BigQuery to Looker Studio</h2>
<p>Here is the recommended implementation path for a team without a dedicated data engineer. Each step is achievable in a few hours spread across a week.</p>
<p><img src="https://cdn.hashnode.com/uploads/posts/6a6513b3e2d2908b72dc5cf1/6ace7e7b-b28f-4417-89d4-1421cd251e46.jpg" alt="Cloudflare Radar AI insights dashboard showing AI crawler traffic breakdown by bot type and crawl purpose" /></p>
<p><em>Cloudflare Radar's</em> <a href="https://radar.cloudflare.com/ai"><em>AI Insights page</em></a> <em>shows real-time crawler traffic by bot and purpose — a useful benchmark for what you should expect to see in your own logs.</em></p>
<h3>Step 1: Collect your raw log data</h3>
<p>Pull access logs from your web server (Apache, Nginx) or CDN. If you use Cloudflare, enable Cloudflare Analytics or use the Logpush feature to stream structured logs to a storage bucket or directly to BigQuery. If you use a managed host, check whether your control panel exposes raw access logs for download.</p>
<h3>Step 2: Filter for known AI user agents</h3>
<p>Run a grep or regex filter against the raw logs to isolate AI-related traffic. A practical starting pattern:</p>
<pre><code>grep -Ei "GPTBot|OAI-SearchBot|ChatGPT-User|ClaudeBot|Claude-SearchBot|PerplexityBot|Perplexity-User|Google-Extended|Amazonbot|meta-externalagent|GrokBot" access.log
</code></pre>
<p>This covers the major declared bots from <a href="https://developers.openai.com/api/docs/bots">OpenAI</a>, Anthropic, Perplexity, Google, Meta, and xAI. Update the list quarterly — new user agents appear regularly and older ones are occasionally deprecated.</p>
<h3>Step 3: Normalize the fields and load into BigQuery</h3>
<p>Structure each log row with at minimum: <code>timestamp</code>, <code>user_agent</code>, <code>url_path</code>, <code>http_status</code>, <code>bytes_sent</code>, and a derived <code>bot_family</code> label (e.g. "OpenAI", "Anthropic", "Perplexity"). Load the cleaned file into a BigQuery table via the console upload or a scheduled Cloud Function.</p>
<h3>Step 4: Query by day, bot family, and crawl purpose</h3>
<p>Once the data is in BigQuery, a basic aggregation query looks like this:</p>
<pre><code>SELECT
  DATE(timestamp) AS date,
  bot_family,
  COUNT(*) AS total_requests,
  COUNT(DISTINCT url_path) AS unique_urls_crawled,
  COUNTIF(http_status BETWEEN 200 AND 299) AS successful_requests
FROM `your_project.your_dataset.ai_crawl_logs`
GROUP BY date, bot_family
ORDER BY date DESC, total_requests DESC
</code></pre>
<p>Run this as a scheduled query and write the results to a summary table. That summary table becomes your Looker Studio data source.</p>
<h3>Step 5: Connect to Looker Studio and build the dashboard</h3>
<p>In <a href="https://lookerstudio.google.com/">Looker Studio</a>, create a new report and add BigQuery as the data source. Point it at your summary table. From there, build scorecards for total requests and unique URLs crawled, a time-series chart by bot family, and a table of top crawled pages. The whole dashboard takes under an hour once the data is flowing.</p>
<h2>What to Include in the Dashboard</h2>
<p>A good AI crawler dashboard is not a data dump. It is a decision-support tool. Keep it focused on the metrics that actually change your behaviour.</p>
<p><strong>Core metrics to track:</strong></p>
<ul>
<li><strong>Total AI bot requests per day</strong> — the baseline volume metric, split by bot family</li>
<li><strong>Unique URLs crawled per period</strong> — shows which content is being prioritized by each bot category</li>
<li><strong>Crawl purpose breakdown</strong> — separate training, retrieval, and user-action rows so trend lines carry meaning</li>
<li><strong>HTTP status codes by bot</strong> — a spike in 404s or 5XXs from a specific crawler often signals a technical issue worth fixing</li>
<li><strong>Top crawled pages</strong> — which content is drawing the most AI crawler attention, and whether that matches your strategic priorities</li>
</ul>
<p><strong>Secondary metrics worth adding once the basics are stable:</strong></p>
<ul>
<li><strong>Crawl frequency over time</strong> — are bots returning more or less often? A drop in retrieval-bot frequency can be an early signal that your content is being deprioritized</li>
<li><strong>New URLs crawled vs. previously seen URLs</strong> — useful for understanding whether bots are discovering fresh content or re-crawling the same pages</li>
<li><strong>AI referral sessions (from GA4)</strong> — add this as a separate data source and display it alongside the crawler view, but never blend the two into a single metric</li>
</ul>
<p><strong>The pairing that matters most:</strong> put your AI crawler trend line next to your AI referral traffic trend line. If crawl activity is rising but referral traffic is flat, the bots are collecting your content for training, not sending buyers your way. That is useful context for how you prioritize content investment.</p>
<h2>Common Mistakes That Make AI Traffic Reports Useless</h2>
<blockquote>
<p><strong>Warning:</strong> The following mistakes are common enough that they deserve their own section. Each one produces a report that looks credible but actively misleads the people reading it.</p>
</blockquote>
<ul>
<li><strong>Treating every AI-tagged request as a potential customer session.</strong> The vast majority of AI crawler activity is training or indexing. A spike in GPTBot hits does not mean ChatGPT is recommending you to buyers. It means OpenAI's training infrastructure visited your site.</li>
<li><strong>Relying on GA4 alone.</strong> Most AI crawlers do not execute JavaScript, so they leave no trace in GA4. The referral sessions you see from <code>chat.openai.com</code> or <code>perplexity.ai</code> are human visitors arriving from those platforms, not crawler activity. These are two different things.</li>
<li><strong>Merging all AI user agents into one bucket.</strong> GPTBot trains models. OAI-SearchBot powers ChatGPT's search results. ChatGPT-User fires during live queries. Lumping them together is like adding organic search, paid search, and direct traffic into a single "Google" row and drawing conclusions from the total.</li>
<li><strong>Doing data cleaning inside Looker Studio.</strong> As log volume grows, calculated fields and filters inside Looker Studio slow dashboards significantly and make maintenance harder. Clean the data upstream in BigQuery or Sheets, and let Looker Studio do what it does well: display pre-aggregated results.</li>
<li><strong>Never updating the user-agent list.</strong> New bots launch regularly. Grok-DeepSearch, MistralAI-User, and several agentic AI systems appeared or expanded significantly in 2025 alone. A grep pattern that was complete six months ago is probably missing bots today.</li>
</ul>
<h2>FAQ</h2>
<h3>Can GA4 track LLM crawls?</h3>
<p>No, not reliably. GA4 depends on JavaScript to fire tracking events. Most AI crawlers do not render JavaScript, so they leave no trace in GA4. What you see in GA4 under sources like <code>perplexity.ai</code> or <code>chat.openai.com</code> are human referral sessions, not bot activity. Server or CDN logs are the only layer that captures all requests regardless of JavaScript execution.</p>
<h3>Which bots should I prioritize tracking first?</h3>
<p>Start with the user-action category: <strong>ChatGPT-User</strong> (OpenAI), <strong>Perplexity-User</strong> (Perplexity), and <strong>Claude-User</strong> (Anthropic). These fire during live user queries and have the closest relationship to whether an AI engine is actively retrieving and surfacing your content. Training crawlers like GPTBot and ClaudeBot are worth monitoring for volume trends, but a spike in training crawl hits carries far less immediate business signal.</p>
<h3>Do I need Cloudflare to do this?</h3>
<p>No. Cloudflare makes the setup faster and adds verified bot classification, but it is not required. Any web server that writes standard access logs gives you enough raw data to build a functional tracking workflow. The trade-off is that without CDN-level bot verification, you rely more heavily on user-agent matching, which is easier to spoof.</p>
<h3>Will blocking GPTBot hurt my visibility in ChatGPT search results?</h3>
<p>Yes, if you block the wrong bot. <a href="https://developers.openai.com/api/docs/bots">OpenAI's documentation</a> makes clear that GPTBot (training) and OAI-SearchBot (search retrieval) are independently controllable. Blocking GPTBot in <code>robots.txt</code> prevents your content from being used in model training, but it does not affect whether you appear in ChatGPT search results. OAI-SearchBot controls that. You can disallow one without disallowing the other.</p>
<h3>How often should the dashboard update?</h3>
<p>Daily is sufficient for most teams. AI crawler patterns do not change hour to hour in ways that require real-time monitoring. A daily scheduled query in BigQuery that refreshes your Looker Studio summary table overnight gives you a clean, stable view without unnecessary infrastructure overhead. Review the dashboard weekly and do a deeper audit of your user-agent list and crawl patterns quarterly.</p>
]]></content:encoded></item><item><title><![CDATA[How Windgrove's AEO Strategy Drives B2B SaaS Pipeline Through Complex Buyer Journeys]]></title><description><![CDATA[B2B SaaS buyers are no longer starting their research on Google. According to G2's 2025 buyer research, 87% of software buyers say answer engines changed how they research solutions, and 51% now defau]]></description><link>https://windgrove.hashnode.dev/how-windgroves-aeo-strategy-drives-b2b-saas-pipeline-through-complex-buyer-journeys</link><guid isPermaLink="true">https://windgrove.hashnode.dev/how-windgroves-aeo-strategy-drives-b2b-saas-pipeline-through-complex-buyer-journeys</guid><category><![CDATA[aeo]]></category><category><![CDATA[service]]></category><dc:creator><![CDATA[Mitko Dimitrov]]></dc:creator><pubDate>Wed, 20 May 2026 00:00:00 GMT</pubDate><enclosure url="https://cdn.hashnode.com/res/hashnode/image/upload/v1785121041851/8d6bf157-55db-45b6-b95c-28c9ae3f7b1f.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><strong>B2B SaaS buyers are no longer starting their research on Google.</strong> According to <a href="https://www.g2.com/">G2's 2025 buyer research</a>, 87% of software buyers say answer engines changed how they research solutions, and 51% now default to AI chat for software shortlisting. When a VP of Operations asks ChatGPT "what's the best workflow automation tool for mid-market SaaS," your brand either shows up in that answer or it doesn't. There is no page two.</p>
<p>The problem is that most B2B SaaS companies are invisible in AI-generated answers, not because their product is weak, but because AI models lack the domain-specific context to recommend them. The content doesn't exist in the right format, the technical signals aren't there, and no one has told the AI what problem this company actually solves and for whom.</p>
<p>That's the gap Windgrove was built to close.</p>
<blockquote>
<p><strong>Executive Summary:</strong> Windgrove is an AI Visibility Agency that builds systematic AEO programs for B2B SaaS companies. The goal is simple: get your brand cited by ChatGPT, Perplexity, Claude, Gemini, and Grok when your ideal buyers ask the questions that matter most to your pipeline.</p>
</blockquote>
<p><strong>Four things that make Windgrove's approach different:</strong></p>
<ul>
<li><strong>Technical foundation first.</strong> AEO without the right infrastructure doesn't compound. Windgrove starts every engagement with a full technical audit before writing a single word of content.</li>
<li><strong>Data-driven content decisions.</strong> Every blog topic is chosen based on query intent data, ICP fit, and AI citation gap analysis, not gut feel or keyword volume alone.</li>
<li><strong>ICP-specific content.</strong> Windgrove helps clients get precise about who their best-fit buyer is, then writes content specifically for that buyer's questions, language, and decision criteria.</li>
<li><strong>Structured accountability.</strong> Weekly updates, biweekly reports, and monthly strategic calls keep every client's program on track and continuously improving.</li>
</ul>
<h2>Why AI Models Struggle with B2B SaaS</h2>
<p>AI models are excellent at synthesizing general information. They are significantly weaker at understanding nuanced B2B domain expertise, complex sales processes, and the specific differentiation between SaaS vendors in crowded categories.</p>
<p>Ask ChatGPT to recommend a CRM for enterprise sales teams and it will name the same three or four brands it sees cited most often across the web. The brands that get cited aren't necessarily the best products. They're the brands whose content is structured in a way AI engines can parse, extract, and trust.</p>
<p><strong>The uncomfortable truth:</strong> most B2B SaaS companies have years of product expertise and zero AI-readable content to show for it.</p>
<p>This creates a real pipeline problem. <a href="https://www.forrester.com/">According to Forrester</a>, 89% of B2B buyers have adopted generative AI as a primary source of self-guided research. If your brand isn't in those AI-generated answers, you're being cut from consideration before a single sales conversation happens.</p>
<p>The gap is widest in complex sales categories: vertical SaaS, multi-stakeholder deals, enterprise workflows, and niche B2B use cases where buyers ask highly specific questions that generic content can't answer. These are exactly the categories where Windgrove's approach has the most leverage.</p>
<h2>Step One: The Technical AEO Foundation</h2>
<p>Content strategy without a solid technical foundation is like running paid ads to a broken landing page. The traffic might show up. The conversions won't.</p>
<p>Windgrove starts every client engagement with a comprehensive technical AEO audit before any content is planned or published. This audit examines the signals AI engines use to evaluate whether a brand is credible, authoritative, and worth citing.</p>
<h3>What the Technical Audit Covers</h3>
<table>
<thead>
<tr>
<th>Signal Area</th>
<th>What Windgrove Evaluates</th>
</tr>
</thead>
<tbody><tr>
<td>Schema markup</td>
<td>Is structured data present and correctly implemented? Do AI engines understand what your company does and who it serves?</td>
</tr>
<tr>
<td>Site architecture</td>
<td>Can AI crawlers navigate and index your content efficiently?</td>
</tr>
<tr>
<td>Entity clarity</td>
<td>Does your brand have a clear, consistent identity across the web?</td>
</tr>
<tr>
<td>Content structure</td>
<td>Are pages formatted with question-answer patterns, headers, and extractable data?</td>
</tr>
<tr>
<td>Backlink authority</td>
<td>Do authoritative third-party sources reference and validate your brand?</td>
</tr>
</tbody></table>
<p>Most B2B SaaS sites fail on at least three of these five dimensions. A site can rank reasonably well on Google while remaining almost entirely invisible to AI answer engines, because the two systems evaluate content differently.</p>
<p><strong>Google ranks pages. AI engines synthesize answers.</strong> The technical requirements are not the same.</p>
<p>Once the audit is complete, Windgrove builds the infrastructure that allows every subsequent piece of content to compound. Schema gets implemented correctly. Site architecture gets cleaned up. Entity signals get strengthened across the web. Only then does content work get started, because without the foundation, even the best-written blog post won't get cited.</p>
<h2>Data-Driven Content: How Windgrove Decides What to Write</h2>
<p>Publishing content without a targeting strategy is one of the most common and expensive mistakes in B2B marketing. Most companies write about what their team finds interesting, or what a competitor just published, or what a keyword tool says has high search volume. None of those inputs tell you what your buyers are actually asking AI models right now.</p>
<p>Windgrove approaches content selection differently.</p>
<h3>Query Intent Research</h3>
<p>Before any blog topic gets approved, Windgrove runs a query intent analysis. This means identifying the exact questions your ideal buyers are typing into ChatGPT, Perplexity, and Google, and mapping those queries to stages in your sales cycle.</p>
<p>A buyer at the awareness stage asks: "What causes [problem]?" A buyer at the consideration stage asks: "What's the best solution for [problem]?" A buyer at the decision stage asks: "How does [your product] compare to [competitor]?" Each of those queries requires a different type of content, and each one represents a different moment in the pipeline.</p>
<p><strong>The goal is to have a credible, citable answer for every question your ICP is asking at every stage of their journey.</strong></p>
<h3>Why Topic Selection Is a Strategic Decision</h3>
<p>Not all blog topics are equal from an AEO standpoint. Windgrove evaluates each potential topic against three criteria:</p>
<ul>
<li><strong>AI citation gap:</strong> Is this a question where AI models currently give a weak or generic answer? If so, there's a real opportunity to become the cited source.</li>
<li><strong>ICP relevance:</strong> Does this query come from the type of buyer the client actually wants to close? High-volume topics that attract the wrong audience waste everyone's time.</li>
<li><strong>Pipeline stage alignment:</strong> Does this content move a buyer closer to a conversation, or does it just generate impressions?</li>
</ul>
<p>Only topics that pass all three filters make it into the content calendar. This is what separates an AEO content program from a generic blog strategy.</p>
<h2>Writing for Your ICP, Not the Algorithm</h2>
<p>Most B2B content is written for a vague, imaginary reader. It tries to appeal to everyone and ends up resonating with no one. AI engines can detect this. Generic content gets deprioritized in favour of content that demonstrates specific, credible expertise for a specific audience.</p>
<p>Windgrove's content work begins with a critical upstream question: who is your best-fit customer, and what do they specifically need to believe to choose you?</p>
<h3>Getting Clear on the ICP</h3>
<p>Before writing a single piece of content, Windgrove works with clients to sharpen their Ideal Customer Profile. This isn't a demographic exercise. It's a strategic one. The questions that matter are:</p>
<ul>
<li>What is the specific problem this buyer is trying to solve, in their own words?</li>
<li>What does their buying committee look like? Who is the economic buyer, and who is the technical evaluator?</li>
<li>What objections do they raise before signing? What do they need to see to feel confident?</li>
<li>What language do they use when they describe the problem? Not your language. Theirs.</li>
</ul>
<p>That last point is where most B2B content fails. Companies write about their product using their internal vocabulary. Buyers search using their own vocabulary. The gap between those two languages is where AI visibility is lost.</p>
<h3>ICP-Specific Content in Practice</h3>
<p>Once the ICP is defined, every piece of content is written to speak directly to that buyer's reality. This means:</p>
<ul>
<li>Using the exact phrasing and terminology your buyer uses in their searches</li>
<li>Addressing the specific objections and comparison questions that come up in your sales cycle</li>
<li>Structuring answers in a way that maps to how that buyer evaluates vendors</li>
</ul>
<p>The result is content that doesn't just rank or get cited. It arrives pre-loaded with the context a buyer needs to take the next step. <a href="https://previsible.io/">According to the Previsible AI Traffic Report</a>, visitors arriving from AI-generated answers convert at 4.4x the rate of standard organic search visitors, because the AI has already done the pre-qualification work.</p>
<p>That conversion advantage only materializes when the content was written for the right buyer in the first place.</p>
<h2>Accountability Built Into the Program</h2>
<p>AEO is not a set-it-and-forget-it channel. AI models update their training data, citation patterns shift, and what gets cited today may need to be refreshed in 90 days. The brands that sustain AI visibility are the ones that treat it as an ongoing program, not a one-time project.</p>
<p>Windgrove builds structured accountability into every client engagement through a three-tier communication cadence.</p>
<h3>The Reporting Cadence</h3>
<p><strong>Weekly updates</strong> keep clients informed on what was published, what technical changes were implemented, and what early signals are showing up in AI citation tracking. No surprises. No black boxes.</p>
<p><strong>Biweekly reports</strong> go deeper. These include AI visibility metrics across ChatGPT, Perplexity, Claude, Gemini, and Grok, query performance data, content that's gaining traction, and content that needs iteration. The report is designed to answer one question: is the program moving the needle?</p>
<p><strong>Monthly strategic calls</strong> are where the bigger picture gets reviewed. Windgrove and the client align on ICP refinements, new query opportunities, competitive shifts in AI citations, and content priorities for the next 30 to 60 days. These calls are not status updates. They are strategic working sessions.</p>
<blockquote>
<p>"Most B2B brands show up in fewer than 20% of relevant AI responses, even when they rank well in Google." — Optimist, 2026</p>
</blockquote>
<p>That number moves. The reporting cadence is what ensures it moves in the right direction.</p>
<h3>What Gets Measured</h3>
<table>
<thead>
<tr>
<th>Metric</th>
<th>What It Tells You</th>
</tr>
</thead>
<tbody><tr>
<td>Share of Model Response</td>
<td>How often your brand appears in AI answers for your key buyer queries</td>
</tr>
<tr>
<td>AI citation count</td>
<td>How many times your content is cited as a source across AI engines</td>
</tr>
<tr>
<td>Query coverage</td>
<td>What percentage of your ICP's key questions you have content for</td>
</tr>
<tr>
<td>Pipeline attribution</td>
<td>Which AI-sourced visits convert to demo requests or qualified leads</td>
</tr>
</tbody></table>
<p>The goal is not impressions. The goal is pipeline. Every metric in Windgrove's reporting framework traces back to that outcome.</p>
<h2>The AI Search Interface Your Buyers Are Using</h2>
<p>This is where your buyers are researching vendors right now. Perplexity alone captures nearly 20% of AI search traffic in the US and is growing 25% every four months. When a buyer asks a question here, the AI synthesizes an answer from sources it deems authoritative. Your content either makes that cut or it doesn't.</p>
<p><img src="https://cdn.hashnode.com/uploads/posts/6a6513b3e2d2908b72dc5cf1/27b5640a-7529-4c67-aaea-21a5b57dccda.jpg" alt="Perplexity AI answer engine showing how B2B SaaS vendors are cited in AI-generated search responses" /></p>
<p>Only 11% of domains are cited by both ChatGPT and Perplexity, which means a multi-model AEO strategy isn't optional. Each engine has different citation preferences, and Windgrove's approach accounts for all of them.</p>
<h2>The Window Is Still Open. Not for Long.</h2>
<p>The AEO landscape in most B2B SaaS categories is still wide open. Your competitors have not built this infrastructure. Most are still running the same SEO playbook they used in 2022, optimizing for blue links while their buyers migrate to AI-generated answers.</p>
<p>That is a narrow window.</p>
<p><a href="https://zeroclicklabs.ai/2026-ai-seo-statistics/">According to ZeroClick Labs' 2026 AI SEO report</a>, only five brands capture roughly 80% of AI responses in any given category. The brands that move first, build the technical foundation, and publish the right content for the right ICP will hold those positions. The brands that wait will find themselves in the same position most companies found themselves in with SEO in 2015: playing catch-up in a game where the early movers already own the territory.</p>
<p>Windgrove exists to help B2B SaaS companies claim that territory before someone else does.</p>
<p><strong>Ready to see where your brand stands in AI-generated answers?</strong> <a href="https://windgrove.ai/">Book a discovery call with Windgrove</a> and get a clear picture of your current AI visibility, your biggest gaps, and what it would take to close them.</p>
<h2>Frequently Asked Questions</h2>
<p><strong>What is AEO and how is it different from SEO?</strong></p>
<p>Answer Engine Optimization (AEO) is the practice of structuring your content and technical infrastructure so that AI engines like ChatGPT, Perplexity, Claude, Gemini, and Grok cite your brand when buyers ask relevant questions. SEO optimizes for ranking in a list of links. AEO optimizes for being selected as the trusted source inside a synthesized answer. The two disciplines overlap but require different strategies, and AEO is increasingly where B2B buyer decisions are being made before a vendor is ever contacted.</p>
<p><strong>How long does it take to see results from AEO?</strong></p>
<p>Most clients begin seeing measurable improvements in AI citation frequency within 60 to 90 days of the technical foundation work being completed. Content compounds over time, meaning the program becomes significantly more effective at the 6-month and 12-month marks as topical authority builds across more of the buyer's query landscape. AEO is a compounding channel, not a one-time campaign.</p>
<p><strong>Do I need to abandon my existing SEO strategy to invest in AEO?</strong></p>
<p>No. AEO is an extension of SEO, not a replacement. The technical foundations overlap significantly, and well-structured content that performs in AI-generated answers typically also performs well in traditional search. Windgrove's approach is built to strengthen both channels simultaneously, which is why the technical audit comes first.</p>
<p><strong>How does Windgrove decide which blog topics to prioritize?</strong></p>
<p>Every topic is evaluated against three filters: whether there is an AI citation gap (meaning AI models currently give a weak answer for this query), whether the query comes from your ICP, and whether the content will move a buyer closer to a pipeline conversation. Topics that don't pass all three filters don't make the calendar. This is a fundamentally different process from standard keyword research.</p>
<p><strong>What does working with Windgrove actually look like week to week?</strong></p>
<p>Clients receive weekly updates on publishing activity and early citation signals, biweekly reports covering AI visibility metrics across all major models, and monthly strategic calls to review performance, refine the ICP, and set content priorities for the next cycle. The program is designed to be transparent, measurable, and continuously improving. You always know what's been done, what's working, and what's next.</p>
]]></content:encoded></item><item><title><![CDATA[How to Track Inbound Leads from AI Recommendations (GA4, Form Fields, and What the Data Is Telling You)]]></title><description><![CDATA[Executive Summary
AI tools like ChatGPT and Perplexity are sending inbound leads to your website right now. Most businesses have no idea. GA4 buries this traffic inside generic "Referral" or misfiles ]]></description><link>https://windgrove.hashnode.dev/how-to-track-inbound-leads-from-ai-recommendations-ga4-form-fields-and-what-the-data-is-telling-you</link><guid isPermaLink="true">https://windgrove.hashnode.dev/how-to-track-inbound-leads-from-ai-recommendations-ga4-form-fields-and-what-the-data-is-telling-you</guid><category><![CDATA[#leads]]></category><category><![CDATA[tracking]]></category><dc:creator><![CDATA[Mitko Dimitrov]]></dc:creator><pubDate>Tue, 19 May 2026 00:00:00 GMT</pubDate><content:encoded><![CDATA[<h2>Executive Summary</h2>
<p>AI tools like ChatGPT and Perplexity are sending inbound leads to your website right now. Most businesses have no idea. GA4 buries this traffic inside generic "Referral" or misfiles it as "Direct," making an entire acquisition channel invisible. This article shows you exactly how to surface it, measure it, and act on it.</p>
<ul>
<li>AI-referred traffic converts at up to 16% — nearly nine times higher than Google Organic — making it the highest-intent inbound channel most teams aren't measuring.</li>
<li>GA4 has no native AI channel. Without a custom channel group, sessions from ChatGPT, Perplexity, and Gemini are scattered across Referral and Direct, invisible in your standard reports.</li>
<li>"How did you find us?" form fields capture the attribution GA4 can't — including dark traffic from mobile AI apps where referrer data is stripped entirely.</li>
<li>Businesses that build this tracking infrastructure now will have a compounding data advantage as AI referral volume grows. The ones that wait start from zero in a more competitive landscape.</li>
</ul>
<h2>The Channel Your Analytics Aren't Showing You</h2>
<p>A prospect types "best AI visibility agency for B2B SaaS" into ChatGPT. ChatGPT recommends your company by name. The prospect clicks through, reads your site, and books a call.</p>
<p>Your GA4 report shows it as "Direct."</p>
<p>This is happening at scale. <a href="https://seranking.com/blog/ai-traffic-research-study/">According to a 2025 study by SE Ranking</a>, visitors referred by AI platforms spend 68% more time on websites than those from traditional organic search. AI tools act as intent filters, sending people who are already deep in the decision process. They're not browsing. They're evaluating.</p>
<p><strong>The conversion numbers are striking.</strong> A case study by Seer Interactive found that traffic referred from ChatGPT converts at 16%, compared to 1.8% from Google Organic. That's nearly nine times higher. The volume is still small relative to total traffic, but the quality signal is extraordinary.</p>
<p>Yet most businesses are flying blind on this channel. GA4 was built for a world of search engines, social platforms, and email campaigns. There is no native "AI" channel. When ChatGPT sends someone to your site, GA4 does one of two things:</p>
<ul>
<li><strong>If the referrer header is passed</strong> (typical on desktop), it files the session under generic "Referral" alongside forum backlinks and directory listings.</li>
<li><strong>If the referrer is stripped</strong> (common in mobile apps and in-app browsers), it records the session as "Direct," indistinguishable from someone typing your URL.</li>
</ul>
<p>The result is that one of your fastest-growing inbound channels is invisible in your standard reports.</p>
<blockquote>
<p>"There is no row in GA4 that says 'AI Traffic' by default. That's where a custom channel grouping comes in." — <a href="https://measureu.com/ai-traffic-ga4/">MeasureU</a></p>
</blockquote>
<p>The good news: fixing this takes about 15 minutes. The better news: once you fix it, you'll also understand which AI engines are sending you the most qualified traffic, and you can start optimizing for them.</p>
<h2>Why This Matters Now: The AI Referral Traffic Landscape</h2>
<p>Before diving into the setup, it's worth understanding the scale of what's already happening.</p>
<p>ChatGPT now has <a href="https://technologychecker.io/blog/chatgpt-statistics">900 million weekly active users</a> as of early 2026, more than double the 400 million reported a year earlier. It holds the #10 global domain ranking, ahead of Amazon, Instagram, and YouTube. ChatGPT accounts for roughly 78% of all AI-driven referral traffic worldwide.</p>
<p>Perplexity is the second-largest source, driving 15% of global AI traffic and nearly 20% in the United States. Gemini is growing fast: its referral traffic surged 388% year-over-year between September and November 2025, <a href="https://digiday.com/media/in-graphic-detail-the-state-of-ai-referral-traffic-in-2025/">according to Similarweb data reported by Digiday</a>.</p>
<h3>The Quality vs. Volume Equation</h3>
<p>Total AI referral volume is still modest. Across 10 major industries, AI platforms drive roughly 1% of overall web traffic, according to research by Conductor. That number will feel small until you look at what that traffic does when it arrives.</p>
<table>
<thead>
<tr>
<th>Traffic Source</th>
<th>Sign-up Conversion Rate</th>
</tr>
</thead>
<tbody><tr>
<td>LLM / AI referral</td>
<td>1.66%</td>
</tr>
<tr>
<td>Social media</td>
<td>0.46%</td>
</tr>
<tr>
<td>Organic search</td>
<td>0.15%</td>
</tr>
<tr>
<td>Direct</td>
<td>0.13%</td>
</tr>
</tbody></table>
<p><em>Source: Microsoft Clarity analysis of 1,200+ publisher websites</em></p>
<p>LLM-referred visitors convert to sign-ups at more than 10 times the rate of organic search visitors. For subscription conversions, the gap is similar: LLM traffic converts at 1.34%, compared to 0.55% for search.</p>
<p><strong>This is the intent filter effect.</strong> When someone asks ChatGPT "which agency should I hire for AI search optimization," they're not in research mode. They've already decided to act. They're asking for a recommendation. The AI gives them one. If your brand is in that answer, the visitor who clicks through is pre-sold in a way that a Google click-through almost never is.</p>
<h3>Which AI Engines Should You Track?</h3>
<p>Not all AI platforms pass referrer data equally. Here's the practical breakdown:</p>
<ul>
<li><strong>ChatGPT (chatgpt.com / chat.openai.com):</strong> Largest source by far. Passes referrer on desktop. Mobile app often strips it.</li>
<li><strong>Perplexity (perplexity.ai):</strong> Strong North American user base. Consistent referrer passing. Highly engaged traffic.</li>
<li><strong>Gemini (gemini.google.com):</strong> Fast-growing. Passes referrer reliably on desktop.</li>
<li><strong>Microsoft Copilot (copilot.microsoft.com):</strong> Embedded in Windows and Edge. Referrer passing is inconsistent.</li>
<li><strong>Claude (claude.ai):</strong> Smaller share but growing. Anthropic's crawl-to-referral ratio is extremely high, meaning it sends far less traffic than it crawls.</li>
</ul>
<p>Track all five. The channel group you're about to build will capture them all in one row.</p>
<h2>Step 1: Build a Custom AI Channel Group in GA4</h2>
<p>This is the foundational fix. It takes 15 minutes and gives you a dedicated "AI Traffic" row in your acquisition reports from that point forward. It does not recover historical data, but it starts your baseline immediately.</p>
<h3>Create the Channel Group</h3>
<ol>
<li>In GA4, click <strong>Admin</strong> (the cog icon, bottom-left of the screen).</li>
<li>Under <strong>Data display</strong>, click <strong>Channel groups</strong>.</li>
<li>Click <strong>Create new channel group</strong> (top right).</li>
<li>Name it something clear: "AI Search" or "AI Traffic" works well.</li>
</ol>
<h3>Add the AI Channel with Regex Matching</h3>
<ol>
<li>Click <strong>Add new channel</strong>.</li>
<li>Name the channel <strong>AI Search</strong>.</li>
<li>Under conditions, set one rule:<ul>
<li><strong>Dimension:</strong> Source</li>
<li><strong>Match type:</strong> Matches regex</li>
<li><strong>Value:</strong> paste the regex below</li>
</ul>
</li>
</ol>
<pre><code>chatgpt\.com|chat\.openai\.com|perplexity\.ai|claude\.ai|gemini\.google\.com|copilot\.microsoft\.com|grok\.x\.com|mistral\.ai|you\.com|phind\.com|deepai\.org|anthropic\.com
</code></pre>
<ol>
<li><strong>Do not restrict the medium to "referral."</strong> Leaving the medium unrestricted is critical. Constraining to <code>medium = referral</code> causes you to miss ChatGPT sessions that arrive as <code>chatgpt.com / (none)</code> or <code>chatgpt.com / (not set)</code>. Based on real-world data, this represents 30-40% of visible ChatGPT traffic.</li>
<li>Click <strong>Save channel</strong>.</li>
</ol>
<h3>Get the Channel Order Right</h3>
<p>This is where most setups fail. GA4 evaluates channel rules top to bottom. The first matching rule wins. If your new AI Search channel sits below "Referral" in the list, the broader Referral rule claims AI traffic first and your new channel never fires.</p>
<ol>
<li>Back on the channel group screen, click <strong>Reorder</strong>.</li>
<li>Drag <strong>AI Search</strong> to the very top of the list, above Organic Search, Direct, and Referral.</li>
<li>Click <strong>Apply</strong>, then <strong>Save Group</strong>.</li>
</ol>
<h3>View Your Data</h3>
<p>Allow 24-48 hours for GA4 to reprocess. Then go to:</p>
<p><strong>Reports → Acquisition → Traffic Acquisition</strong></p>
<p>At the top-left of the table, change the channel grouping dropdown from "Session default channel group" to your new custom group. You'll now see AI Search as its own row.</p>
<blockquote>
<p><strong>What you're looking at:</strong> sessions, engaged sessions, conversions, and revenue attributed to AI referrals. Segment by source to see which AI platform is driving the most traffic. Compare conversion rates across AI platforms vs. organic search.</p>
</blockquote>
<h2>Step 2: Check Your Referral Report for Existing AI Traffic</h2>
<p>Before your new channel group is fully populated, check what AI traffic is already sitting in your existing Referral data. You may have months of unattributed AI sessions hiding in plain sight.</p>
<h3>Build a Dedicated Referral Report</h3>
<ol>
<li>Go to <strong>Reports → Library</strong>.</li>
<li>Select <strong>Create new report → Create detail report</strong>.</li>
<li>Choose <strong>Traffic Acquisition Report</strong> as your template.</li>
<li>In the top-right corner, click <strong>Dimensions</strong>. Remove all dimensions except <strong>Session source/medium</strong>, then save.</li>
<li>Click <strong>+ Filter</strong>.</li>
<li>From the Dimensions dropdown, select <strong>Session default channel group</strong>.</li>
<li>Set Match type to <strong>Exactly matches</strong> and select <strong>Referral</strong>.</li>
<li>Apply the filter and save the report with a clear name like "Referral Traffic Report."</li>
<li>Publish it to your GA4 Library so it's accessible without rebuilding each time.</li>
</ol>
<h3>What to Look For</h3>
<p>Once the report is live, scan the source/medium column for AI domains. Common entries include:</p>
<ul>
<li><code>chatgpt.com / referral</code></li>
<li><code>chat.openai.com / referral</code></li>
<li><code>perplexity.ai / referral</code></li>
<li><code>claude.ai / referral</code></li>
<li><code>gemini.google.com / referral</code></li>
</ul>
<p><strong>If you see these domains, you already have AI traffic.</strong> It's been there. You just weren't measuring it. Cross-reference those sessions against your conversions to understand what they were doing on your site.</p>
<p>If no AI domains appear in your referral report, it doesn't necessarily mean AI engines aren't recommending you. It may mean your referrer data is being stripped, which brings us to the layer GA4 cannot fix on its own.</p>
<h2>Step 3: Add "How Did You Find Us?" to Every Lead Form</h2>
<p>GA4 can only see what the browser tells it. When an AI assistant strips the referrer header, or when someone reads a ChatGPT recommendation on their phone and then visits your site from their laptop later, that attribution is gone. No analytics tool can recover it.</p>
<p>The only way to capture this signal is to ask.</p>
<p>A simple "How did you find us?" field on your contact form, demo request page, or intake questionnaire captures attribution data that no analytics platform can. It also captures something more valuable: the exact language the lead used to find you.</p>
<h3>How to Set It Up</h3>
<p>Add a dropdown or free-text field to your primary lead forms. Keep it optional but prominent. Suggested dropdown options:</p>
<ul>
<li>Google search</li>
<li>LinkedIn</li>
<li>ChatGPT or another AI tool</li>
<li>Perplexity</li>
<li>Referral from a colleague</li>
<li>Podcast or webinar</li>
<li>Other</li>
</ul>
<p><strong>If you use a free-text field instead of a dropdown</strong>, you'll get more nuanced data. Leads will often write things like "ChatGPT recommended you" or "I asked Perplexity for the best [service] and you came up." That language is gold. It tells you exactly which queries are driving recommendations, which you can then use to optimize your content and positioning.</p>
<h3>Connect the Data</h3>
<p>Route form responses into your CRM and tag them by source. Over time, you'll be able to:</p>
<ul>
<li>Compare close rates between AI-referred leads and organic search leads</li>
<li>Identify which AI platforms produce the highest-value customers</li>
<li>Track whether your AI visibility efforts are translating into revenue, not just sessions</li>
</ul>
<p><strong>This is the attribution layer GA4 cannot provide.</strong> GA4 tells you someone arrived from ChatGPT. Your form field tells you what they asked ChatGPT to find them. Those are two different pieces of intelligence, and you need both.</p>
<h3>The "Dark Traffic" Problem</h3>
<p>A meaningful portion of AI-referred traffic will never show a referrer. Mobile app sessions, in-app browsers, and certain link-handling behaviours strip the referrer entirely. This traffic lands in your "Direct" bucket.</p>
<p>If you notice your "Direct" traffic converting at unusually high rates, especially for sessions with engagement patterns that don't look like branded searches (long session duration, multiple pages visited, direct navigation to your pricing or contact page), some of it is almost certainly AI-referred. The form field is the only reliable way to confirm it.</p>
<h2>Step 4: Interpret the Data and Act on It</h2>
<p>Tracking is only valuable if it changes what you do. Once your GA4 channel group and form field data are running, here's how to turn the numbers into decisions.</p>
<h3><a href="https://windgrove.ai/blog/aeo-competitor-benchmarking">Benchmark Your</a> AI Traffic Monthly</h3>
<p>Set a recurring calendar reminder to review your AI Search channel data every 30 days. Track:</p>
<ul>
<li><strong>Total sessions from AI sources</strong> (month-over-month trend)</li>
<li><strong>Conversion rate from AI traffic</strong> vs. organic search and direct</li>
<li><strong>Which AI platform</strong> is sending the most sessions</li>
<li><strong>Which pages</strong> AI-referred visitors land on and engage with most</li>
</ul>
<p>If your AI traffic is growing but your conversion rate is low, the problem is likely on-page. AI engines are recommending you, but your landing experience isn't matching the intent of the recommendation. That's a fixable problem.</p>
<h3>Identify Your Best-Performing AI Queries</h3>
<p>From your form field data, look for patterns in how leads describe finding you. Group responses by:</p>
<ul>
<li><strong>Platform mentioned</strong> (ChatGPT vs. Perplexity vs. Gemini)</li>
<li><strong>Query type</strong> (comparison queries, "best X for Y" queries, specific problem queries)</li>
<li><strong>Lead quality</strong> (did this lead close? what was the deal size?)</li>
</ul>
<p>The queries that produce closed deals are the ones you need to own in AI responses. That means creating content that directly and authoritatively answers those exact questions.</p>
<h3>Use AI Traffic Data to Justify Visibility Investment</h3>
<p>This is the business case your leadership team needs. AI referral traffic is not a vanity metric. It is a measurable acquisition channel with documented conversion rates that outperform every other inbound source in most industries.</p>
<p>When you can show that AI-referred leads convert at 16% compared to 1.8% from Google, the conversation about investing in AI visibility shifts from "is this real?" to "how much should we allocate?"</p>
<p><strong>Track this data from day one.</strong> Even if your current AI traffic volume is low, establishing the baseline now means you'll have the before/after comparison when your visibility investments compound.</p>
<h3>Watch for the Compounding Effect</h3>
<p>AI visibility compounds in a way that paid search does not. When your brand is consistently cited in AI responses, it builds a feedback loop: more citations lead to more training data exposure, which leads to more recommendations, which leads to more inbound traffic, which generates more brand signals that reinforce future citations.</p>
<p>The businesses that start tracking and optimizing now will have a structural advantage that is very difficult to close later.</p>
<h2>What to Do This Week: Actionable Summary</h2>
<p>If you read nothing else, do these five things.</p>
<p><strong>1. Set up your GA4 channel group today.</strong> It takes 15 minutes. Use the regex in Step 1, drag AI Search to the top of the channel order, and you'll have clean data flowing within 48 hours. Every week you wait is a week of baseline data you'll never get back.</p>
<p><strong>2. Pull your existing Referral report.</strong> Before your new channel group is populated, scan your current referral data for AI domains. You may already have months of ChatGPT or Perplexity traffic sitting in your analytics unnoticed.</p>
<p><strong>3. Add "How did you find us?" to your primary lead form.</strong> One field. Optional. Dropdown or free-text. This single change will surface attribution data that no analytics platform can provide, including the exact queries driving recommendations.</p>
<p><strong>4. Tag AI-sourced leads in your CRM.</strong> Once form responses start coming in, tag them by source. Track close rate, deal size, and time-to-close separately for AI-referred leads. The numbers will make the business case for you.</p>
<p><strong>5. Review monthly and act on what you find.</strong> AI visibility is not set-and-forget. Check which platforms are growing, which pages AI-referred visitors engage with most, and which queries are producing closed deals. Build content around those queries. Repeat.</p>
<blockquote>
<p>The businesses that build this infrastructure now will have a compounding advantage that is very difficult to close later. The ones that wait will be building from scratch in a more competitive landscape.</p>
</blockquote>
<hr />
<h2>The Attribution Gap Is a Strategy Gap</h2>
<p>Most businesses treat AI visibility as a brand awareness play. Something nice to have, hard to measure, and difficult to justify in a quarterly review.</p>
<p>The data tells a different story.</p>
<p>AI-referred visitors convert at rates that make every other inbound channel look ordinary. The problem has never been the channel's value. The problem has been attribution. Without a clean way to see AI traffic in your analytics, you cannot make the case for investing in it, and you cannot optimize what you cannot measure.</p>
<p>The GA4 channel group and the form field together close that gap. They give you the quantitative signal (sessions, conversions, revenue by AI source) and the qualitative signal (which queries, which platforms, which intent patterns). Used together, they build the business case for AI visibility as a core acquisition channel, not a side experiment.</p>
<p><strong>The window to build this advantage is still open.</strong> Most of your competitors have not set up this tracking. Most have not asked their leads how they found them. Most are still treating AI recommendations as unmeasurable.</p>
<p>That's your edge. Use it.</p>
<hr />
<p><strong>Ready to see what AI engines are saying about your brand?</strong> Windgrove AI offers a free consultation to audit your current AI visibility, identify where you're being cited (and where you're being overlooked), and build a roadmap to compound your inbound from AI recommendations. <a href="https://app.searchable.com/content/f7c18bdc-dbe7-439d-a641-907141d5800d#">Book your free call here</a> and we'll show you exactly where you stand.</p>
<h2>Frequently Asked Questions</h2>
<h3>1. Will the GA4 channel group capture all my AI traffic?</h3>
<p>No, and it's important to understand why. When an AI platform strips the referrer header (common in mobile apps and certain in-app browsers), GA4 records the session as "Direct" with no source information. The custom channel group only captures sessions where the AI platform passes a referrer. In practice, this covers the majority of desktop AI traffic. Mobile traffic from AI apps is largely invisible to GA4, which is why the "How did you find us?" form field is an essential complement, not an optional add-on.</p>
<h3>2. How do I know if my brand is actually being recommended by AI engines?</h3>
<p>The most direct method is to test it yourself. Open ChatGPT, Perplexity, Claude, and Gemini and ask the kinds of questions your target buyers would ask. If your brand appears in the responses, you're being recommended. If it doesn't, you have an AI visibility gap. You can also monitor your GA4 referral data for AI domains appearing as traffic sources. A simpler signal: if leads are showing up in your form field data mentioning AI tools by name, your brand is in the conversation.</p>
<h3>3. My AI traffic volume is very low. Is it worth tracking?</h3>
<p>Yes, for two reasons. First, the conversion quality of AI-referred traffic is so high that even small volumes can represent meaningful revenue. A handful of AI-referred leads converting at 16% is worth more than hundreds of organic search visitors converting at 1.8%. Second, you need a baseline. AI referral traffic is growing fast. The businesses that establish clean tracking now will have comparative data that proves ROI when volume increases. Starting late means starting without context.</p>
<h3>4. What's the difference between tracking AI traffic in GA4 and tracking AI visibility?</h3>
<p>They measure different things. GA4 tracks the traffic that AI platforms send to your website, which requires a user to click a link in an AI response. AI visibility refers to whether your brand appears in AI-generated answers at all, including answers that don't include clickable links. Many AI responses recommend brands without linking to them directly. A user might read a ChatGPT recommendation, close the app, and search your brand name on Google. That session shows as organic branded search in GA4, not AI referral. Tracking both GA4 referrals and form field responses gives you a more complete picture, but neither captures 100% of the influence AI engines have on your inbound pipeline.</p>
<h3>5. How do I get AI engines to recommend my brand more often?</h3>
<p>AI engines recommend brands that appear authoritative, well-cited, and clearly relevant to the query. The core levers are: structured, authoritative content that directly answers the questions your buyers ask; consistent brand mentions across credible third-party sources; schema markup that makes your brand's expertise machine-readable; and a content strategy built around the specific queries where you want to appear. This is the discipline of Answer Engine Optimization (AEO). It's where Windgrove AI specializes. If you want a concrete assessment of where your brand stands in AI-generated answers, <a href="https://app.searchable.com/content/f7c18bdc-dbe7-439d-a641-907141d5800d#">book a free consultation</a> and we'll walk you through it.</p>
]]></content:encoded></item></channel></rss>