All articles

Your FAQ Schema Died in May. Reddit and YouTube Are What's Getting You Into ChatGPT.

Everyone spent two years marking up FAQ pages for AI search. Google killed the feature on May 7, 2026, then published guidance saying the markup was never the point. Here is where the visibility actually comes from.

Short answer: FAQ schema markup never drove AI visibility. It powered a Google display feature that was deprecated on May 7, 2026. What actually determines whether an AI names your business is the retrieval corpus, and for most categories that corpus is dominated by two platforms you do not own: Reddit and YouTube.

Key takeaways

  • Google deprecated FAQ rich results for all site types on May 7, 2026. Search Console reporting ended in June, API data in August.
  • On May 15, 2026, Google confirmed AI Overviews and AI Mode use the same ranking systems as normal Search, and that no special structured data, llms.txt file, or AI-specific rewriting is required.
  • Reddit is the most-cited domain overall, but its weight ranges from roughly 47% of citations in Perplexity to about 3% in Gemini.
  • YouTube mentions carry a 0.737 correlation with brand appearances in ChatGPT, AI Mode, and AI Overviews, the strongest single signal in Ahrefs' 75,000-brand study.
  • Views and subscriber counts have near-zero correlation with AI citations. Transcript quality is what matters.
  • On your own site, the peer-reviewed levers are citing external sources (up to +40%), adding statistics (+30 to 41%), and adding quotations (+22 to 28%). Keyword stuffing produces roughly nothing.

What exactly did Google change on May 7, 2026?

Google stopped showing FAQ rich results in Search for every site type. The change arrived as a deprecation banner on a developer documentation page, with no blog post and no explanation.

Those were the expandable question-and-answer dropdowns that appeared under your search listing. Search Console reporting for them was removed in June 2026. The API data goes dark in August 2026.

Eight days later, on May 15, Google published its first official guidance on optimizing for generative AI features. Its position was blunter than the industry expected: AI Overviews and AI Mode run on the same core ranking and quality systems as regular Search. There is no separate AI algorithm. And you do not need llms.txt files, AI-specific content chunking, machine-only rewrites, or special structured data to appear in them.


Do I still need FAQ schema markup?

Not for Google rich results, because that display feature no longer exists. The markup remains harmless and other crawlers still parse it, so there is no urgency to remove it.

FAQPage is still a valid Schema.org type, and Google has confirmed that unused markup causes no problems. But it was never the lever. It was a display feature wearing a strategy costume.

The distinction that matters: FAQ markup is deprecated. FAQ content, meaning real customer questions answered directly in your page text, is more valuable than ever, because that is the format answer engines lift from.


How does my content actually reach an AI model?

Three ways, and only two are things you can influence this quarter. Conflating them is where most GEO advice goes wrong.

  1. Training data. Baked into model weights, years stale, and effectively unaddressable in a quarterly plan.
  2. Licensed data feeds. Commercial agreements giving specific vendors structured, real-time access to specific platforms.
  3. Live retrieval, also called RAG or grounding. The model runs searches at query time, pulls documents, and grounds its answer in them. Google calls this retrieval-augmented generation and confirms its AI features use it, alongside "query fan-out," which issues multiple related sub-searches across subtopics before writing a single answer.

Numbers 2 and 3 are the addressable ones. Both reward the same thing: being present, specific, and quotable on the domains those systems reach for.

Traditional SEO competes for a position on a results page. AI search competes for something structurally different: being the source a model selects, quotes, and attributes when it composes an answer.


Why is Reddit the most-cited domain in AI answers?

Because Reddit threads contain what brand websites structurally lack: named products, direct comparisons, stated trade-offs, first-person outcomes, and community-validated ranking through upvotes.

When a model needs to answer "is X actually worth it," a marketing page offers adjectives. A Reddit thread offers evidence.

The scale of Reddit's position is large, and while the numbers vary by methodology, the direction never does:

  • A synthesis of six citation-tracking studies covering 680M+ citations (Everything-PR Research, April 2026) put Reddit at roughly 40% of aggregate multi-engine citation frequency, the single most-cited domain.
  • Peec AI's analysis of 30 million directly-cited sources (May 2026) also ranked Reddit first across engines.
  • SE Ranking's 129,000-domain study found domains with heavy Reddit brand-mention volume averaged 7 ChatGPT citations versus 1.8 for domains with minimal Reddit presence, a 3.9x multiplier.

Does Reddit work equally well across every AI engine?

No, and this is where most GEO advice is wrong. Reddit's weight ranges from roughly 47% of citations in Perplexity down to about 3% in Gemini.

Reddit signed licensing agreements with Google and OpenAI in 2024, reportedly worth around $60M annually. It took the opposite approach with everyone else: Reddit sued Anthropic in June 2025 over unauthorized scraping to train Claude, and in October 2025 sued Perplexity along with data intermediaries SerpApi, Oxylabs, and AWMProxy, alleging they reconstituted Reddit content by scraping Google's search results rather than licensing it directly.

EngineReddit's share of citationsSource
Perplexity~46.7% of top-source citationsThe Stacc
Google AI Overviews~21%The Stacc
ChatGPT11% to 27% depending on studyDiscovered Labs / Evertune
Gemini~3%, favors Medium among social sourcesTinuiti, Q1 2026
ClaudeNear-absent in some datasetsMachine Relations Index v2

What this means practically: if your buyers research in Perplexity or Google AI Overviews, Reddit is close to mandatory. If they use Gemini or Claude, Reddit is a weak channel and your budget belongs elsewhere. "Do Reddit" is not a strategy. "Do Reddit because our buyers use Perplexity" is.

Update, August 2026

Reddit's ChatGPT citation share collapsed roughly 86% in a single day on August 14, 2026, while Perplexity citations rose over the same window. See the follow-up: Is a Reddit Strategy Still Worth It for AI Search?


Why is YouTube the most underused GEO asset most companies already own?

Because spoken words become indexed text. A 12-minute walkthrough is a 2,000-word document that happens to have a video attached, and by several 2026 datasets YouTube has overtaken Reddit as the most-cited social platform overall.

  • LLM Pulse's rolling data (May 2026) ranked YouTube first at 26.47% of citations, with Reddit second at 17.39%.
  • OtterlyAI's study of 100M+ AI citations found YouTube represented 31.8% of all social and video platform citations, cited roughly 200x more than any other video platform.
  • Ahrefs' December 2025 study of 75,000 brands found YouTube mentions carried a 0.737 correlation with brand appearances in ChatGPT, AI Mode, and AI Overviews, the strongest single correlation in the dataset.

Three mechanisms are doing the work:

  • Transcripts as indexed text. Spoken words are transcribed, indexed, and retrieved as ordinary text.
  • Multimodal grounding. Gemini analyzes video frames directly. Demonstrating a physical product or a step-by-step interface workflow creates proof that page copy cannot replicate.
  • Ecosystem privilege. YouTube is Google-owned, and in Gemini and AI Overviews it reportedly accounts for a substantial share of all citations to Google properties. That is a structural advantage no amount of on-site optimization buys you.

Do views, likes, and subscribers affect AI citations?

No. OtterlyAI's data found near-zero correlation between engagement metrics and citation frequency.

Every conventional YouTube guide tells you to chase watch time and engagement. For citation purposes that advice is beside the point. A 300-view video with a clean, specific, well-spoken transcript answering a real question outperforms a 300,000-view video of vibes and b-roll.

That should be liberating for B2B and small-cap companies. You do not need a channel. You need answers on record.


Should I prioritize Reddit or YouTube?

Prioritize by the engine your buyers use and the question type you want to win. Reddit wins evaluation questions. YouTube wins procedural ones.

RedditYouTube
Primary GEO functionSocial proof, sentiment verification, third-party validationProcedural depth, visual proof, transcript indexing
Query types it wins"Best tools for X," "is Y worth it," "alternatives to Z," comparisons"How to do X," step-by-step, setup, troubleshooting, demos
Strongest enginesPerplexity, Google AI OverviewsGemini, AI Overviews, AI Mode
Weakest enginesGemini, ClaudeText-only ChatGPT sessions
Indexing speedNear real-time via licensed feedsFast, automated transcription
What optimizes itGenuine expertise, disclosed affiliation, specificity, trade-offsClear speech, clean uploaded transcripts, chapters, specific titles
Failure modeDetected as promotional, removed, and negative sentiment cachedVague narration that transcribes into noise

What else belongs in the citation map?

Wikipedia, LinkedIn, Medium, Forbes, and industry review platforms. Together with Reddit and YouTube, the top 15 domains capture roughly two-thirds of all AI citations.

For B2B software that means G2 and Capterra. For finance and mining it means trade press, exchange filings, and analyst coverage. Third-party "best X" listicles are consistently among the highest-leverage placements in the entire discipline. An inclusion in someone else's roundup often outperforms a year of your own blog.


What still works on my own website?

Data density and source provenance, not volume. The peer-reviewed benchmark is the GEO study from Princeton, Georgia Tech, the Allen Institute for AI, and IIT Delhi (KDD 2024), which tested content modifications across 10,000 queries.

ModificationEffect on visibility
Citing external sourcesUp to +40%, and far more for lower-ranked content
Adding statistics+30% to +41%
Adding credible quotations+22% to +28%
Keyword stuffing~3%, effectively nothing
Simply adding more wordsNo improvement

The practical instruction is short: write the number, name the source, date the claim.


Can AI crawlers actually read my site?

Check your robots.txt today. Many sites block AI crawlers by accident through a default CDN rule, then wonder why they are invisible.

Confirm these agents are handled deliberately rather than by default:

  • Google-Extended, which gates Gemini and AI Overviews grounding
  • GPTBot and OAI-SearchBot for OpenAI
  • PerplexityBot
  • ClaudeBot
  • User-triggered fetchers such as ChatGPT-User and Perplexity-User, which fire when someone pastes your URL into a chat

This is fifteen minutes of work and the highest-return item on any GEO list if something is misconfigured.


How do I measure whether any of this is working?

Track each engine separately and run a fixed prompt set monthly. A blended "AI visibility score" hides exactly the divergence that matters.

  • Separate referral traffic from chatgpt.com, perplexity.ai, and gemini.google.com in analytics.
  • Watch AI crawler user-agents in server logs.
  • Run 25 to 40 buyer-intent prompts across all five engines monthly and log who gets cited.
  • Use the generative-AI performance report in Google Search Console for Google's own surfaces.

Should I seed Reddit threads to get mentioned?

No. It is explicitly discouraged, operationally fragile, and the academic evidence says manipulation tactics mostly do not work.

  • It is on the record. Google's May 2026 guidance tells site owners not to pursue inauthentic mentions across the web.
  • The downside is permanent. Subreddit moderators identify astroturfing quickly. The risk is not a removed post. It is a permanently indexed thread in which real users describe your brand as the one that got caught spamming. That sentiment gets retrieved too.
  • The research is unkind to it. C-SEO Bench (Puerto et al., 2025), the first systematic benchmark of conversational-SEO tactics, found most of them do not help and several actively hurt, while plain source relevance keeps working.

The version that works is unglamorous: show up as a named expert, disclose who you work for, answer questions where you are genuinely the best-informed person in the thread, and be willing to say when a competitor is the better fit. Models are trained on the difference between advice and advertising. So are the humans voting.


What should I do in the first 90 days?

  1. Map your engines before your platforms. A mining-investor audience and a SaaS-buyer audience have almost inverted engine mixes, and therefore inverted Reddit weightings.
  2. Audit your crawler access. Fifteen minutes, highest return on this list if something is broken.
  3. Baseline your citations. 25 to 40 buyer-intent prompts, five engines, logged. You cannot improve a number you have never measured.
  4. Publish 8 to 12 answer-first videos. Two to six minutes each, one specific question per video, keywords spoken aloud, chapters marked, clean transcript uploaded manually rather than left to auto-captioning.
  5. Rewrite your top 10 pages for extractability. Question as the heading, direct answer in the first two sentences, then the statistic, the source, and the date. Apply this to what already ranks, not to new thin pages.
  6. Earn third-party inclusion. Analyst roundups, review platforms, trade-press listicles, podcasts. Someone else's page describing you is worth more than your own page describing you.
  7. Build a real Reddit presence. One account, disclosed, expert, patient. Measured in quarters, not weeks.
  8. Re-baseline at day 90. Citation share is volatile on a scale of weeks. Treat this as a monitored position, not a completed project.

Frequently asked questions

Is FAQ schema markup dead?

The Google rich result it powered is gone as of May 7, 2026. The markup itself is still valid Schema.org, still parsed by other crawlers, and harmless to leave in place. It simply never influenced AI visibility.

Do I need an llms.txt file?

No. Google stated on May 15, 2026 that no AI-specific file, chunking format, or structured data is required to appear in its generative features.

Is GEO different from SEO?

For Google's surfaces, largely no. The same crawlability, quality systems, and fundamentals apply. The difference is that roughly half the problem now lives off your domain, on Reddit, YouTube, Wikipedia, LinkedIn, and review sites, which conventional SEO never had to address.

Which single action has the highest return?

Auditing your robots.txt for AI crawler access. It takes fifteen minutes and, if it is misconfigured, nothing else you do will matter.

How long before I see results?

Citation share moves on a scale of weeks and is volatile. Baseline at day zero, re-measure at day 90, and expect the number to fluctuate for reasons outside your control, including platform licensing disputes.

Does this replace my SEO work?

No. It extends it. Crawlability, page quality, and technical fundamentals still govern whether you are eligible to be retrieved at all. The additional work is off-domain presence and answer-first page structure.

Google's framing is that this is all still SEO. That is true of Google's surfaces, and a useful antidote to the acronym industry. It is also incomplete, because it describes only the half of the problem that lives on your domain. You cannot mark up your way into the conversation. You have to actually be in it.

We run this playbook for our clients.

Content, video, third-party placement, and investor outreach, built and operated as one system.

Book a Demo