Keyword Research for AI Search Engines: How to Show Up in Perplexity and ChatGPT

Last updated: 29 July 2026
Keyword research for Perplexity and ChatGPT differs from Google SEO because these AI systems prioritize cited sources that directly answer conversational queries. The five-step workflow involves mapping user questions to your content's answer quality, identifying gaps where AI assistants lack authoritative sources, and optimizing for citation likelihood rather than search ranking. This approach positions your content as a trusted reference that AI models actively recommend to users.
Word count: 72
Traditional keyword lists were built for a different system. You found high-volume, low-competition terms, optimized a page around them, and waited for rankings. AI answer engines work differently. Perplexity and ChatGPT do not rank ten blue links; they synthesize a single answer and cite two or three sources inline. A keyword list that tells you "monthly search volume: 4,400" gives you no signal about whether your article will be one of those cited sources. The gap between the two approaches is structural, not cosmetic.
The numbers reflect how fast this shift is moving. ChatGPT now processes over one billion queries per week, and Perplexity has crossed 100 million monthly users, a milestone the company confirmed in early 2024. A meaningful share of those queries are research-oriented, the kind where the engine cites sources rather than just answering from its training data. Otterly's 2026 AI keyword research framework puts it plainly: winning in AI search means mapping prompts, entities, and sub-questions, not chasing exact-match terms.
The five-step workflow covers:
- Identifying the query types AI engines actually cite sources for
- Extracting the sub-questions and entities buried inside each query
- Scoring your existing content against citation-readiness signals
- Structuring new articles so the answer appears in the first paragraph
- Tracking which prompts cite you versus a competitor, then closing the gap
One honest caveat up front: this workflow improves your probability of being cited. It does not guarantee it. AI engines make probabilistic retrieval decisions, and a well-structured article still competes against established domains with more backlink authority. The workflow tilts the odds; it does not remove them.
Before You Start: What You Need
To research keywords for AI search engines, you need four things: access to Perplexity (free tier works) and ChatGPT, a spreadsheet to log prompt patterns and citation signals, and one existing content asset to use as a test case. You also need to be comfortable reading a Search Console report and know what a featured snippet is.
That baseline matters because AI keyword research borrows from both disciplines. You are not starting from zero, but you are applying familiar concepts in a different environment. Search Console shows you what Google indexed; it tells you nothing about what Perplexity cited last Tuesday. Those are separate data streams.
On the tool side, the free tier of Perplexity covers basic query testing, but Perplexity's Deep Research mode adds multi-step reasoning and source synthesis that surfaces citation patterns you will not see in a standard search. If you plan to run more than a handful of test queries, the Pro tier is worth the cost.
The trade-off here is time. Manual prompt testing across two AI engines, logged in a spreadsheet, is slow. A complete AI keyword research workflow for 2026 typically involves semantic clustering and intent mapping on top of basic query logging, which means the spreadsheet approach scales to maybe 30 to 40 target prompts before it becomes unwieldy. If your content program runs larger than that, you will need a more systematic tracking layer eventually.
One content asset to test against is essential. Abstract keyword research for AI engines without a real page to pressure-test produces lists, not insight. Pick something you already rank for in Google. That gives you a comparison point: you know the topic has search demand, and you can observe whether AI engines cite you, cite a competitor, or cite nobody at all.
Step 1: Map the Query Fanout for Your Topic

When an AI search engine receives a query, it rarely answers from a single document. It decomposes the question into a cluster of sub-queries, retrieves sources for each, then synthesizes a response. That decomposition is the fanout. A seed topic like "project management software" can spawn 30 or more distinct sub-queries in a single session, each pulling different citations. Mapping that fanout before you write is how you identify which questions actually need coverage.
Why One Keyword Becomes Dozens of Sub-Queries
Traditional keyword research assumes a one-to-one relationship: one query, one intent, one page. AI engines break that assumption. Perplexity, for example, runs follow-up retrieval passes to fill gaps in its initial answer, a behavior Ethan Lazuk documents in his analysis of Perplexity content topics. Each pass is effectively a new query. A page that answers only the surface question gets cited once, if at all. A page that pre-empts the follow-up questions gets cited repeatedly across the thread.
How to Prompt for the Full Fanout Tree
Open Perplexity or ChatGPT and run this prompt against your seed topic:
"List every sub-question a researcher would need to answer to fully understand [topic]. Group them by intent: definitional, comparative, procedural, and evaluative."
Then run a second prompt:
"For each sub-question above, what follow-up questions would a reader still have after reading a standard answer?"
That second pass surfaces the second-order fanout, the questions your competitors are probably not covering. A practical walkthrough of Perplexity's research prompting patterns shows how iterative follow-up prompts consistently surface angles that a single broad query misses.
Run this across both engines. ChatGPT and Perplexity decompose differently. ChatGPT tends to generate more procedural sub-queries; Perplexity skews toward comparative and evaluative ones. Running both gives you a fuller tree.
Scoring Sub-Queries by Answer Frequency
Not every sub-query in the fanout is worth a dedicated page. To prioritize, run each candidate sub-query three separate times across two engines and log which URLs appear in the citations each time. Sub-queries that consistently pull the same two or three sources are already "owned" by those pages. Sub-queries that return inconsistent citations, or pull from weak sources, are the gaps worth targeting.
A rough scoring rule: if a sub-query returns the same cited domain in at least two of three runs, that domain has a citation hold on it. If the citations vary across all three runs, the topic is contestable. Prioritize contestable sub-queries first.
The trade-off is time. Running three separate sessions per sub-query across two engines is manual and slow, especially for a fanout tree that might contain 40 or 50 nodes. The approach breaks down when you are working with a broad category topic rather than a specific niche, because the fanout becomes too wide to score exhaustively by hand. In that case, narrow your seed topic before you start, or accept that you are sampling the fanout rather than mapping it completely.
Step 2: Classify Queries by AI Search Intent

Not every query on your fanout list has the same chance of earning a citation in Perplexity or ChatGPT. Generative engines pull from content differently depending on what the query is asking for. Classifying each query into one of four intent buckets before you write anything lets you prioritize the formats most likely to get cited, and skip the formats that almost never do.
The Four Buckets
Definitional. The query asks what something is. "What is retrieval-augmented generation?" or "What does churn rate mean?" These queries produce encyclopedia-style answers. AI engines often synthesize definitions from multiple sources rather than citing one page directly, which limits your citation odds.
Comparative. The query asks which option is better. "Perplexity vs. ChatGPT for research" or "PostgreSQL vs. MySQL for high-write workloads." Generative engines handle these by weighing trade-offs, and AI search intent research from 2026 shows that ChatGPT and Perplexity increasingly ask clarifying questions before narrowing down options, meaning your comparative content needs to address multiple user scenarios, not just one.
Procedural. The query asks how to do something. "How to set up a webhook in Zapier" or "How to calculate customer lifetime value." Step-by-step structure maps directly to how generative engines construct instructional answers.
Sourced-fact. The query asks for a specific, verifiable number or claim. "What is the average SaaS churn rate?" or "How many tokens does GPT-4 support?" These queries force AI engines to cite a source because the answer requires attribution to be credible.
How to Tag Each Query
Run every query from your fanout list through a single decision rule:
- Does it ask for a definition? Tag it definitional.
- Does it ask which option, or compare two or more things? Tag it comparative.
- Does it ask how to do something, step by step? Tag it procedural.
- Does it ask for a specific number, date, rate, or named fact? Tag it sourced-fact.
Queries that feel ambiguous usually fit two buckets. Pick the one that matches the dominant intent in the top results for that query, not the one that feels more comfortable to write.
Why Procedural and Sourced-Fact Win More Citations
Procedural and sourced-fact queries produce the most citable content because they require structure and specificity that AI engines cannot easily synthesize on their own. A step-by-step guide with numbered actions gives a generative engine a clean passage to extract and attribute. A page with a named statistic and a clear methodology gives it a reason to cite rather than paraphrase.
Definitional content, by contrast, gets absorbed and rewritten. Comparative content gets partially cited but often loses attribution when the engine builds its own pros-and-cons summary.
One trade-off worth acknowledging: procedural content ages faster. A "how to" guide for a tool that ships a major update in six months can become inaccurate quickly, and AI engines do factor recency into citation decisions. If your procedural content covers a fast-moving product, build a review cadence into your publishing calendar, not just a one-time publish.
Sourced-fact content has its own limitation. If the statistic you cite is widely available from a primary source (a government database, a major industry report), AI engines will often cite that primary source instead of your page. The stronger play is to produce original data or aggregate multiple sources into a single, clearly attributed figure that exists nowhere else.
Step 3: Audit Existing Content Against Your Query List

Before writing a single new page, check what you already have. Auditing existing content against your AI query list tells you whether your current pages appear in Perplexity or ChatGPT answers, and if not, exactly why. Most pages fail for one of three diagnosable reasons: vague claims with no numbers, missing named sources, or assertions that no LLM can verify. Fix those first. Net-new content comes after.
Run a Manual Citation Check in Perplexity
Take the top 20 queries from your list and paste each one directly into Perplexity. Scan every response for your domain in the cited sources panel. Log the result in a simple spreadsheet: cited, not cited, or competitor cited instead.
This takes about 45 minutes for 20 queries if you move efficiently. The output is a citation rate for your existing content, expressed as a percentage. If you are cited in fewer than 3 of 20 queries on topics you already rank for in Google, your content has a structural problem, not a distribution problem.
The Three Failure Modes
Vague claims. A sentence like "many companies use this approach" gives an AI engine nothing to attribute. Replace it with "According to a 2025 Gartner survey, 61% of enterprise teams adopted RAG-based retrieval by Q4." Specificity is what makes a passage citable.
Missing named sources. AI engines are more likely to cite pages that themselves cite credible sources. A page with no outbound citations reads as opinion. A page that links to primary research reads as synthesis, and synthesis gets cited.
Unverifiable assertions. If your page makes a claim that cannot be cross-referenced anywhere else on the web, a generative engine will skip it. This is especially common with proprietary frameworks or internal data presented without methodology. Add a methods note, even a brief one.
Prioritizing Fixes vs. New Content
Score each existing page on a simple three-point scale:
- 1 point for having at least one specific statistic with a named source
- 1 point for having a clear procedural or comparative structure
- 1 point for covering at least two of the second-order sub-questions from your fanout tree
Pages scoring 0 or 1 are candidates for a content refresh before you write anything new. Pages scoring 2 or 3 that still are not getting cited likely have a domain authority problem, not a content problem, and new pages will face the same ceiling.
Step 4: Structure New Content for AI Citation

Once you know which queries to target and which existing pages to fix, the next task is writing new content in a format that generative engines can actually use. The structure of a page matters as much as its substance. An AI engine retrieving a passage to cite needs to find a clean, self-contained answer near the top of the document, not buried in paragraph eight.
The Answer-First Format
Put the direct answer to the query in the first 40 to 60 words of the article. Not a teaser, not a hook, the actual answer. If the query is "how to calculate net revenue retention," the first paragraph should contain the formula and a one-sentence explanation of what each variable means.
This runs counter to traditional SEO writing, where the introduction builds context before delivering the payoff. AI engines do not read for narrative arc. They retrieve passages. A passage that opens with the answer is far more likely to be extracted and cited than one that opens with background.
Named Sections with Descriptive Headings
Use H2 and H3 headings that describe the content of the section, not clever labels. "How to Set Up Webhook Authentication in Zapier" is a citable heading. "Getting Started" is not. Generative engines use heading text as a signal for what the section covers, and that signal feeds into retrieval decisions.
Each section should be able to stand alone as an answer to a specific sub-query. If a section only makes sense in the context of the sections before it, it is too dependent on narrative flow and will not extract cleanly.
Specificity Markers
Seed every section with at least one of the following:
- A named statistic with a source and date
- A step-by-step numbered list with concrete actions
- A named comparison between two specific options
- A defined term with a precise, quotable definition
These are the elements that give a generative engine something to attribute. Prose without any of these markers tends to get paraphrased without citation.
One Limitation to Keep in Mind
Formatting for AI citation can make content feel dense or clinical if you overdo it. A page that is nothing but numbered lists and statistics loses the explanatory connective tissue that helps a human reader understand context. The goal is to embed citable elements inside readable prose, not to replace prose with structured data. Readers who arrive from a cited AI answer still need to find the page useful enough to stay.
Step 5: Track Citation Performance and Close the Gap
Publishing optimized content is not the end of the workflow. AI citation patterns shift as engines update their retrieval models, as competitors publish new content, and as your own domain authority changes. Tracking which prompts cite you, which cite a competitor, and which cite nobody gives you the feedback loop that makes the workflow repeatable.
Building a Citation Tracking Spreadsheet
Set up a spreadsheet with these columns:
- Query text
- Engine tested (Perplexity or ChatGPT)
- Date of test
- Your domain cited? (yes/no)
- Competitor domain cited (if applicable)
- Citation position (first, second, third)
- Notes on answer structure
Run the full list of target queries once per month. A monthly cadence is frequent enough to catch meaningful shifts without consuming more time than the signal justifies. Weekly tracking is only worth the effort if you are actively publishing new content or running a refresh campaign.
Reading the Gap
When a competitor is cited instead of you on a query you have content for, compare the two pages directly. Look for three things:
- Does their page answer the query in the first paragraph and yours does not?
- Do they have a named statistic or primary source you are missing?
- Is their page more recently updated?
In most cases, the gap is fixable with a targeted edit rather than a full rewrite. Add the missing statistic, move the answer higher, update the date. Then re-test the query in two to three weeks.
When the Gap Is Not Fixable by Content Alone
Sometimes a competitor holds a citation position because of domain authority, not content quality. A page on a domain with 50,000 referring domains will often outrank a better-written page on a domain with 500. In that case, the content fix is necessary but not sufficient. You also need to build the authority signals that make your domain a credible retrieval target, which is a longer-term project outside the scope of a single content audit.
Acknowledge that ceiling early. If you are consistently losing citation positions to the same two or three high-authority domains across multiple queries, the bottleneck is not your writing.
Frequently Asked Questions
Does Google SEO keyword research transfer directly to AI search?
Partially. Volume and competition metrics from Google Keyword Planner tell you that a topic has demand, but they do not tell you whether AI engines cite sources for that query or synthesize an answer from training data alone. You need to run the query in Perplexity or ChatGPT to see whether citations appear at all. If they do not, traditional SEO optimization is the right tool. If they do, the AI keyword research workflow in this article applies.
How often do AI engines update which sources they cite?
There is no published update schedule. Perplexity retrieves sources in real time, so citation results can shift day to day as new content is indexed. ChatGPT's browsing-enabled responses also vary by session. In practice, a page that earns a citation position tends to hold it for weeks to months unless a competitor publishes something significantly more specific or more recent. Monthly tracking is enough for most content programs.
Can you do this research without paying for any tools?
Yes, with limits. The free tier of Perplexity and the free tier of ChatGPT cover basic query testing. You will not have access to Deep Research mode or advanced citation analytics, but you can still run the fanout prompts, classify intent, and log citation results manually. The ceiling is roughly 30 to 40 queries before the manual process becomes too slow to be practical. Beyond that, a paid tier or a dedicated AI search tracking tool is worth evaluating.
What content format gets cited most often in Perplexity?
Based on observed citation patterns, pages with a direct answer in the opening paragraph, at least one named statistic with a source, and a clear procedural or comparative structure get cited more consistently than pages built around narrative or opinion. Perplexity also tends to favor pages that have been cited by other credible sources, so domain authority still plays a role even in AI-native search.
Is there a minimum word count that helps with AI citation?
No specific word count threshold has been confirmed by either Perplexity or OpenAI. What matters is whether the page contains a self-contained, citable passage near the top. A 600-word page with a precise answer in the first paragraph will often outperform a 3,000-word page where the answer is buried. Longer content helps when it covers more sub-questions from the fanout tree, not because length itself is a signal.
How do you know if an AI engine is citing you from live retrieval vs. training data?
In Perplexity, every cited source appears in the sources panel with a URL and retrieval timestamp, so you can confirm live retrieval directly. In ChatGPT with browsing enabled, cited URLs appear inline. If ChatGPT answers without any citations, it is drawing from training data, and no amount of content optimization will change that for that specific query. Focus your AI keyword research on queries where citations appear consistently, because those are the ones where your content can actually influence the outcome.
If you want help applying this workflow to your specific content program, visit Seorav to see how the team approaches AI search visibility for content-driven businesses.
Keep reading

Healthcare SEO: Ranking When Trust and Compliance Matter Most
Learn how SEO for the healthcare industry works across YMYL standards, HIPAA-compliant analytics, E-E-A-T requirements, and local search. Practical guidanc

Financial Services SEO: Ranking When Regulations and Trust Are Everything
Learn how financial services sites rank on Google by turning compliance requirements into E-E-A-T signals. Covers YMYL, schema, disclaimers, and technical

How Real Estate Agents Win With SEO
Learn how to do SEO for real estate: optimize your Google Business Profile, build neighborhood content, fix technical issues, and rank for local buyer sear