How to get your pages cited by AI assistants

More people now ask an assistant before they search. When ChatGPT, Perplexity or Google's AI Overviews answer a question, they often link to the pages they drew on. Being one of those links is valuable, and there is no secret formula for it. What there is, is a set of basics that make a page easy to find, easy to read and safe to quote.
Start with being findable
An assistant can only cite a page its systems can reach. For Google, the rules are the same as for ordinary search: Google's documentation on AI features says a page must be “indexed and eligible to be shown in Google Search with a snippet” to appear as a supporting link in AI Overviews or AI Mode, and that “there are no additional requirements” or special optimisations.
Other assistants run their own search crawlers, so they need to be allowed in too. Pages that are rendered on the server, so the text is in the HTML without running JavaScript, are the safest choice for every crawler.
Search crawlers and training crawlers are different
The major AI companies publish separate user agents for search and for model training, and each can be allowed or blocked in robots.txt on its own:
- OpenAI: OAI-SearchBot is “used to surface websites in search results in ChatGPT's search features”; GPTBot crawls content “that may be used in training our generative AI foundation models”. OpenAI states each setting is independent, so a site can allow OAI-SearchBot while disallowing GPTBot, and sites that opt out of OAI-SearchBot are not shown in ChatGPT search answers.
- Anthropic: ClaudeBot collects content that could contribute to model training; Claude-SearchBot crawls to improve search results, and Anthropic notes that blocking it may reduce a site's visibility in Claude's search results. Claude-User fetches pages when a person asks Claude to.
- Perplexity: PerplexityBot surfaces and links websites in Perplexity's results and, according to Perplexity, is not used to crawl content for AI foundation models. Perplexity-User fetches pages in response to a user's question.
- Google: Google-Extended controls whether content Google crawls may be used to train Gemini models and for grounding in Gemini Apps and Vertex AI. Google states it does not affect inclusion or ranking in Google Search, and AI Overviews rely on ordinary Googlebot access.
That gives a business a real choice. If you want to be cited but would rather your content were not used for training, allow the search agents and disallow the training ones. Blocking everything with an AI-sounding name usually removes you from the answers as well.
Write pages that are easy to quote
- Lead with the answer. A clear first sentence that answers the question (“A full blood count at our Westlands branch costs KES 1,500 and results are ready the same day”) is easier to quote than a paragraph of preamble.
- Put facts in real structure: headings, lists and HTML tables, not images of text.
- Say who wrote it and who checked it. Google's guidance on helpful content asks whether authorship is self-evident, and warns against made-up names or false credentials. Use real people.
- Cite your sources with links, especially for figures and rules that come from a regulator or official body.
- Keep it current. Show when the page was last updated, and update it when prices or rules change rather than leaving old figures in place.
Structured data helps, but is not a ticket in
Structured data (schema.org markup in JSON-LD) describes the page in machine-readable form: the organisation, the author, the questions and answers, the breadcrumb trail. It helps search engines understand a page and can make it eligible for rich results. But Google is explicit that “there's also no special schema.org structured data that you need to add” to appear in its AI features. Treat it as good housekeeping, and make sure it matches what is visible on the page.
What about llms.txt?
llms.txt is a proposal, made by Jeremy Howard in September 2024, for a Markdown file at the root of a site (/llms.txt) that gives language models a short summary and a curated list of links to the site's most useful content, ideally in plain Markdown. Many documentation sites and tools now publish one.
It is worth being honest about its status. It is a community convention, not a standard from a body such as the IETF or W3C, and Google's AI features documentation says you do not need to create “new machine readable files, AI text files, or markup” to appear in AI Overviews. Publishing an llms.txt is cheap and does no harm; just do not expect it to do the work that clear, crawlable, well-sourced pages do.
No one can guarantee a citation
Assistants choose sources by their own systems, which change often and are not published in detail. Anyone who promises that your business will be “recommended by ChatGPT” is promising something they do not control. What you can control is the page: reachable, clear, attributed, sourced and current.
Searchlode builds pages this way by default: an answer first, named authors and reviewers in the structured data, a sources block on every page, FAQ markup, a Markdown version of each page and an llms.txt index. Each client chooses whether AI training crawlers are allowed; Googlebot is never blocked.