A growing share of searches no longer ends in a list of links but in a finished answer. Somebody asks ChatGPT what to look for in a good kitchen planner and gets a paragraph with three named sources. Whoever is among those sources gets visitors. Whoever is not does not exist for that search.
This is a new kind of visibility, and it follows different rules from the classic results page. Not entirely different, but different enough that it pays to understand them while most competitors still do not. This piece answers both halves of the question: how these systems pick their sources, and what to change in your copy and your markup so that your page shows up as a source in ChatGPT.
What "ChatGPT as a source" actually means
When ChatGPT answers a question using web search, the answer carries a list of links, either inline or at the end: the sources it leaned on. Appearing in the ChatGPT sources simply means your URL is one of those links. It is also the only way an AI answer turns into a visit, because without a linked source the person asking gets the text and no reason to click.
What matters is the distinction between the three ways content enters an answer, because only one of them is yours to influence. First, the model's training knowledge: it may hold something about your industry, but it is not linked and you cannot write yourself into it after the fact. Second, a live web search the system runs in the background while it answers — this is where the linked sources come from, and where all the work happens. Third, a URL the user pastes in themselves; that is luck, not a strategy.
When people say "ChatGPT as a source", they almost always mean the second case. And it works closer to classic search than most expect: there is a query, an index, and a selection from results. Only the last step is new.
How these systems pick their sources
ChatGPT with web search, Perplexity, and Google's AI overviews work in much the same way at their core: they run a search in the background, read the best results, and compose an answer that cites some of them. Between the question and the source list sit four stages, and something drops out at each one. Knowing which stage your page fails at also tells you what to work on.
First, one question becomes several search queries. The system does not search with the sentence somebody typed; it turns it into two to five shorter, differently phrased queries — "what should I look for in a kitchen planner" becomes searches about the consultation process, measuring up, and cost. Optimizing a page for one exact wording therefore buys you little. What is wanted is a topic covered across its breadth.
Second, the results come out of a search index. That index is not magic, it is filled by crawlers: OpenAI maintains it with its own OAI-SearchBot and supplements it with bought-in search results. For you that means your page has to rank high enough in an ordinary index for the queries from stage one, because only the first results get read. A page sitting at position 90 for its topic in classic search rarely becomes the source of an AI answer. The groundwork in finding SEO problems is not a separate job, then, but the precondition.
Third, those results are fetched and cut into passages. This is where the technology bites: these bots generally do not execute JavaScript and do not wait for anything to load later. Whatever only comes into being in a browser — text inside accordions, sections pulled in by script, content behind a cookie banner or a login — simply does not exist for them. What remains is cut into sections along the headings.
Fourth, passages compete, not websites. From all the sections of all the pages it read, the model picks the ones that answer its question most directly, writes the answer out of them, and attaches to every borrowed statement the source it came from. That is why the strongest website does not win at this point; the clearest paragraph does. A paragraph that answers the question in full, without needing the rest of the page for context, gets cited far more often than the same content spread across four sections full of "as mentioned above". A small supplier with a precise answer regularly ends up standing beside an industry giant whose page only circles the same question.
What a page needs to survive all four stages
That boils down to a fairly short list of requirements:
- Findable in classic search, for the topic rather than for one phrasing. What sits on page ten never gets read.
- Fetchable without a browser: the text you want cited is in the delivered HTML — not behind JavaScript, a login, or a consent dialog.
- A section that answers one question completely, under a heading that names that question.
- Checkable specifics in the copy: a number, a unit, a date, a proper name, instead of "depends on the job".
- Visible currency: a date on the page saying when the details were last true.
- Mentions elsewhere, so the source does not rest on its own claim alone.
The first two points are craft carried over from classic search; the rest is decided in the copy. The next sections walk through exactly that rest.
What to change about your writing
The first lever is the self-contained answer. For every important question you want to be found for, picture a paragraph that carries the answer within it. Not "that depends on many factors, more on that shortly", but the answer, directly, with the elaboration after it. Journalism calls this the inverted pyramid: the most important thing first. AI answer systems love this shape, because that is exactly how they read.
The second lever is structure. Clear subheadings that name a real question or a real topic help the system find the right section. A heading like "What does it cost" is more useful than "Pricing", because it resembles the way people ask. That structure pays twice, because it also helps real visitors skim, as described in reducing bounce rate.
The third lever sits outside your page: mentions elsewhere. These systems also weigh whether a source appears elsewhere on the web. A specialist article that a directory links to and a forum quotes counts as more trustworthy than a page nobody else mentions. That is more effort than a copy change, but it is the same trust mechanism that counts in classic search too.
How content actually gets into an AI answer
The three levers above describe the attitude. Here are the three concrete moves that turn a well-written page into a citable one.
Clear facts somebody can copy down
An answer system prefers statements it does not have to rephrase and can stand behind. Those are sentences with a number, a unit and a reference: "The call-out fee is 45 euros within 20 kilometres" instead of "travel costs depend on effort". A date instead of "for many years". A proper name instead of "the market leader". Every one of those specifics turns your sentence into a checkable claim, and checkable claims get cited, because the model can attach a source to them.
Part of this is not hiding the fact. If the number lives in an image caption, inside a graphic, or behind an accordion that only loads on click, it does not exist for many systems. Whatever you want cited belongs in the visible page text.
Structured data that confirms those facts
Structured data is the same facts a second time, in machine-readable form: a small JSON block in the source that states what the page actually contains. Who wrote it, when it was last changed, what the product costs, which question is answered in which section.
<script type="application/ld+json">
{
"@context": "https://schema.org",
"@type": "FAQPage",
"mainEntity": [{
"@type": "Question",
"name": "How do I appear in the ChatGPT sources?",
"acceptedAnswer": {
"@type": "Answer",
"text": "By being findable in classic search for the question and having a section that answers it completely on its own."
}
}]
}
</script>
The markup is not an incantation that lifts a weak page. It removes ambiguity: "there is a year somewhere on this page" becomes "the publication date is 14 March 2026". Which types are worth it for a small site, and how to add and test them, is in structured data for AI answers.
llms.txt as a map
The third offering is a file at your-domain.com/llms.txt: a plain Markdown list of your best pages with one sentence of description each. It replaces nothing, but it is written in half an hour and needs no maintenance afterwards. Honestly placed, llms.txt is a young proposal and not a standard everyone is guaranteed to read — its shape, its use and its limits are in creating llms.txt.
The contradiction with pure keyword optimization
Here the new visibility parts ways with an old habit. For the classic results page, people often built a page around a keyword and worked it in as often as possible. For AI answers that helps little to nothing. These systems understand meaning, not word density. A page that repeats "rent a plate compactor" fifteen times does not get cited more. A page that cleanly answers the question behind the term does.
That is good news for anyone who would rather write clearly than write optimized anyway. The text that helps a human is the same one an answer system is happy to quote. The two goals that sometimes fought each other in old search pull in the same direction here.
How to check whether you appear in the ChatGPT sources
The test costs two minutes and is uncomfortably honest. Ask one of these systems the question you want to be found for, phrased the way a customer would phrase it, and make sure web search is actually active — without it the model answers from memory and names no sources at all. Then look at who is in the source list. If the same three competitors keep showing up, you know which pages you are measured against, and you can look at how their answer to the question is built.
Repeat the test every few weeks with the same questions. Unlike classic search, there are barely any tools yet that track this visibility for you. Your own spot check is, for now, the most honest instrument you have.
Are you even letting the AI bots in?
A point that matters mostly to the technically minded. These systems read your page through their own bots, and OpenAI splits the jobs across several: GPTBot collects material for training, OAI-SearchBot maintains the index behind web search, and ChatGPT-User fetches a page when a conversation asks for it right now. Perplexity uses PerplexityBot accordingly. Any of them can be blocked through robots.txt, the same way you could block Google.
The distinction has practical consequences: if you do not want to supply training data but do want to show up in answers, block GPTBot and let OAI-SearchBot and ChatGPT-User through. Blocking everything is a decision against being named — a legitimate stance, but it has a price. For a small site only just building visibility, the maths is usually clear: the citation and its visitors are worth more than the content you would be protecting. Check once whether your robots.txt accidentally carries such a block, because some builders add it unasked.
Where this fits in the larger sum
AI visibility is one of several routes to more visitors, and on its own it delivers small volumes. Its quality is unusually high in return, because these people have half-decided before they arrive. Where this lever sits alongside the others is in increasing website traffic.
Where to start
Take the one question you would most like to be found for. Find the paragraph on your page that answers it. Does the answer sit there whole and readable, or is it scattered across half the page? Rebuild it into a paragraph that stands on its own, and put a heading above it that names the question. That is the one change that makes the biggest difference for both humans and machines. Markup and llms.txt come after that, not before.
Frequently asked questions
How do I appear in the ChatGPT sources?
Your page has to meet two conditions. It has to be findable in web search for the question asked, because ChatGPT picks its sources from search results. And it has to contain a section that answers exactly that question on its own, under a heading that names the question. Miss the first and you are never read; miss the second and you are read but not cited.
Where does ChatGPT get its sources from?
From a web search that runs in the background while the answer is being written. The system turns your question into one or more queries, reads the best results, and cites some of them as sources. Answers without web search come from the model's training knowledge and therefore carry no links.
How does ChatGPT choose its sources?
In four steps. The question becomes several search queries; those queries hit a search index OpenAI maintains with its own OAI-SearchBot plus bought-in results; the best hits are fetched and cut into sections along their headings; and from all those sections the model picks the ones that answer its question most directly, attaching the matching source to every statement it borrows. What gets chosen, in other words, is not the strongest website but the clearest paragraph.
Can I cite ChatGPT as a source in my own work?
For evidence, generally no. ChatGPT is not a reference work but a model that composes from training data and search results, and it can be wrong. The citable thing is the page it points to — follow the link in the source list and verify the claim there. Citing ChatGPT itself only makes sense when the conversation is the subject, for instance in a piece about AI tools.
Does structured data help you get cited in AI answers?
It helps indirectly and reliably, but it is not a switch. Markup makes facts that are already in the text unambiguous: price, date, author, question and answer. It does not substitute for missing content and will not lift a page that contributes nothing. The effort pays off once the copy is already good.
How long does it take before my page shows up as a source?
That depends on when the page enters the search index the system draws on — for new pages, typically days to weeks. The only way to speed it up is the usual one: submit a sitemap, link internally, and check indexing in Google Search Console.