Blog
How to Get ChatGPT to Recommend Your Business
There are three concrete technical changes that raise the odds of
ChatGPT, Claude or Perplexity citing your website: publishing an
llms.txt file, marking up content with structured data
(schema.org) and explicitly allowing AI crawlers in
robots.txt. None of the three guarantees a mention:
they're a necessary condition, not a sufficient one. Here's how each
one works.
How does a language model decide which website to cite?
A model like ChatGPT doesn't "browse" your website the way a person would: when it answers by citing a source, it does so from content it was able to crawl and understand clearly, or that its built-in search engine retrieves at the moment it answers. Two prior conditions decide whether your website even enters that pool: whether the model's crawler can access your pages, and whether the content is written so a specific answer can be extracted without having to interpret an entire page. Without those two things, it doesn't matter how good your website is: there's nothing a model can cite with confidence.
What is llms.txt and what's it for?
llms.txt is a text file at the root of a domain, built
specifically for language models: instead of a model having to crawl
and interpret an entire website to understand what a business does,
llms.txt summarises in plain text what the site is, what
services it offers and which pages are most relevant for common
questions. It's the equivalent, for a language model, of what a
sitemap is for a traditional search engine: it doesn't replace the
real content of the pages, but it makes finding and summarising it
much easier.
What do structured data (schema.org) add?
Schema.org is a shared vocabulary for marking up, within the HTML, what each thing is: that a price is a price, that a frequently asked question is a question with its answer, that an article has an author and a publication date. A language model —just like a search engine— can read that markup without having to guess the structure from the page's visual design. The practical difference is this: a price written as plain text in a paragraph is just one figure among many; a price marked up with the right schema.org type is a data point a model can extract reliably and attribute correctly to whoever published it.
Why does robots.txt need attention?
By default, some crawlers used by AI tools aren't allowed under the
standard configuration of many sites, or weren't even considered when
the file was written. If a site's robots.txt blocks, on
purpose or by omission, the crawlers used by OpenAI, Anthropic or
Perplexity, that model will never be able to index the content to
cite it, no matter how well it's written. Allowing those crawlers
isn't a promise of being cited: it's simply removing the obstacle
that stops the possibility from even being considered.
What makes a paragraph citable rather than just marketing copy?
These are the concrete differences between content a model treats as information and content it treats as marketing copy, which this very blog tries to apply:
| Element | Citable content | Marketing copy |
|---|---|---|
| Opening paragraph | Answers the question with a concrete fact from the first sentence | Opens with a brand line before saying anything useful |
| Market figures | Cites the exact source next to the figure | Uses figures with no origin, or made up |
| Limitations | Explains when what's being sold is NOT the right fit | Only lists advantages |
| Structure | Headings that are real questions, data tables | Vague headlines, continuous text with no data |
Does this guarantee ChatGPT will recommend my business?
No, and anyone who promises it as a guaranteed outcome isn't being
honest. Publishing llms.txt, marking up content with
schema.org and allowing AI crawlers are technical conditions
necessary for a model to be able to cite you, not a guarantee that it
will: the model still decides, for each specific question, which
sources it considers most reliable and relevant, and that also
depends on competition, the exact question asked, and changes to the
model itself that no one outside the companies that train it fully
controls. What can be said, because it's verifiable, is that without
those technical conditions the probability is flatly zero: you can't
cite what you can't crawl or understand.
What has actually been done on this website?
This very Zemog website applies the three changes this article
covers: it publishes its own llms.txt, marks up its
pages with schema.org (prices, FAQs and articles, each with its
matching type) and its robots.txt explicitly allows
OpenAI's, Anthropic's and Perplexity's crawlers. On top of that, pages
like this one are written following the same table from the previous
section: the first sentence answers the title's question with a
figure, every market figure cites its source next to the data, and
every article includes an explicit section on when the option we
sell isn't the right one. This isn't a promise that it will result in
mentions; it's a description of concrete technical work that anyone
can verify by looking at this page's source code.
Can I apply this to my own website?
Yes, and it doesn't require rebuilding the whole site:
llms.txt and robots.txt are text files that
can be added to any existing site, and structured data can be added
incrementally, page by page, starting with the ones you most want a
model to understand well (prices, FAQs, articles). Zemog's
SEO and AI search visibility service
includes exactly this technical preparation, together with content
written to answer what people actually ask, with a minimum
three-month commitment because there isn't enough data before that to
know if something is working.
If you'd like me to review where your website stands on these three points, tell me in the contact form: I reply within 24 hours.