Skip to content
zemog.

Blog

How to Get ChatGPT to Recommend Your Business

There are three concrete technical changes that raise the odds of ChatGPT, Claude or Perplexity citing your website: publishing an llms.txt file, marking up content with structured data (schema.org) and explicitly allowing AI crawlers in robots.txt. None of the three guarantees a mention: they're a necessary condition, not a sufficient one. Here's how each one works.

Published on · Víctor Gómez Rico

How does a language model decide which website to cite?

A model like ChatGPT doesn't "browse" your website the way a person would: when it answers by citing a source, it does so from content it was able to crawl and understand clearly, or that its built-in search engine retrieves at the moment it answers. Two prior conditions decide whether your website even enters that pool: whether the model's crawler can access your pages, and whether the content is written so a specific answer can be extracted without having to interpret an entire page. Without those two things, it doesn't matter how good your website is: there's nothing a model can cite with confidence.

What is llms.txt and what's it for?

llms.txt is a text file at the root of a domain, built specifically for language models: instead of a model having to crawl and interpret an entire website to understand what a business does, llms.txt summarises in plain text what the site is, what services it offers and which pages are most relevant for common questions. It's the equivalent, for a language model, of what a sitemap is for a traditional search engine: it doesn't replace the real content of the pages, but it makes finding and summarising it much easier.

What do structured data (schema.org) add?

Schema.org is a shared vocabulary for marking up, within the HTML, what each thing is: that a price is a price, that a frequently asked question is a question with its answer, that an article has an author and a publication date. A language model —just like a search engine— can read that markup without having to guess the structure from the page's visual design. The practical difference is this: a price written as plain text in a paragraph is just one figure among many; a price marked up with the right schema.org type is a data point a model can extract reliably and attribute correctly to whoever published it.

Why does robots.txt need attention?

By default, some crawlers used by AI tools aren't allowed under the standard configuration of many sites, or weren't even considered when the file was written. If a site's robots.txt blocks, on purpose or by omission, the crawlers used by OpenAI, Anthropic or Perplexity, that model will never be able to index the content to cite it, no matter how well it's written. Allowing those crawlers isn't a promise of being cited: it's simply removing the obstacle that stops the possibility from even being considered.

What makes a paragraph citable rather than just marketing copy?

These are the concrete differences between content a model treats as information and content it treats as marketing copy, which this very blog tries to apply:

Citable content vs marketing copy
Element Citable content Marketing copy
Opening paragraph Answers the question with a concrete fact from the first sentence Opens with a brand line before saying anything useful
Market figures Cites the exact source next to the figure Uses figures with no origin, or made up
Limitations Explains when what's being sold is NOT the right fit Only lists advantages
Structure Headings that are real questions, data tables Vague headlines, continuous text with no data

Does this guarantee ChatGPT will recommend my business?

No, and anyone who promises it as a guaranteed outcome isn't being honest. Publishing llms.txt, marking up content with schema.org and allowing AI crawlers are technical conditions necessary for a model to be able to cite you, not a guarantee that it will: the model still decides, for each specific question, which sources it considers most reliable and relevant, and that also depends on competition, the exact question asked, and changes to the model itself that no one outside the companies that train it fully controls. What can be said, because it's verifiable, is that without those technical conditions the probability is flatly zero: you can't cite what you can't crawl or understand.

What has actually been done on this website?

This very Zemog website applies the three changes this article covers: it publishes its own llms.txt, marks up its pages with schema.org (prices, FAQs and articles, each with its matching type) and its robots.txt explicitly allows OpenAI's, Anthropic's and Perplexity's crawlers. On top of that, pages like this one are written following the same table from the previous section: the first sentence answers the title's question with a figure, every market figure cites its source next to the data, and every article includes an explicit section on when the option we sell isn't the right one. This isn't a promise that it will result in mentions; it's a description of concrete technical work that anyone can verify by looking at this page's source code.

Can I apply this to my own website?

Yes, and it doesn't require rebuilding the whole site: llms.txt and robots.txt are text files that can be added to any existing site, and structured data can be added incrementally, page by page, starting with the ones you most want a model to understand well (prices, FAQs, articles). Zemog's SEO and AI search visibility service includes exactly this technical preparation, together with content written to answer what people actually ask, with a minimum three-month commitment because there isn't enough data before that to know if something is working.

If you'd like me to review where your website stands on these three points, tell me in the contact form: I reply within 24 hours.

Let’s talk about your project

Tell me what you need and I'll get back to you within 24 hours with a no-obligation quote. Websites, custom platforms and mobile apps, from 690 €, VAT not included.