Skip to main content

Technical GEO

What are llms.txt and structured data, and do they help with AI answers?

Last reviewed: · By Kraffic.ai

Short answer

llms.txt is a proposed convention: a plain text file at the root of a website that lists its most important pages for AI systems. It is not an adopted standard, and AI assistants are not known to obey it. Structured data, meaning schema.org markup, is well established and helps machines interpret what a page describes. Neither guarantees a mention in AI answers.

Key takeaways

  • llms.txt is a proposal, not a standard; adding one is low cost but should not be expected to change AI answers by itself.
  • robots.txt is the established way to allow or block crawlers, including the AI crawlers that companies document.
  • Structured data labels facts on a page so machines do not have to infer them.
  • Markup must match what is visible on the page; marking up things that are not there creates risk and no benefit.
  • Clear visible content comes first, because markup describes content and cannot replace it.

Prompts this guide answers

  • “what is llms.txt”
  • “does llms.txt work”
  • “schema markup for AI search”
  • “structured data for ChatGPT”
  • “Should I add an llms.txt file to my website, and will ChatGPT, Claude or Gemini actually read it?”
  • “Which schema types should I add to my site so AI assistants understand my products and services correctly?”
  • “llms.txt vs robots.txt”
  • “technical GEO checklist”

What is llms.txt?

llms.txt is a proposed file, placed at the root of a website, that gives AI systems a short guide to the site in Markdown: what the site is and links to its key pages with brief descriptions. It was proposed in 2024 as a community convention.

The idea is to help a language model find the most useful pages without working through navigation, scripts and layout. Some sites also publish plain-text versions of key pages alongside it.

Do AI assistants actually use llms.txt?

AI assistants are not known to rely on llms.txt when generating answers. It is not a formal standard, and the major AI companies have not committed to reading or obeying it in their assistants.

That makes it an optional extra. Adding the file is quick and harmless if it is accurate, but it should not be sold or bought as a fix for AI visibility. If a provider claims llms.txt will get you named in answers, ask for the evidence.

How is llms.txt different from robots.txt and a sitemap?

The three files do different jobs. robots.txt is a long-established way to tell crawlers what they may fetch. A sitemap lists URLs for search engines to discover. llms.txt is a proposed summary for language models.

robots.txt, XML sitemap and llms.txt compared
FilePurposeStatusWhat to expect
robots.txtAllow or block crawlers from parts of a siteEstablished and widely followed by major crawlersDocumented AI crawlers can be allowed or blocked by user agent
XML sitemapList URLs for discoveryEstablished for search enginesHelps search engines find pages; AI use is not documented in detail
llms.txtSummarize a site and point to key pages for language modelsProposed convention, not a standardNo assurance that any assistant reads it

What is structured data and why does it help?

Structured data is code added to a page, usually as JSON-LD using the schema.org vocabulary, that states facts in a form machines can read without interpretation. It can say that a page is about a product, what the product is called, its price and its rating.

It helps because it removes guesswork. Search engines have used it for years to understand pages. For AI answers the benefit is indirect and not guaranteed: clearer facts in the systems that feed assistants reduce the chance of your business being described wrongly.

Which schema types are worth adding?

Add the types that match what your pages actually contain. The ones below cover most business websites.

  • Organization: your name, logo, website, contact details and official profiles.
  • LocalBusiness: address, hours and service area, for businesses serving specific places.
  • Product: name, description, brand, price, availability and ratings for items you sell.
  • Service: what a service is and who provides it.
  • FAQPage: questions and answers that are visible on the page.
  • Article: headline, author, publication and modification dates for guides and posts.
  • BreadcrumbList: where a page sits in your site.
  • SoftwareApplication: for software products, including category and operating system.

What are the rules for using structured data properly?

The main rule is that markup must describe what a visitor can see. Marking up reviews, prices or answers that are not on the page is misleading and can lead search engines to ignore your markup.

  1. Write the visible content first

    State the facts in plain text on the page.

  2. Add matching markup

    Use JSON-LD with the schema.org type that fits the page.

  3. Validate

    Test the markup with a schema validator and fix errors.

  4. Keep it current

    Update markup when prices, hours or products change.

What should you not expect from technical GEO?

Do not expect a file or a tag to produce mentions. Technical work makes your site easier to access and interpret; it does not make an assistant choose your brand.

Mentions depend on the whole picture: what your content says, what other websites say about you, and how the assistant's model weighs it. Treat technical GEO as removing obstacles, then measure what the assistants say to see whether anything changed.

FAQ

Questions and answers

Is llms.txt an official standard?

No. llms.txt is a community proposal from 2024, not a standard issued by a standards body, and AI assistants are not known to obey it. Some websites and tools have adopted it voluntarily. It is reasonable to add one, but there is no assurance that ChatGPT, Claude or Gemini will read it or change their answers because of it.

Should I add llms.txt to my website?

It is optional. The file takes little effort to create and does no harm if it is accurate and kept current. Do it after the work that matters more: making sure crawlers are not blocked, that key pages state your facts plainly, and that structured data is valid. Do not expect it to change AI answers by itself.

Can llms.txt block AI from using my content?

No. llms.txt is meant to guide AI systems to content, not to restrict them. To allow or block crawlers, use robots.txt with the user agent names each AI company documents. Check each company's current documentation for what its crawlers do and which controls it honors.

Does FAQ schema get my answers quoted by AI?

FAQPage markup labels question-and-answer pairs that are visible on your page, which makes them easy for machines to identify. That does not guarantee an assistant will quote them. What helps most is the content itself: a clear question, and an answer that is complete and makes sense without the rest of the page.

Do AI crawlers run JavaScript?

It differs by crawler and is not always documented, so the safe assumption is that some do not. If your main content appears only after scripts run, a crawler that does not run scripts may see an almost empty page. Serving the main text in the initial HTML avoids the problem for every crawler.

What is the most important technical step for GEO?

Make sure your pages can be reached and read. Check that robots.txt and bot protection are not blocking search and AI crawlers you want, that the main text is in the HTML, and that each key page states your core facts plainly. Structured data comes next. Files such as llms.txt are optional extras.

Get started

See what AI assistants say about your brand

Start with your website address. Kraffic runs the prompts your buyers type and shows you each saved answer. No one can guarantee what an AI assistant will say, so results are measured scan by scan.