Table of Contents
llms.txt vs sitemap.xml vs RSS: What AI Crawlers Actually Fetch (October 2026)
Google's John Mueller says llms.txt does not work as a sitemap for Google, and that the files he has actually seen AI crawlers fetch in his own server logs are sitemap.xml and RSS. On the October 1, 2026 episode of Google's Search Off the Record podcast, he said Google's systems can't use llms.txt as a sitemap because it lacks the strict format. Then he added: "I think the hope is bigger than the reality" (per Search Engine Journal). On whether search systems read those Markdown files today, his answer was "Currently none of this happens."
For a site that wants AI crawlers to find its pages, that turns a vague GEO to-do into a short decision. Serve a sitemap at the default /sitemap.xml path. Link an RSS feed from the head of every page. Treat llms.txt as optional, and do not expect a Google Search effect from it. The comparison table below shows where each of those calls comes from, and the last section is a three-item checklist you can hand to whoever owns your site's build.
What Mueller said on October 1, and where his evidence stops
The episode is titled "Do sitemaps still matter?" and pairs Mueller with Martin Splitt on the Google Search Central YouTube channel. Search Engine Journal's Matt G. Southern wrote it up on October 5, and every quote below comes from that write-up.
Google and Bing give you a console to submit a sitemap. AI training crawlers "usually don't have any kind of a Console or any setup where you can submit a sitemap file," Mueller said. His advice for sites that want their content in AI systems: "either you stick to the generic naming, call it sitemap.xml, or you focus on RSS feeds" (per SEJ).
Mueller also said he has seen it in his own logs: "I've seen that happen in my server logs where some AI crawler accesses my sitemap file," he said, adding that he has seen the same with his RSS files (per SEJ).
Read that evidence at its actual size. It is one person's logs. He did not name the crawlers, and he said he doesn't know whether AI companies document this behavior or what they do with the files afterward. It shows that some AI crawler fetched those two files on his sites. It does not tell you which companies do it, how often, or whether a fetch leads to a citation.
His earlier data point is narrower still. In an August 2026 Reddit thread about serving Markdown to bots, Mueller wrote: "On my test sites the only crawlers who claim to accept markdown are SEO tools" (per SEJ, August 31, 2026). That was about Markdown page versions requested through the HTTP Accept header, a close cousin of llms.txt, and it covered his test sites only.
Three files compared: sitemap.xml, RSS, and llms.txt
Each row cites the spec or the statement it rests on. "Mueller's logs" means exactly that: his own sites, crawlers unnamed.
| sitemap.xml | RSS or Atom feed | llms.txt | |
|---|---|---|---|
| What it is | An XML list of a site's URLs for crawlers (sitemaps.org) | A feed of recent items, which the sitemap protocol also accepts as a limited sitemap (sitemaps.org) | A Markdown file with a site summary and curated links, proposed by Jeremy Howard on September 3, 2024 (llmstxt.org) |
| Seen fetched by an AI crawler in Mueller's logs | Yes (SEJ) | Yes (SEJ) | Not reported; on his test sites, only SEO tools claimed to accept Markdown (SEJ) |
| Format rules | Strict XML, capped at 50,000 URLs and 50MB per file (sitemaps.org) | RSS 2.0 or Atom, usually recent items only (sitemaps.org) | Loose Markdown; an H1 is the only required section (llmstxt.org) |
| How a crawler finds it | Guessing the default name, a robots.txt Sitemap: line, or console submission (SEJ) | A link rel="alternate" tag in each page's head (RSS Board) | Convention only: the site root or any path below it (llmstxt.org) |
| Search engine support | Submit it to Google directly (SEJ) | Accepted as a limited sitemap under the protocol (sitemaps.org) | Ignored by Google Search (Google Search Central) |
| Our call | Ship it at /sitemap.xml | Ship it and link it from every page | Optional; no Google Search effect |
Two of the three files have a published format, a discovery path that works without a console, and a direct sighting in Mueller's logs. The third has none of those, at least for Google.
sitemap.xml: the filename is the discovery mechanism
A crawler with no submission console has two ways to find your sitemap. It can read the Sitemap: line in robots.txt, which the protocol treats as separate from any user-agent block, so it applies to every bot that supports the directive (per sitemaps.org). Or it can request the obvious path and see what comes back. Mueller's "stick to the generic naming" advice is aimed at that second group.
He also described the reverse case. A site that wants a private sitemap can give it an unusual filename, leave it out of robots.txt, and submit it only to Google; the cost is that other systems can't find it, and Bing would probably need its own submission (per SEJ). Plenty of sites are in a milder version of that situation by accident, because their CMS picks the name.
WordPress core has served its sitemap index at /wp-sitemap.xml since version 5.5, and references it in robots.txt (per WordPress core). Astro's sitemap integration writes sitemap-index.xml and sitemap-0.xml (per Astro docs). Pondero runs on Astro, and when we checked our own site on October 6, 2026, /sitemap.xml returned a 404 while robots.txt pointed to /sitemap-index.xml. Any crawler that reads robots.txt finds it. A crawler that only tries the default name does not.
Fixing it takes a redirect and a robots.txt line. Keep whatever your CMS generates, and make /sitemap.xml resolve to it, either as a copy of the index or as a redirect. Then list it in robots.txt:
User-agent: *
Allow: /
Sitemap: https://example.com/sitemap.xml
On Cloudflare Pages, Netlify, or any host that reads a _redirects file, one line covers an Astro site:
/sitemap.xml /sitemap-index.xml 301
RSS: every page already advertises it
Mueller said feeds are easier to find because they are usually linked from a page's HTML head (per SEJ). That is the RSS autodiscovery convention, and the spec requires the tag to sit inside the head element with rel="alternate" and the feed's MIME type (per the RSS Advisory Board):
<link rel="alternate" type="application/rss+xml" title="Example articles" href="https://example.com/rss.xml">
Because the tag rides along on every HTML page, a crawler that lands anywhere on your site gets a pointer to the feed without guessing a filename or needing a submission console.
A feed has one structural weakness. The sitemap protocol notes that feeds usually carry only recent items, so they may not tell a crawler about every URL (per sitemaps.org). Run both: the feed signals what is new, and the sitemap covers the back catalog.
llms.txt: a file for browser agents, ignored by Google Search
The llms.txt spec describes a Markdown file at the site root (or any path) with an H1, a short blockquote summary, and sections of curated links, positioned as a curated overview for LLMs where sitemap.xml lists every page (per llmstxt.org). It was never designed as a sitemap replacement, which is the comparison Mueller rejected.
Google's own products disagree about it. The Search team's AI optimization guide, last updated July 10, 2026, says you don't need new machine-readable files, AI text files, or Markdown to appear in Google Search, including its generative AI features, and that Google Search ignores them (per Google Search Central). Meanwhile Lighthouse 13.3 added an experimental Agentic Browsing category with an llms.txt audit (per SEJ, May 20, 2026). That audit fails a page only when fetching llms.txt throws a server error, and marks a 404 as not applicable because "providing the file is optional at the moment" (per Chrome for Developers).
So the file has a defined job: helping browser-based agents get oriented on a site. It has no job in Google Search ranking or AI Overview citations. Whether OpenAI's, Anthropic's, or Perplexity's crawlers read it is a separate question, and nothing Mueller said answers it in either direction. He spoke about his own logs and test sites. His stated position on llms.txt is that he's fine with sites trying it but wouldn't rely on it (per SEJ).
If you already have one, leave it up. Pondero serves one at /llms.txt. Just make sure it returns a 200 or a clean 404, never a 5xx, since a server error is the one thing Lighthouse flags.
Check your own logs before anyone builds a generator
Mueller's August advice still applies: if you want to know whether AI bots request Markdown, set up logging for the Accept header and check the numbers before investing (per SEJ). The same logic covers llms.txt. Your access logs already say which user agents fetch which discovery files. On a server writing the standard combined log format, this counts requests per file and user agent:
grep -E '"GET /(sitemap[^ ]*\.xml|wp-sitemap\.xml|rss\.xml|feed/?|llms\.txt) ' /var/log/nginx/access.log \
| awk -F'"' '{split($2, req, " "); print req[2] "\t" $6}' \
| sort | uniq -c | sort -rn | head -40
Look for two things. First, whether any AI crawler user agent appears against /sitemap.xml or your feed, which would match what Mueller saw. Second, whether anything other than SEO tools touches /llms.txt. User-agent strings are self-reported, so treat a match as a lead to verify against the operator's published crawler documentation, and not as proof. If you would rather not parse raw logs, our comparison of AI search visibility tools covers platforms that track the citation side instead.
The verdict: three things to ship this week
Mueller's comments settle the Google side of this, and the specs settle the rest.
- Make
/sitemap.xmlresolve. If your CMS writes/wp-sitemap.xmlor/sitemap-index.xml, add a redirect or copy so the default name works, and keep theSitemap:line in robots.txt pointing at it. That is the filename Mueller recommended (per SEJ). - Link an RSS feed from every page's head. One
link rel="alternate"tag in your base layout covers the whole site, and it is the second file Mueller reported seeing AI crawlers fetch. - Do not fund an llms.txt project for Google citations. Google Search ignores the file (per Google Search Central). If one already exists, keep it error-free and move on.
By site type: a content or marketing site should ship items 1 and 2 and skip llms.txt work entirely. A documentation site whose readers send browser agents to it has a real reason to maintain llms.txt, because that is the use case Lighthouse audits. An agency that lists llms.txt as a GEO must-have should drop it from Google-facing deliverables, since Google's own Search documentation says the file does nothing there.
Discovery only gets a crawler to your pages. Whether an answer engine quotes them once it has them is a different problem, and our GEO citation playbook covers the content changes that move that number.
