Stratezik, Toronto

llms.txt: Does It Actually Get You Cited by AI? (We Checked 50 Toronto Startups)

Toronto startups adopt llms.txt at triple the global rate. No major AI engine documents support for it, and Google ignores it outright. Here is what the data says.

Shah Md. Rifat
By Shah Md. Rifat
Updated 2026-07-19
llms.txt: Does It Actually Get You Cited by AI? (We Checked 50 Toronto Startups)

The short version:

  • One in three funded Toronto startups (33%) publishes an llms.txt file. The global rate is 10.13%. Toronto is adopting it at roughly triple the pace.
  • Meanwhile only 5% of those same companies publish FAQ schema, which AI engines actually parse. The city is enthusiastically doing the thing that does not work and skipping the thing that does.
  • No major AI engine documents support for llms.txt. Google has said outright it does not use the file. Perplexity, OpenAI, and Microsoft have published nothing supporting it either.
  • When researchers modelled what predicts AI citations, removing llms.txt from the model made it more accurate. The file was noise, not signal.
  • Our verdict: llms.txt is cheap and harmless, so ship it if you like, but ship it last. It is not a citation lever, and treating it as one is why a lot of Toronto sites are still invisible.

We keep finding llms.txt files on Toronto sites that fail every check that matters. That pattern is interesting enough to write up properly, so here is what we found and what the wider data says.

What is llms.txt, exactly?

It is a plain markdown file placed at yourdomain.com/llms.txt that lists your key pages with short descriptions, intended as a map for large language models. Think of it as a robots.txt for meaning rather than permission. It is a proposed community standard, not something any search engine requires or has committed to using.

The idea is genuinely reasonable. If a model is going to read your site, why not hand it a clean index instead of making it crawl your navigation? The problem is not the concept. The problem is that adoption ran far ahead of any evidence that the file changes outcomes, and a lot of businesses now believe they have done their AI homework because they shipped one.

Does llms.txt actually get you cited by AI?

There is no evidence that it does. When researchers built a model to predict how often a domain gets cited by AI systems, removing the llms.txt variable improved the model's prediction accuracy. The file contributed noise rather than signal. No study we can find demonstrates a measurable citation lift from adding one.

That result is worth sitting with, because it is stronger than "we found no effect." A variable that makes a model worse is a variable that was misleading it. If llms.txt were quietly helping, you would expect it to carry at least a weak positive signal. It did not.

We want to be fair here: absence of measured lift is not proof that the file can never help. It is proof that nobody has shown it helping, while a lot of people sell it as though someone has.

Does Google use llms.txt?

No. Google has stated it does not use the file, and that position covers AI Overviews and AI Mode as well as classic Search. It carries no weight as a ranking factor, plays no part in AI answers, and publishing one will not change how Google treats your site.

This is the cleanest fact in the whole debate, and it is still the one most often glossed over in the guides telling you to add the file this afternoon. Google was explicit. Anyone implying otherwise is either behind or selling something.

Which AI tools actually read llms.txt?

None of the major ones document support for it. Perplexity has published no support in its crawler documentation, OpenAI's documentation makes no reference to it, and Microsoft and Bing have not announced support. The file's realistic audience today is developer and research agents that fetch it directly when pointed at a documentation site, not the consumer assistants your customers use.

That distinction matters for deciding whether you personally should bother. If you run a developer-facing product with real docs, an llms.txt is a sensible courtesy to the coding agents that will read them. If you run a dental clinic in Scarborough, the agents fetching llms.txt are not the ones your patients are asking for a recommendation.

How many sites actually use llms.txt?

SE Ranking analysed nearly 300,000 domains and found 10.13% had one. Adoption is oddly higher among mid-traffic sites (10.54%) than high-traffic ones (8.27%), which suggests the biggest, best-resourced sites are the least convinced it is worth doing.

That inversion is the tell. Normally a genuine ranking or visibility lever gets adopted fastest by the sites with the most to gain and the most staff to implement it. Here the pattern runs backwards, which is what you would expect from a practice spreading through advice columns rather than through measured results.

What we found in Toronto

Here is our own data, from auditing 50 funded Toronto and GTA startups on a machine-verified 20-point AEO test:

SignalToronto funded startupsGlobal benchmark
Publishes llms.txt33% (14 of 42)10.13%
Publishes FAQPage schema5% (2 of 42)n/a
Missing basic Organization schema52% (22 of 42)n/a
Passes the answer-first test30% (12 of 40)n/a

Read those rows together and the story writes itself. Toronto startups are roughly three times more likely than the global average to publish the file that no engine commits to reading, and only one in twenty publishes the structured Q&A that AI engines demonstrably do parse and quote.

We do not think this is stupidity. We think it is a rational response to bad information. llms.txt is a single file you can ship in an hour and then tell your board you are "AI ready." FAQ schema means arguing with your CMS, writing real customer questions, and validating markup. One of those feels like progress and the other is progress.

Should I add llms.txt to my site?

If you run a documentation-heavy or developer-facing site, yes, it is cheap and the coding agents that read docs will use it. For everyone else it is optional and low priority. Add it only after your site renders without JavaScript, allows the AI crawlers, and carries Organization and FAQ schema. Doing it first is the mistake.

We still score llms.txt in our own audit, and we still check it in our free tool. Not because we think it drives citations, but because adoption is a useful year-over-year signal of whether a market is paying attention to AI search at all. Toronto's 33% tells us founders here are listening. The 5% FAQ-schema number tells us they are listening to the wrong people.

What should I do instead to get cited by AI?

Three things, in order: be reachable (allow OAI-SearchBot and the other AI crawlers, and check your CDN is not blocking them), be readable (your core content must be in the raw HTML, because most AI crawlers do not run JavaScript), and be quotable (a direct answer-first passage plus Organization and FAQPage schema). Those are the levers with evidence behind them.

If you want the specific version for ChatGPT, we wrote that up separately in how to appear in ChatGPT's answers, including the Bing indexing step people miss. And if you would rather see your own gaps than read another list, run your site through our free AEO checker. It tests the render, the crawler access, the schema, and yes, the llms.txt, so you can see where the file sits relative to everything else that is broken.

The honest framing we give clients: llms.txt is a nice-to-have that costs an hour. The render and schema problems are the ones costing you answers right now.

How do I write an llms.txt file, if I decide to?

Put a markdown file at yourdomain.com/llms.txt with an H1 of your company name, a one-line description, and a short list of your most important pages as markdown links with brief context for each. Keep it current. A malformed or stale file is worse than none, because it misrepresents you to anything that does read it.

That is genuinely the whole specification in practice. If someone quotes you a price for llms.txt implementation that sounds like a project, you are being oversold.

The bigger point

The reason we published this is not to dunk on a file. It is that llms.txt has become a proxy for effort in AI search, and proxies for effort are how markets waste years. A third of this city's funded startups can point at their llms.txt. One in twenty can point at FAQ schema. The businesses that flip that ratio over the next twelve months are the ones that will be in the answers.

We will re-run the audit in 2027 and report whether Toronto's llms.txt adoption keeps climbing while the structural work stays flat. Our honest guess is that it will, at least for another year, and that the gap will be worth a lot of money to whoever closes it first.

Sources

  1. Stratezik Toronto Startup Website Audit 2026: 50 funded Toronto and GTA startups scored on a machine-verified 20-point AEO test, May and June 2026. Dataset available on request.
  2. SE Ranking llms.txt adoption study across nearly 300,000 domains, with traffic-band breakdown and the citation-prediction modelling result, as analysed by OrganiKPI: organikpi.com/blog/distribution/llms-txt-adoption-impact.
  3. Google's position that llms.txt is not used for Search, AI Overviews, or AI Mode: baselinelabs.ai/blog/llms-txt-google-search and ecorpit.com/does-llms-txt-help-seo-google-2026.
  4. Stratezik Toronto AI Citation Tracker: monthly measurement of which AI engines name Toronto businesses. Available at stratezik.com/blog/toronto-ai-citation-tracker-july-2026.

Want the optimization playbook, not just the platform overview? Get our free ChatGPT Ads cheat sheet — context hints, bid-floor tests, industry readiness, and the measurement stack from practitioners running real budgets.

Quick answers

Shah Md. Rifat

Shah Md. Rifat
Content Strategist · Stratezik · Toronto, ON · LinkedIn