How to write an llms.txt that earns its place
The llms.txt convention, why A4 gives three points for three described links and one for a bare list, and a complete example file for a small SaaS company.
Checks A4
Blog category
Pillar A: robots.txt, AI crawlers, edge blocking, sitemaps, llms.txt and server-rendered parity.
The llms.txt convention, why A4 gives three points for three described links and one for a bare list, and a complete example file for a small SaaS company.
Checks A4
Google-Extended is a robots.txt control token, not a crawler. What it switches off, what it leaves alone, three common misconfigurations, and what A2 records.
Checks A2
A vendor-neutral tour of the edge layer, bot rulesets, bot scores, rate limits, JavaScript challenges and slow origins, and how each shows up in the A5 probes.
Checks A5
Anthropic's crawler and its user-triggered fetcher share a prefix and get blocked together by accident. The robots.txt for each intent, and its A2 score.
Checks A2
OpenAI documents three tokens so training, indexing and user fetches are decided separately. The robots.txt that opts out of training only, and its A2 score.
Checks A2
Which of the six A2 tokens a shop, a SaaS product, a publisher and a services firm should allow, what each combination scores, and what each block gives up.
Checks A2
A line-by-line rewrite of a typical legacy robots.txt. What checks A1, A2 and A3 read from it, the REP rules a parser applies, and a finished file to adapt.
Checks A1, A2, A3
Cloudflare's default bot rules turn GPTBot, ClaudeBot and PerplexityBot away before robots.txt is read. How to check in thirty seconds, and what to change.
Checks A2, A5