The heaviest signal on the whole scale is one line of a text file
Sixteen of the seventy-five points in the AI Discoverability reading come down to whether robots.txt lets AI crawlers in. No other single signal in the registry is worth that much, and most people have not read theirs since 2023.
Updated
Whether AI crawlers are allowed is worth 16 of the 75 points available in the AI Discoverability reading, which makes it the heaviest single signal in the entire registry. It is one line in a plain text file, and a large number of sites are failing it by accident.
Why it weighs that much
Because it is the only signal where the failure is total rather than partial.
Almost everything else this thermometer reads is a matter of degree. A missing meta description costs you a better snippet. Thin content costs you against a deeper competitor. No structured data means a machine has to infer rather than read.
A site-wide Disallow against GPTBot is not a degree. The page is not fetched, so nothing else on it is ever evaluated. Every other signal becomes decorative at once.
Weights in this registry are ordinal judgements, not measurements, and I have said so on the registry page. But they encode a real distinction: a page blocked from retrieval is not slightly less visible than a page missing an og:image.
The accidental failures
Three, in rough order of frequency.
The 2023 block nobody revisited. There was a wave of reasonable-sounding advice about blocking AI crawlers. Plenty of people followed it. Very few have reopened the file since, and in the meantime the same directives came to cover live retrieval as well as training collection.
The staging leftover. User-agent: * / Disallow: /, written to keep a development site out of the index, shipped to production with everything else. This costs the search reading too and it is the single most expensive typo in the business.
Google-Extended confusion. This is the one that catches careful people, because they read the documentation and still got it backwards. Google-Extended governs training and grounding use. It has nothing to do with whether Googlebot indexes you. People block it believing they are opting out of AI training while staying in search, which is roughly correct, and then are surprised not to appear in AI Overviews, which is the part they did opt out of.
What this is not
It is not an argument that you should let AI crawlers in.
There are real reasons not to. If your content is your product, if you license it, if you have a genuine commercial position on machine consumption of your work, blocking is a legitimate decision and I am not going to pretend it is a mistake.
The thermometer will still charge you 16 points for it, and that is correct too, because the reading measures visibility rather than virtue. A site that has deliberately chosen to be invisible to answer engines is invisible to answer engines. The number is not a judgement of the decision. It is a description of the consequence.
What I would push back on is the third category: sites where nobody made a decision at all, and a file is quietly doing something its owner has not thought about in two years.
How to check, in ninety seconds
- Open
yoursite.com/robots.txtin a browser. It is a public file, there is nothing to install. - Look for any
Disallow: /at all. Note which user-agent block it sits under. - Look specifically for
GPTBot,ClaudeBot,PerplexityBot,OAI-SearchBot,CCBot,Google-Extendedandanthropic-ai. - For each one you find, answer honestly: did I decide this, or did I inherit it?
That is the whole audit. It is free, it needs no tool including this one, and it is worth more than any other ninety seconds you will spend on this subject.
Take a reading afterwards if you want the number, but the file is the thing.
Filed under: robots, ai-crawlers, weighting