What each machine could actually read
Read from your robots rules and from what your server answered. No opinions in it.
-
A search engine crawlerWanted the pages and the sitemap
Reached everything it asked for. This is the reader almost every site is already set up for, because it is the one everybody has heard of.
Reached -
An assistant fetching liveWanted the one page a person just asked about
Turned away by the bot rules in front of the site. Nobody chose this: the rule was set to stop scrapers and it stops the assistant answering a question about you.
Blocked -
A training crawlerWanted the public pages
Disallowed in the robots file. That is a legitimate decision and it has a cost, so it belongs on a page where somebody decided it rather than in a file copied from a template.
Blocked -
A machine reading your factsWanted the name, address and hours
Found three versions across the site and the listings. It will pick one, and which one is not up to you until they agree.
Reached -
A machine reading a service pageWanted a sentence it could quote
Reached the page and found no definition in the first eighty words. A page that describes itself is a page that can be quoted, and this one describes a company.
Reached

