Agent readiness report
bhpioneer.com
What an AI agent can and cannot do on this site, measured by an unauthenticated crawl and scored against the published rubric.
Scanned · rubric v1.0.1 · scored as a e-commerce site · re-scan (results are cached for seven days)
What to fix first
Ranked by impact first, then by how cheap the fix is — the same ordering the engine uses.
robots.txt blocks GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot, Google-Extended. If that is deliberate, nothing to do. If it is inherited from a template or a platform default, add explicit Allow directives — these agents are how a growing share of buyers now reach you.
1 page(s) contain JSON-LD that fails to parse — malformed structured data scores zero, exactly as if it were absent. Validate the output of whatever template generates it.
Fix: 15 of 75 inputs have no associated label. A placeholder is not a label — an agent filling your form has no way to know what an unlabelled field wants.
Found API documentation link. A second machine surface — an OpenAPI spec or an MCP server — takes this to full marks.
Publish an llms.txt at the root: an H1, a one-line summary, then described links to your most useful pages. It is a text file and it takes an hour — the cheapest points on this rubric.
Every check, with the evidence
No finding without evidence: each result shows what we actually captured during the crawl.
ADiscovery & access
- A1Pass
robots.txt exists and parses
2/2What we captured
{ "groups": 31, "status": 200, "sitemaps": 28 } - A2Partial
Major AI crawlers allowed
1/6robots.txt blocks GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot, Google-Extended. If that is deliberate, nothing to do. If it is inherited from a template or a platform default, add explicit Allow directives — these agents are how a growing share of buyers now reach you.
What we captured
{ "of": 6, "allowed": 1, "blocked": [ "GPTBot", "OAI-SearchBot", "ClaudeBot", "PerplexityBot", "Google-Extended" ] } - A3Pass
XML sitemap valid and referenced
3/3What we captured
{ "exists": true, "urlCount": 5, "referencedInRobots": true } - A4Fail
llms.txt present
0/4Publish an llms.txt at the root: an H1, a one-line summary, then described links to your most useful pages. It is a text file and it takes an hour — the cheapest points on this rubric.
What we captured
{ "status": 429, "present": false } - A5Pass
Bot user-agent HTTP health
3/3What we captured
{ "probes": [ { "name": "GPTBot", "status": 200, "ttfbMs": 725 }, { "name": "ClaudeBot", "status": 200, "ttfbMs": 77 }, { "name": "PerplexityBot", "status": 200, "ttfbMs": 80 } ], "medianTtfbMs": 80, "edgeBlockingDespiteRobots": false } - A6Fail
No hard interstitial
0/2Detected: CAPTCHA. An agent cannot dismiss a consent dialog or solve a challenge — whatever sits behind it is unreachable. Cookieless analytics removes the need for a banner entirely.
What we captured
{ "markers": [ "CAPTCHA" ] } - A7Not applicable
Server-rendered content parity
0/5No headless render was captured, so parity could not be measured. Re-run with rendering enabled.
What we captured
{ "rendered": false }
BUnderstanding
- B1Partial
JSON-LD present and valid
2/51 page(s) contain JSON-LD that fails to parse — malformed structured data scores zero, exactly as if it were absent. Validate the output of whatever template generates it.
What we captured
{ "pagesCrawled": 3, "pagesWithValidJsonLd": 2, "homepageHasValidJsonLd": false, "pagesWithInvalidJsonLd": 1 } - B2Pass
Correct schema types for the site type
5/5What we captured
{ "matched": [ "Product" ], "missing": [], "siteType": "ecommerce", "typesFound": [ "CreativeWork", "NewsArticle", "NewsMediaOrganization", "Organization", "Product" ] } - B3Fail
Semantic HTML structure
0/4Headings and landmarks are how a machine builds an outline of the page. One h1, no skipped levels, and real main/nav/footer elements — all template-level fixes, applied once.
What we captured
{ "problems": [ "https://bhpioneer.com: 0 h1 elements", "https://bhpioneer.com: skipped heading level", "https://bhpioneer.com: missing main", "https://bhpioneer.com/opinion: skipped heading level", "https://bhpioneer.com/opinion: missing main", "https://bhpioneer.com/a_and_e: skipped heading level", "https://bhpioneer.com/a_and_e: missing main" ], "singleH1": 2, "allLandmarks": 0, "pagesCrawled": 3, "noSkippedLevels": 0 } - B4Partial
Meta and Open Graph
1/3Fix: 3 description(s) outside 50–160 characters; 3 page(s) missing og:title, og:description or og:image.
What we captured
{ "goodTitles": 3, "pagesCrawled": 3, "uniqueTitles": true, "goodDescriptions": 0, "completeOpenGraph": 0 } - B5Partial
Clean text ratio
2/3Visible text is 6.9% of your markup. Deeply nested wrapper elements make a page expensive to parse and dilute the content an agent extracts. Script and style contents are already excluded from this measure, so this is markup weight, not framework overhead.
What we captured
{ "perPage": [ 0.053, 0.07, 0.084 ], "averageRatio": 0.069 } - B6Fail
Alt text coverage
0/2172 of 307 content images have no alt text. An agent cannot see the image; the alt attribute is the only description it gets.
What we captured
{ "withAlt": 135, "coverage": 0.44, "contentImages": 307 } - B7Pass
Canonical and duplication hygiene
3/3What we captured
{ "hostVariants": [ { "url": "https://www.bhpioneer.com", "status": 200 }, { "url": "http://bhpioneer.com", "status": 301, "location": "https://www.bhpioneer.com/" } ], "pagesCrawled": 3, "selfConsistentCanonicals": 3 }
CActionability
- C1Partial
Form usability
4.4/5Fix: 15 of 75 inputs have no associated label. A placeholder is not a label — an agent filling your form has no way to know what an unlabelled field wants.
What we captured
{ "forms": 51, "inputs": 75, "realSubmit": true, "unlabelled": 15, "wrongTypes": 0 } - C2Partial
Navigable link graph
2/4Fix: 14% of links use contextless text like "click here"; no breadcrumbs or clear URL hierarchy.
What we captured
{ "anchors": 370, "genericLinks": 52, "genericRatio": 0.14, "hasBreadcrumb": false, "clickableNonAnchors": 0 } - C3Pass
Interactive element semantics
4/4What we captured
{ "realButtons": 39, "positiveTabindex": 0, "clickableNonButtons": 0, "keyboardTrapsTested": false } - C4Partial
Machine endpoints and agent manifests
3/5Found API documentation link. A second machine surface — an OpenAPI spec or an MCP server — takes this to full marks.
What we captured
{ "found": [ "API documentation link" ] } - C5Partial
Commerce actionability
1/4Fix: no product page reachable within the crawl depth; no Offer schema carrying price and availability; cart or checkout path did not return 200 without login. Agentic commerce depends on an agent being able to read a price and reach a checkout without an account.
What we captured
{ "offerSchema": false, "cartReachable": false, "expressCheckout": true, "productPageFound": false } - C6Pass
On-site search
3/3What we captured
{ "param": "q", "usesGet": true, "searchForm": true, "searchAction": true }
DTrust & freshness
- D1Partial
Transport security
3/4Fix: 3 insecure subresource reference(s).
What we captured
{ "hsts": true, "tlsValid": true, "mixedContentReferences": 3 } - D2Pass
Verifiable identity
4/4What we captured
{ "sameAsCount": 2, "contactDiscoverable": true, "hasOrganizationSchema": true } - D3Partial
Policies discoverable
2/4Fix: shipping and returns pages not both linked.
What we captured
{ "hasTerms": true, "siteType": "ecommerce", "hasPrivacy": true, "hasReturns": true, "hasShipping": false } - D4Partial
Freshness signals
1/4Fix: sitemap has no lastmod values; only 0% of sitemap URLs modified in the last 90 days. Staleness is a ranking signal for models deciding what to cite.
What we captured
{ "sitemapUrls": 5, "withLastmod": 0, "visibleDates": true, "modifiedLast90Days": 0 } - D5Partial
Identity consistency
1/3Fix: site name differs across title, og:site_name and schema name. An unexplained name mismatch reads as a signal the site may not be what it claims.
What we captured
{ "title": "bhpioneer.com | Local & Independent Since 1876", "hasFavicon": true, "hasOgImage": true, "ogSiteName": "Black Hills Pioneer", "schemaName": "Black Hills Pioneer" } - D6Fail
security.txt
0/2Publish /.well-known/security.txt with a Contact field and a future-dated Expires field. It is a five-line text file.
What we captured
{ "status": 404, "present": false } - D7Partial
Multi-page stability
2/4Fix: 3 page(s) did not return 200: https://bhpioneer.com/site/about.html (429), https://bhpioneer.com/site/contact.html (429), https://bhpioneer.com/video (429).
What we captured
{ "nonOk": [ { "url": "https://bhpioneer.com/site/about.html", "status": 429 }, { "url": "https://bhpioneer.com/site/contact.html", "status": 429 }, { "url": "https://bhpioneer.com/video", "status": 429 } ], "slowPages": 0, "pagesCrawled": 6, "consistentTemplate": true }
Keep this report
Get the permanent link and the ranked fix list by email. One message, no list.
Want this fixed?
We audit in depth and implement the fixes — in your codebase or your CMS, scoped by architecture.
Own this site?
This report and its leaderboard entry come from a public crawl. Prove you own the domain with an email address on it, and you can hide the report or publish it again yourself, in minutes.
To block future scans instead, disallow AgentFriendlyRankBot in your robots.txt — see how our crawler works.
Scan your own site
Same rubric, same evidence standard, about a minute.