Skip to content

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

aivscan

Check how ready a website is for AI crawlers and answer engines — ChatGPT, Perplexity, Gemini, Claude and friends. One static Go binary, zero dependencies.

$ aivscan https://auracite.de

aivscan report for https://auracite.de
HTTP 200   server: cloudflare   title: "KI-Empfehlungen in ChatGPT & Co. prüfen | AuraCite"

AI crawler rules (robots.txt present)
  OK        GPTBot                 OpenAI model training
  OK        OAI-SearchBot          ChatGPT search index
  OK        ClaudeBot              Anthropic model training
  OK        PerplexityBot          Perplexity answer engine
  ...

Machine-readable layer
  OK        llms.txt (12996 bytes)
  OK        llms-full.txt (15362 bytes)
  OK        sitemap.xml
  OK        JSON-LD structured data (Organization, SoftwareApplication, FAQPage, ...)
  OK        Open Graph tags (og:title + og:description)
  MISSING   single <h1> (found 2)
  OK        meta robots allows AI use

Score: 95/100   Grade: A

What it checks

Area Checks Points
robots.txt Per-bot verdict for 19 AI crawlers/fetchers using Google-style matching: most specific user-agent group wins, longest matching path wins, allow beats disallow on ties. Supports * wildcards and $ anchors. 30
llms.txt / llms-full.txt Presence and size of the emerging machine-readable content layer 25
sitemap.xml Presence 10
meta robots Surfaces noai / noimageai opt-outs and max-image-preview 10
JSON-LD Collects @type values, including inside @graph 15
Open Graph og:title + og:description 5
Structure Exactly one <h1> (answer engines like unambiguous pages) 5

Grades: A ≥ 85, B ≥ 70, C ≥ 55, D ≥ 40, else F.

Install

go install github.com/G5317/aivscan@latest

Or build from source:

git clone https://github.com/G5317/aivscan && cd aivscan
go build -o aivscan .

Flags

--json <url>          machine-readable report
--markdown <url>      report for pasting into issues/PRs
--fail-under 70 <url> exit 1 below threshold — use in CI to guard regressions
--timeout 20s         per-request timeout

Example: two very different sites

The tool is honest about trade-offs. A publisher that deliberately blocks training bots scores low even though it is a healthy site:

  • aivscan https://www.spiegel.de → 46/100 (D) — GPTBot, ClaudeBot, CCBot, Bytespider and others blocked; PerplexityBot and Googlebot allowed. That is an editorial choice, not an accident — aivscan makes it visible.
  • aivscan https://auracite.de → 95/100 (A) — all AI crawlers allowed, full machine-readable layer present.

CI usage

- run: go install github.com/G5317/aivscan@latest
- run: aivscan --fail-under 70 https://yoursite.example

Honest limits

  • One page per scan (the URL you pass, plus root-level robots/llms/sitemap). It does not crawl the whole site.
  • HTML is parsed with targeted regexes, not a full DOM engine. Good enough for meta/JSON-LD/h1 detection, not a browser-grade renderer.
  • The bot list reflects commonly declared user-agents as of 2026-09. Vendors add and rename crawlers; a no-rule verdict means unmanaged, not safe.
  • Eligibility is not visibility. Passing every check here means answer engines can read and cite you — not that they do. Measuring whether they actually mention, rank, and cite you is what AuraCite does over time.

Why

Built by the founder of AuraCite, an AI-visibility (GEO) analytics platform. The checks come straight from real customer audits.

License

MIT

About

Check how ready a website is for AI crawlers and answer engines (ChatGPT, Perplexity, Gemini): per-bot robots.txt rules, llms.txt, JSON-LD, meta robots. Single Go binary, zero deps.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages