Trust, Methodology & Privacy

Last updated: August 2026 ยท This page is the standard we hold ourselves to.
What Tolljar measures What we do NOT know Methodology Limitations Corrections & opt-out Privacy Terms

What Tolljar measures

Tolljar reads publicly available signals a website chooses to expose about machine access: its robots.txt, discoverable RSL licensing links, llms.txt, and homepage HTTP/HTML metadata. From these we describe a site's AI-access posture โ€” which known AI crawlers are allowed, restricted, or blocked, and whether a machine-readable license is declared.

What we do NOT know

Tolljar's public scanner does not tell us how many times an AI company actually accessed a site. It measures publicly observable access rules and machine-readable rights signals โ€” not traffic. Verified traffic data will be reported separately, and only for sites that install Tolljar Observe.

We also do not, and will not, publish: a site's economic value, how much an AI "earned" from it, how many times it was actually read, a "most valuable site" ranking, copyright-infringement estimates, or any amount an AI company supposedly "owes."

Methodology

Sample

Our first published dataset is 110 independent US websites, selected as well-known independent publishers, blogs, and niche-data sites (not Fortune-500 or platform-owned properties). All figures are reported as "of the 110 sites scanned", never as "of the independent web."

Fetch

For each domain we retrieve robots.txt (following the published HTTP status), check for RSL association via robots License:, HTML <link rel="license">, HTTP Link headers, and probe llms.txt. We check a fixed list of 16 named AI crawler identities (e.g. GPTBot, ClaudeBot, Google-Extended, CCBot, PerplexityBot, and others).

Classification

A transparent score (0โ€“100) combines: AI-specific rules present, training crawlers controlled, and a machine-readable license declared. The exact formula is shown on every scan result.

Limitations

Corrections & opt-out

If a site owner believes a Tolljar profile is inaccurate, or wants their site removed from the public Index, email hello@tolljar.com with the domain. We show a last scanned timestamp on every profile and will re-scan, correct, or remove on request. Verified owners can manage their own profile directly in the console.

Privacy

Scanning a domain requires no signup. When you claim a site we store the email you provide and a session identifier to bind ownership; we apply basic per-IP rate limiting to prevent abuse. We do not sell personal data, and we do not place personal data in URLs. Verification tokens prove domain control only. You can request deletion of your account data at hello@tolljar.com.

Terms (plain-English summary)

Questions, corrections, or opt-out: hello@tolljar.com. ๐Ÿซ™