TrustVendorBot — Crawler Policy

User agent: TrustVendorBot/1.0 (+https://trustvendor.co/bot)

Opt-out: [email protected] — honoured within 5 business days.

TrustVendor is a vendor intelligence platform. We crawl publicly accessible vendor pages — trust centres, subprocessor lists, security pages, DPAs, changelogs, and filings — to build an evidence-grade record of what vendors claim and how it changes over time. We sell to compliance teams, not to advertisers. We follow the policy below exactly.

Policy

  • Identify the bot with a contact URL in the user agent. Every request carries TrustVendorBot/1.0 (+https://trustvendor.co/bot) so you can identify and block us if you choose.
  • Honor robots.txt and Crawl-delay. We cache robots.txt per eTLD+1 for 24 hours, parse it with a standard library, and record every check. A disallowed path is skipped and logged with robots_ok = false so we can prove compliance.
  • Never authenticate, never bypass a paywall, never solve a CAPTCHA, never use residential proxies to evade blocks. If a page requires a login, we do not fetch it.
  • Rate limit per apex domain. Default: 1 request per 2 seconds per host, configurable lower per source. We back off on 429 and 503 responses with exponential jitter and do not resume until the host recovers.
  • Prefer official APIs and RSS over HTML scraping wherever both exist. Structured feeds reduce load on your origin and give us cleaner data.
  • Honour opt-out within 5 business days. Email [email protected] with your domain. Opted-out vendors still get a profile page built from public registry data (NVD, EDGAR, Certificate Transparency), clearly labelled as such. We will not crawl your origin again.
  • Never re-publish substantial verbatim content. We store snapshots for evidence integrity; the public vendor page shows claims and short cited spans with attribution and a link to the source. Full snapshots are visible only to tenants who monitor that vendor, in an evidence viewer, for audit purposes.

What we crawl

We focus on content vendors have already chosen to publish: trust centres, subprocessor pages, security pages, data processing addenda, changelogs, and SEC filings. We do not scrape news publishers (we license that feed). We do not crawl anything behind authentication.

Contact

Questions, opt-out requests, or concerns: [email protected]. We respond within 2 business days.