The getdomaindata.com crawler

getdomaindata.com operates an automated crawler (which identifies itself as getdomaindata/1.0) that visits the homepage of publicly reachable websites to detect the technologies and infrastructure they run. This page explains what it does and how to opt out.

How to identify it

Every request our crawler makes carries this User-Agent:

Mozilla/5.0 (compatible; getdomaindata/1.0; +https://getdomaindata.com/bot)

What it does

  • Requests the homepage of publicly reachable sites, over standard web ports (TCP 80 and 443) only.
  • Identifies itself on every request with the User-Agent above.
  • Honors robots.txt - if your site disallows our crawler (User-agent getdomaindata, or *), we will not request its pages.
  • Detects the technologies a site uses; page content is not retained.
  • Visits any given site infrequently, at a low, polite request rate.

What it does not do

  • No port scanning, vulnerability probing, brute-forcing, or exploitation of any kind.
  • No logins, form submissions, or attempts to access non-public areas.
  • No connections to anything other than public web servers on ports 80/443.

Opt out

The simplest way to opt out is your robots.txt: disallow the User-agent getdomaindata (or *) and our crawler will not request your pages.

Or use the form below to have a domain permanently excluded - one apex domain (covers all its subdomains) or one specific subdomain (your apex stays scannable) per request. Exclusions persist across all future crawls, and you can optionally have existing data about the domain removed from the site as well. We verify your email address; an address at the domain itself gets the request granted automatically.

Use an address at the domain you're excluding and the request is granted automatically; other addresses go to a short manual review.

Prefer email, need to exclude an IP range, or excluding many domains at once? Write to optout at this domain with the domains or network ranges you want removed. We add exclusions promptly and they persist across future crawls.

Network operators & abuse reports

If you operate a network and observed traffic from our crawler, please reach out to abuse at this domain. We honor exclusion requests, maintain a permanent do-not-contact list for reported sensor and honeypot ranges, and are happy to coordinate on attribution or whitelisting.

The same address takes reports of illegal or obscene content in our results. Send the domain and what you saw, not the content. The Terms say what we remove and what we withhold.

Operated by getdomaindata.com. Abuse: abuse at this domain · Opt-out: the form above or optout at this domain.