The Cartiq crawler
If you found this page in your server logs, this is what we were doing and how to stop us.
How we identify ourselves
Cartiq/1.0 (+https://cartiqhq.com/bot)What we fetch
A scan is triggered by a person entering your domain. We fetch at most 14 URLs, spaced 500ms apart, within a 20 second budget:
- robots.txt
- your homepage
- sitemap.xml and, if present, a product sitemap
- /products.json, if it exists
- one product page
- llms.txt and one /.well-known/ discovery path
- your shipping and returns policy pages, if linked from the homepage
We never submit forms, never log in, never add anything to a cart, and never execute JavaScript. We read a maximum of 3MB per response.
How to opt out
Add this to your robots.txt. We honour it strictly and immediately — after which we fetch only robots.txt itself from your domain, and nothing else.
User-agent: Cartiq
Disallow: /You can also email bot@cartiqhq.com and we will exclude your domain at our end.
What we store
Parsed signals only — never page bodies, never raw HTML. We store whether structured data exists and what it contains, not copies of your pages. The full method is published here.