Who we are
ArgandBot is the crawler for Argand. It fetches public pages for Argand's independent web index. Search results link to the source page; the crawler does not fetch content to generate replacement answers.
Every request ArgandBot makes identifies itself with this User-Agent:
Mozilla/5.0 (compatible; ArgandBot/1.0; +https://argand.org/bot)What ArgandBot fetches, and why
- HTML pages only. It checks
Content-Typebefore reading a body and never downloads images, video, audio, or archives it can't index. Bodies are capped at 5 MB. - robots.txt — before anything else on your host, cached at most 24 hours.
- Sitemaps listed in your robots.txt, fetched with conditional requests so unchanged sitemaps cost you a 304 and no bytes.
- Recrawls of pages already in the index, to keep results fresh — also conditional (
If-None-Match/If-Modified-Since), so an unchanged page costs almost nothing.
It never logs in, never submits forms, never bypasses paywalls or members-only areas, and carries no cookies. If a page isn't visible to a logged-out visitor, ArgandBot doesn't see it either.
Politeness guarantees
- robots.txt is enforced. ArgandBot parses it under RFC 9309 and refreshes the cached rules within 24 hours. If robots.txt fails with a server error, crawling is paused.
Crawl-delayhonored (up to 30 seconds between requests, per directive).- At most ~1 request per second per site by default — usually far less — with one connection per host, and all subdomains of your site sharing a single budget.
- Backs off on trouble. HTTP 429 and
Retry-Afterare respected; errors and slow responses widen the delay automatically. - A 403 means we stop asking. If your site refuses ArgandBot three consecutive times and has never served it a page, the domain is added to a refusal list. ArgandBot then checks only once a week in case the block was temporary. A single successful response clears it immediately. You do not have to do anything else to stop the normal crawl.
- Conditional GET on recrawls — unchanged pages answer 304 and transfer no body. If your server sends no
ETagorLast-Modifiedof its own, we still askIf-Modified-Sincethe time we last saw the page, so you can answer 304 rather than resend it. - Compressed transfer (gzip / brotli) on every request, keeping your bandwidth bill small.
- Page-level directives respected:
noindex(meta tag orX-Robots-Tagheader) keeps a fetched page out of the index entirely;nofollowkeeps its links out of the crawl frontier.
How to opt out
Site-wide: add this to your robots.txt.
It takes effect within the 24-hour robots cache, usually much sooner:
User-agent: ArgandBot
Disallow: / (A User-agent: Argand section works too, and directives
for * are honored when no ArgandBot-specific section
exists.)
Per page: either of these keeps a page out of Argand's index:
<meta name="robots" content="noindex"> X-Robots-Tag: noindex A noindex page is never published to our index — the
directive is honored at fetch time, not filtered later. You can also
scope it to us alone with <meta name="argandbot" …> or X-Robots-Tag: argandbot: noindex.
Manual removal: use the crawler request form — opt-outs apply automatically the moment you submit, no robots.txt changes and no waiting on a human. Email crawler@argand.org works too.
Verifying it's really us
- The User-Agent is exactly the string shown above, linking to this page.
- ArgandBot currently crawls from Argand's own infrastructure at
135.181.231.218(Hetzner, Finland). A request claiming to be ArgandBot from any other address is an impostor — block it freely. - DNS verification (forward-confirmed, like Googlebot): reverse-resolve the connecting IP —
dig +short -x 135.181.231.218returnscrawler.argand.org— then forward-resolve that name back to the same IP. Both directions must match; anything else is an impostor. - The real ArgandBot always fetches
/robots.txtbefore your pages and obeys it. Impostors rarely bother.
If the crawl origin changes or grows, this page is updated first — it is the canonical record of ArgandBot's identity.
Contact
Anything crawler-related — removal requests, rate concerns, odd behavior in your logs — goes through the crawler request form. Opt-outs apply automatically; everything else reaches the person who wrote the crawler. Prefer email? crawler@argand.org lands in the same place.