Who we are

ArgandBot is the crawler for Argand. It fetches public pages for Argand's independent web index. Search results link to the source page; the crawler does not fetch content to generate replacement answers.

Every request ArgandBot makes identifies itself with this User-Agent:

Mozilla/5.0 (compatible; ArgandBot/1.0; +https://argand.org/bot)

What ArgandBot fetches, and why

  • HTML pages only. It checks Content-Type before reading a body and never downloads images, video, audio, or archives it can't index. Bodies are capped at 5 MB.
  • robots.txt — before anything else on your host, cached at most 24 hours.
  • Sitemaps listed in your robots.txt, fetched with conditional requests so unchanged sitemaps cost you a 304 and no bytes.
  • Recrawls of pages already in the index, to keep results fresh — also conditional (If-None-Match / If-Modified-Since), so an unchanged page costs almost nothing.

It never logs in, never submits forms, never bypasses paywalls or members-only areas, and carries no cookies. If a page isn't visible to a logged-out visitor, ArgandBot doesn't see it either.

Politeness guarantees

  • robots.txt is enforced. ArgandBot parses it under RFC 9309 and refreshes the cached rules within 24 hours. If robots.txt fails with a server error, crawling is paused.
  • Crawl-delay honored (up to 30 seconds between requests, per directive).
  • At most ~1 request per second per site by default — usually far less — with one connection per host, and all subdomains of your site sharing a single budget.
  • Backs off on trouble. HTTP 429 and Retry-After are respected; errors and slow responses widen the delay automatically.
  • A 403 means we stop asking. If your site refuses ArgandBot three consecutive times and has never served it a page, the domain is added to a refusal list. ArgandBot then checks only once a week in case the block was temporary. A single successful response clears it immediately. You do not have to do anything else to stop the normal crawl.
  • Conditional GET on recrawls — unchanged pages answer 304 and transfer no body. If your server sends no ETag or Last-Modified of its own, we still ask If-Modified-Since the time we last saw the page, so you can answer 304 rather than resend it.
  • Compressed transfer (gzip / brotli) on every request, keeping your bandwidth bill small.
  • Page-level directives respected: noindex (meta tag or X-Robots-Tag header) keeps a fetched page out of the index entirely; nofollow keeps its links out of the crawl frontier.

How to opt out

Site-wide: add this to your robots.txt. It takes effect within the 24-hour robots cache, usually much sooner:

User-agent: ArgandBot
Disallow: /

(A User-agent: Argand section works too, and directives for * are honored when no ArgandBot-specific section exists.)

Per page: either of these keeps a page out of Argand's index:

<meta name="robots" content="noindex">
X-Robots-Tag: noindex

A noindex page is never published to our index — the directive is honored at fetch time, not filtered later. You can also scope it to us alone with <meta name="argandbot" …> or X-Robots-Tag: argandbot: noindex.

Manual removal: use the crawler request form — opt-outs apply automatically the moment you submit, no robots.txt changes and no waiting on a human. Email crawler@argand.org works too.

Verifying it's really us

  • The User-Agent is exactly the string shown above, linking to this page.
  • ArgandBot currently crawls from Argand's own infrastructure at 135.181.231.218 (Hetzner, Finland). A request claiming to be ArgandBot from any other address is an impostor — block it freely.
  • DNS verification (forward-confirmed, like Googlebot): reverse-resolve the connecting IP — dig +short -x 135.181.231.218 returns crawler.argand.org — then forward-resolve that name back to the same IP. Both directions must match; anything else is an impostor.
  • The real ArgandBot always fetches /robots.txt before your pages and obeys it. Impostors rarely bother.

If the crawl origin changes or grows, this page is updated first — it is the canonical record of ArgandBot's identity.

Contact

Anything crawler-related — removal requests, rate concerns, odd behavior in your logs — goes through the crawler request form. Opt-outs apply automatically; everything else reaches the person who wrote the crawler. Prefer email? crawler@argand.org lands in the same place.