INDEPENDENT SYSTEMS JOURNALEDITION 01 / 2026
OPEN SYSTEMS
FIELDBOOK
RSS
WEB OPERATIONS
WEB OPERATIONS • RESEARCHED 2026-09-24

Bot traffic on a small site: measure before you block

A practical way to separate crawlers, broken clients and abuse using request logs, rate patterns and a clear response policy.

A busy access log does not tell you who is reading. One crawler can generate more requests than a thousand human visitors, while a single address can represent many people behind a shared network. Begin with a measurement window and a question: what resource is actually under pressure?

Bot-traffic triage showing discovery, looping requests, and load pressure as distinct patterns.
Three log patterns call for different responses; a bot label alone is not enough.

1. Establish a baseline

Measure requests per minute, distinct paths, error rate, response size and bandwidth over a representative week. Segment by status code and user agent, but do not treat the user-agent string as proof of identity. Compare a normal period with the suspected spike. If the site stays responsive, the cost may be more important than the request count.

2. Classify behavior, not labels

A bot may fetch robots.txt and a few changed pages; a scraper may sweep the entire archive repeatedly; a broken client may loop on one missing URL. Look for path sequences, intervals and concentration. A repeated 404 on an old page suggests a link or redirect problem, not automatically a hostile actor.

A SIMPLE TRIAGE
Discovery

Many distinct paths, slow pace → check crawl policy and cache.

Loop

One path, high frequency → investigate client or redirect.

Pressure

High rate plus errors or load → rate limit or challenge carefully.

3. Choose the lightest intervention

Cache static assets, fix accidental infinite redirects and return an honest 404 for missing pages. Use robots.txt to state a crawl preference. RFC 9309 makes clear that robots rules are requests to crawlers, not access control. When necessary, apply rate limits or challenges to the observed pattern, then check that legitimate readers and search crawlers still reach the site.

4. Measure the result

Record the rule, start time, target path and expected change. Compare origin load, error rate and real page delivery afterward. Revert a rule that reduces legitimate access without meaningfully reducing cost. Keep logs only as long as operationally necessary and avoid publishing raw IP addresses.

For Caddy operators

Caddy's log directive enables HTTP access logging and supports JSON output. Verify what your current configuration records before making claims about bot share or bandwidth.

REFERENCE DESK

Sources & scope

These sources support the technical principles above. Commands and compatibility can change; verify the current version before changing a machine.

  1. RFC 9309: Robots Exclusion Protocol
  2. Caddy documentation: log directive
  3. Caddy documentation: how logging works
A note on this domain

This is a new, independent edition. Its articles are original editorial work inspired by the domain's technical history; we are not the former author. Meet the earlier author in the archive.

KEEP EXPLORINGBrowse all field notes