Bot traffic on a small site: measure before you block
A practical way to separate crawlers, broken clients and abuse using request logs, rate patterns and a clear response policy.
A busy access log does not tell you who is reading. One crawler can generate more requests than a thousand human visitors, while a single address can represent many people behind a shared network. Begin with a measurement window and a question: what resource is actually under pressure?

1. Establish a baseline
Measure requests per minute, distinct paths, error rate, response size and bandwidth over a representative week. Segment by status code and user agent, but do not treat the user-agent string as proof of identity. Compare a normal period with the suspected spike. If the site stays responsive, the cost may be more important than the request count.
2. Classify behavior, not labels
A bot may fetch robots.txt and a few changed pages; a scraper may sweep the entire archive repeatedly; a broken client may loop on one missing URL. Look for path sequences, intervals and concentration. A repeated 404 on an old page suggests a link or redirect problem, not automatically a hostile actor.
Many distinct paths, slow pace → check crawl policy and cache.
One path, high frequency → investigate client or redirect.
High rate plus errors or load → rate limit or challenge carefully.
3. Choose the lightest intervention
Cache static assets, fix accidental infinite redirects and return an honest 404 for missing pages. Use robots.txt to state a crawl preference. RFC 9309 makes clear that robots rules are requests to crawlers, not access control. When necessary, apply rate limits or challenges to the observed pattern, then check that legitimate readers and search crawlers still reach the site.
4. Measure the result
Record the rule, start time, target path and expected change. Compare origin load, error rate and real page delivery afterward. Revert a rule that reduces legitimate access without meaningfully reducing cost. Keep logs only as long as operationally necessary and avoid publishing raw IP addresses.
Caddy's log directive enables HTTP access logging and supports JSON output. Verify what your current configuration records before making claims about bot share or bandwidth.
Sources & scope
These sources support the technical principles above. Commands and compatibility can change; verify the current version before changing a machine.
This is a new, independent edition. Its articles are original editorial work inspired by the domain's technical history; we are not the former author. Meet the earlier author in the archive.