Mitigation Measures for Self-Hosted Gitea Instances Overwhelmed by Crawler Traffic

Contents

Symptoms

A self-hosted Gitea started consuming abnormal bandwidth. The provider panel showed the traffic allowance draining at an alarming rate, with the monthly quota running out fast. Aggregating access logs by response size, the git domain accounted for the vast majority of traffic, dominated by /owner/repo/compare/ version-comparison pages with single responses up to 10MB.

The crawler pattern was obvious: Amazonbot, Lightpanda, and a rotating pool of IPs faking old Chrome/Firefox user agents, repeatedly fetching blame and compare pages. Most unique IPs made only one or two requests — classic distributed scraping.

Causes

  1. Crawler user agents vary widely, from well-known bots (Amazonbot, Lightpanda) to rotating IPs faking browser UAs.
  2. Heavy pages are large — compare pages exceed 10MB uncompressed — so every crawl burns real bandwidth.
  3. Per-IP rate limiting does nothing against a rotating IP pool, where most IPs make one or two requests and move on.

Fixes

1. Split HTTP and SSH onto separate domains

Once HTTP goes through the Cloudflare proxy, SSH cannot follow: the proxy layer only forwards HTTP(S) ports, so port 22 must either go gray-cloud (DNS only) or use the paid Spectrum product. Split into two domains:

  • git.example.com: web traffic, DNS proxied (orange cloud), all through Cloudflare.
  • ssh.example.com: SSH traffic, DNS only (gray cloud), connecting straight to the server.

Configure DOMAIN and SSH_DOMAIN separately in Gitea so clone URLs render correctly:

[server]
DOMAIN = git.example.com
SSH_DOMAIN = ssh.example.com
SSH_PORT = 22

2. Cloudflare WAF: Interactive Challenge on heavy paths

Rotating IPs defeat every per-IP counter, but no matter how often the IPs change, crawlers keep hitting the same pages. Move those paths to the Cloudflare edge: suspicious requests get a challenge (Turnstile), and requests that fail never reach the origin. Create the rule in the Cloudflare dashboard → your domain → Security → WAF → Custom rules:

(http.request.uri.path contains "/compare/"
 or http.request.uri.path contains "/archive/"
 or (http.request.uri.path contains "/search"
     and not starts_with(http.request.uri.path, "/repo/search")
     and not starts_with(http.request.uri.path, "/api/v1/repos/search"))
 or http.request.uri.path contains "/blame/")
→ Interactive Challenge
  • Path selection: compare computes diffs and blame walks history line by line — both CPU-heavy; archive generates archives on the fly and search hits the database. These four paths are exactly where the traffic concentrated.
  • The expression uses contains instead of regex: the matches operator requires Business/Enterprise, while Free and Pro plans only have string matching.
  • git commands are unaffected: smart-HTTP request paths are /info/refs, /git-upload-pack and the like, none of which contain those segments.

The contains "/search" clause has a trap: it also matches Gitea’s frontend XHR endpoints /repo/search and /api/v1/repos/search. The repository list on the right of the signed-in homepage loads via fetch("/repo/search?count_only=1"); when the challenge turns that request into a 403, fetch silently fails on the challenge page and the list just shows “no repositories”, with no error anywhere — it looks like a Gitea bug. A full-page navigation hit by a challenge can be passed manually; an XHR never gets that chance. That is why the expression excludes these two endpoints. Before adding path challenges for Gitea, walk through the XHR paths its frontend depends on (/repo/search, /api/v1/*, /user/events).

Interactive Challenge requires every matching request to complete a Turnstile verification — there is no risk-based pass-through. Friendly search-engine crawlers are auto-validated by Cloudflare and are not affected by the challenge.

3. Caddy config

With web traffic behind Cloudflare, Caddy becomes the second line of defense, with the UA blacklist and heavy-page rate limit as a backstop. The three rules in the site block run in order:

git.example.com {
	# Logging: per-domain access log
	import log git.example.com

	# 1. UA blacklist: known crawlers get 403
	@block_bots header_regexp User-Agent "(?i)(SemrushBot|Lightpanda|GPTBot|ClaudeBot|Bytespider|AhrefsBot|MJ12bot|DataForSeoBot|Amazonbot|DotBot|CCBot|cohere-ai|Diffbot|Google-Extended|Meta-ExternalAgent|Meta-ExternalFetcher|Applebot-Extended|Timpibot|ImagesiftBot|Crawl4AI|Scrapy|ChatGLM-Spider|DeepSeekBot|cohere-training-data-crawler|AI2Bot|TikTokSpider|omgili|BLEXBot|Exabot|360Spider|80legs|MegaIndex|Barkrowler|DataCha0s|Nikto|Nmap|Nessus|OpenVAS|Sqlmap|WPScan|Nuclei|Masscan|Shodan|Acunetix|Dirbuster|Whatweb|Havij|Fimap|Jbrofuzz|Zgrab|Censys)"
	handle @block_bots {
		respond 403
	}

	# 2. Rate limit heavy pages (compare/blame etc.) per IP
	@heavy path_regexp heavy_path ^/[^/]+/[^/]+/(?:compare|blame|commit|commits|archive|pulls|issues)(?:/|$)
	handle @heavy {
		rate_limit {
			zone dynamic_zone {
				key {client_ip}
				events 2
				window 10s
			}
			log_key
		}
		reverse_proxy <internal-ip>:3000
	}

	# 3. robots.txt: cooperative crawlers read the rules first
	handle /robots.txt {
		respond `User-agent: *
Disallow: /blame/
Disallow: /compare/
Disallow: /commit/
Disallow: /commits/
Disallow: /archive/
Disallow: /pulls/
Disallow: /issues/`
	}

	# 4. Everything else goes to Gitea
	reverse_proxy <internal-ip>:3000
}

4. Enable origin gzip in Gitea

Add ENABLE_GZIP = true to the [server] section of app.ini, then restart Gitea:

[server]
ENABLE_GZIP = true

ENABLE_GZIP compresses runtime-generated content only; static assets are unaffected. Enable it at the origin rather than at the reverse proxy because the path is origin → proxy → visitor, and under bidirectional billing origin compression saves on both legs, while proxy-only compression saves on one. Caddy’s reverse_proxy adds Accept-Encoding: gzip for clients that omit it and passes the upstream’s compressed response through untouched. No encode directive is needed on the Caddy side.

References

Edit this page

Contents