Conversation

K. Ryabitsev-Prime ๐Ÿ

Scrapers are out in force in the past week or so. I'm sorry if i've not been responsive to your emails -- I'm trying to figure a way to deal with that. We're currently sustaining about 10,000 requests per minute just to git.kernel.org per single Toronto node, or about 40,000 requests per minute across all git.kernel.org nodes. These are not light hits -- these are random commit-id requests that require running git, or diffs, or plain patch downloads.

Like I said, the bots are solving Anubis difficulty 5, though we manage to weed out about 45% of all requests through Anubis itself, plus some analysis of the requests (some cgit query string combinations are unlikely to be ever hit by humans). Still, we're looking at about 7M requests per day that are attributable to bots. That's a guesstimate, because these bots look and act just like real browsers.

If you answer with "why don't you just..." I am going to be cross with you. There is no "why don't you just" here.
10
46
44

@monsieuricon ๐Ÿ˜ฒ so many, this is like an attack! Is that even legal?

1
0
0
@energisch_ "legal" depends on jurisdiction. For sure, it's not ethical.
1
0
3
It's not overwhelming the systems, but it's like having a 30% "background radiation" of cpu load -- at any given time one third of our node capacity is processing scraper bot requests.
0
2
9

@monsieuricon absolutely not ethical or legitimate, agreed. Those bot armys can destroy the best services, reminds me of DDOS attacks.

0
0
0

@monsieuricon i wonder if a tool like p0f that can fingerprint TCP/IP stacks would be able to find some commonalities between the bots based on stuff like ordering of TCP options

2
0
0

@jann @monsieuricon I'm going to have to resort to TLS fingerprinting at this point. It's a losing game.

1
0
0

@monsieuricon I feel your pain. I've been dealing with something similar. Residential proxy networks are a scourge.

0
0
0

@cadey @jann @monsieuricon claude code and codex are already running selenium/webdriver, and disable the `navigator.webdriver` API. The TLS fingerprint of those is completely uniform.

1
0
0

@monsieuricon That sounds really frustrating and demoralising. I hope we find a solution to this plague soon.

0
0
0

@cadey @jann รค or am I wrong and most traffic is from the data centers that run (pre)training, benchmarks etc?

1
0
0

@freddy @jann we don't know. At some level AI training pipelines is a best guess based on limited available information.

0
0
0

@monsieuricon

Sounds like both fedora and sourceware were getting hit with the same thing.
It has been at least 3 weeks for sourceware even.

0
0
0

@monsieuricon like a find out phase after f... around i would say. I'm not worried, the linux foundation millions will help you sort it out.

0
0
0

@monsieuricon It's very frustrating because if there was a "why don't you just" then we'd be doing that, then scrapers would find a way around it.

The real "why don't you just" is making scraping an explicit crime.

0
0
0

@jann @monsieuricon the bots use hacked home PCs or even your own TCL or LG TVs as a proxy.

0
0
0

@monsieuricon Automating Nick Krause doesn't happen for free.

0
0
0