@monsieuricon that's a pretty good share, even if it looks bad!
@monsieuricon I don't understand why there's so much scraping traffic. I assume it's AI bots which are mostly revisiting pages which have already been scraped, but why?
@monsieuricon Oh so it's maybe "if we don't revisit every page every day we might miss a change which our competitor does have in their training data"?
@monsieuricon I think the hours/day unit can be simplified to an average number of full-time core, "promoting" the annotations to the X-axis.
This reminded me all those annoying "GW.h/year" in the press that one must divide by 8760 to get a meaningful GW average. Because W.h does not confuse people enough :-)
@monsieuricon I'm curious, if you can distinguish scrapers from legit traffic, why not block them or serve them garbage?
@jgrg @monsieuricon You're assuming there's logic in the mix. Stop that. There's no logic.
Hmm, wonder how big a Stagit version of the linux kernel project would be. Not sure I want to find out though. That could actually be a project where generating the diff pages and all dynamically could be more efficient than generating all possible ones upfront even with scrapers...
@monsieuricon i wonder if its worth introducing per-ip concurrent-request throttling just to slow em down? like im sure the logs and timestamps will let you see what a real browser looks like vs a scraper and you can use those timing differences to quietly tarpit scrapers :D
@monsieuricon that is super peculiar. im an apache guy and i cant rightly think of a 'gate' there. have you considered that anubis thing? or like, if they only make a few requests, surely the first is like, index.html, then the rest are elements from that - they'd have referrers? maybe enforcing referrer checking could help there? ive had to do a ton of this sorta 'ban evasion stuff'. (i helped run a camsite in like 2005. its how i learned about mod_rewrite)
@monsieuricon In my case it's gitweb and they reorder the url parameter alphabetically which makes it easy to filter but they'll probably figure it out at some point.
@monsieuricon i don't know how this would be done but I can't shake the feeling that the technical responses aren't the right ones. This has the feeling of a legal issue, even if it technically isn't and the guilty parties are well hidden.
Well if you find something better let me know. (or if you end up patching git that you can just clone off of a git checkout with .git directory on a webserver*)
There are javascript libraries for parsing it, so technically you could just do all of that clientside too then.
* not sure if it can already do that, never tried.
@jgrg @monsieuricon combo of dumb, lazy, careless, malicious, ignorant -- its a soup of suck that blurs together into same downstream effect: a de facto tax on the rest of us, one we never agreed to in advance
@monsieuricon oof, man, those git clone numbers... Can't wait until those bundle-uri announcements get folded into the protocol and honored by default by clients.
@monsieuricon Actually that's less ludicrous than the load generated by bots trying to scrape my near-barren corner of the internet websites.