An update on Wayback Machine access
- simonw - 66629 sekunder sedan> Here’s what’s going on. The Internet Archive’s Wayback Machine has been hit by waves of high-volume automated traffic, and we’ve put protections in place to keep the service running.
I'm pretty certain this is scrapers that are trying to workaround blocks on accessing original sites by hitting the Wayback Machine copy instead. Appalling behavior.
In addition to the load it puts on this vital non-profit piece of Internet infrastructure, we've also already seen some sites opt out of the Wayback Machine to prevent their content from being scraped via this alternative route.
- basilikum - 63121 sekunder sedanMad props to the people at the Archive. You are the heros we need in a formerly open internet that is surrendering to evil big corps and closing down free access.
The Internet Archive is in a really bad spot being attacked from multiple sides at once. But — while service has not been consistent — they have maintained open access. I can still access anonymously from Tor without Cloudflare or some other centralized gatekeeper showing me the middle finger.
If you got some money to spare, consider donating to them. They need it.
- robotmay - 51114 sekunder sedanUnrelated, but this week I've been on a memory binge with the Wayback Machine, trying to find old content of mine from the early 2000s. Took me a while but I've finally put together a good bit of info about myself at the time that I'd completely forgotten, and it's all thanks to the Internet Archive storing my little gaming review website from when I was 16. I could barely remember any of the other stuff, it's been genuinely surprising figuring out what I'd forgotten. I couldn't even remember most domains I owned aside from one, which I used as the starting point.
Still can't remember what my Tripod site address was, but that might be lost to time.
Thank you, Archive.org.
- userbinator - 27773 sekunder sedanThe Internet Archive’s Wayback Machine has been hit by waves of high-volume automated traffic
Thank you for not immediately blaming it on "AI bots". I suspect there's some entity manufacturing consent for strong identity/age verification/sanctioned-browser-OS "walled garden" Internet, and these random DDoSes are part of that.
I knew something was up when a few alternative YouTube front-ends I use suddenly put up the 'nubis and complained about the high volumes of traffic they were getting flooded with; of course someone actually going after that data would be aiming their "AI bots" at YouTube directly instead of trying to suck it through a tiny little-known proxy-site, so it really strained the credibility of the argument.
- BeetleB - 65754 sekunder sedanWow, but I wonder if there's more to it.
I've not been able to access web.archive.org from my work computer - I always get the 429 error.
But I then pull out my phone and can access it just fine. All along I was assuming my company was blocking it. Still weird that it happens every time from my work PC and never from my home one.
- timpera - 65644 sekunder sedanI really appreciate the Archive team's efforts to make the Wayback Machine more responsive, and have donated a few times to support them.
Unfortunately, the restrictions have been way too strict for the last few months: from my residential IP, simply moving the mouse on the calendar for a specific URL is enough to get stuck on 429 error messages for a while; and from corporate ISPs (for example, on airport WiFi), you often can't access the WM at all. I hope they'll find a way to relax those.
- CqtGLRGcukpy - 66413 sekunder sedan> We’re getting better at telling abusive bots apart from the people who depend on the Wayback Machine every day. If you think you were blocked in error, email info@archive.org with your operating system, browser, and IP address, and we’ll look into it.
- emaro - 60828 sekunder sedanIt's shame that the AI arms race causes such collateral damage. Free resources were always exploited, but the stakes ($T) and capabilities around AI allow unprecedented abuse. I wish we could go back... :/
I really don't see any solution to this; the scrapers probably wouldn't even mind destroying sources like IA too much, which would leave them as the only "authorative" source of knowledge in the end. Best way is likely regulation incl. hefty (!) fines, but politics are too slow and too fragmented to be effective. So... Enjoy it while it lasts, I guess.
- delis-thumbs-7e - 29608 sekunder sedanI recently remembered a wonderful comic blog from 2010’s that is not online anymore. It was a sonderful Finnish LGTG-thened comic blog that I use to read, then forgot completely until few weeks ago. WM had it stored of course, so I could read through this amazing piece of internet art again.
I really so through some money their way, they do wonderful work.
- roughly - 21332 sekunder sedanBonus points for anyone who’d like to guess how the tragedy of the commons was resolved in the times before the enclosure movement.
- pelican0 - 61369 sekunder sedanIs it established that the scraping scourge of late is primarily driven by AI companies? Anyone aware of any relevant studies?
Beginning to think that the difficulty to browse most websites nowadays due to throttling, is yet another negative externality of AI development that society is forced to bear.
- - 3870 sekunder sedan
- 1vuio0pswjnm7 - 33905 sekunder sedan"Here's what's going on."
Thank you
https://news.ycombinator.com/item?id=49571448
I had a feeling it was due to "AI" companies and developers using "agents"
Not surprised
- thimabi - 60663 sekunder sedanI wonder why doesn’t the Internet Archive require logging-in prior to accessing the Wayback Machine. It would probably help them distinguish humans from bots, at a very little cost to humans.
- xacky - 58241 sekunder sedanThe anti virus industry needs to crack down on crawler and proxy malware, plus ISPs FINALLY need to replace CGNATs with iov6 to stop crawlers banning everyone behind a NAT.
- sicktriple - 42389 sekunder sedanAnyone else feel like making a new internet and starting over
- ilamont - 56801 sekunder sedanShouldn't the solution be to gate bulk access for automated services for a price? Not just the wayback machine, any personal or corporate website?
My blogs are getting slammed and there are issues with cloudflare or captchas.
- petterroea - 38914 sekunder sedanI'd be happy to pay a 5$/month donation to get a higher rate limit/more lenient filter put on me
- tech234a - 65758 sekunder sedanI wonder if they'll end up behind Anubis at some point. I'm surprised it hasn't happened already.
- potato-peeler - 33080 sekunder sedanWayback can’t be accessed through vpn, atleast on proton. Heck, most sites simply block you for using vpn.
- msephton - 45289 sekunder sedanWhy can't they capture OS, Browser, and IP address at the time of error?
All that information is available at the point of failure, the user should not need to email it in.
- xbar - 29183 sekunder sedanThank you for the Wayback Machine. It is immensely powerful for good.
- vlyan - 63942 sekunder sedanunrelated: if a website gets hit with "This URL has been excluded from the Wayback Machine", do existing snapshots get purged or may they still be preserved somewhere?
- int32_64 - 60376 sekunder sedanAre any AI companies using residential proxies to scrape?
- tgtweak - 53805 sekunder sedanCan't wayback machine just offer direct access to the archive for a premium and in doing so, pay for the service?
- MattCruikshank - 64180 sekunder sedanThere was a feature on Amazon Web Services for a while, and I wish it was still there...
Downloader pays.
I make some content and upload it. When you want to download it, you pay Amazon the egress fees. And maybe I get to charge just a bit more, to help me with the Ingress, storage, content creation, etc.
I mean, I know that there's going to be problems with rate limiting, etc. And yes, we have those problems with LLM tokens today. But this just feels like such a useful thing that it baffles me that it doesn't exist already.
- hubraumhugo - 62195 sekunder sedanThere is a HN article on abusive AI crawlers on the front page almost every week, but we rarely talk about the path forward. Web scraping has been around for as long as the internet, and it was fine because we had established best practices (rate limiting, self-identification, robots.txt, etc.) that the industry agreed upon. Now we have AI labs and their crawlers that don't care about any of this gentlemen's agreement. So how do we go from here? Is adding more difficult Anubis and Cloudflare bot protection really the solution? How many millions of human hours and billions in infra costs are we willing to spend on this arms race?
Some approaches that I think are promising:
- A robots.txt V2[0] as a standard way for website owners to state how bots and AI crawlers can use their online content and where to go (e.g. distinguish search from AI training use cases, point to a downloadable file instead of crawling everything, etc.).
- Something like Web Bot Auth[1] as a non-centralized standard for self-identifying bots and agents cryptographically. This would allow websites to allow or deny bots very precisely.
- what else?
[0] https://datatracker.ietf.org/doc/draft-vaughan-machine-reada...
[1] https://datatracker.ietf.org/doc/html/draft-meunier-http-mes...
- brador - 57855 sekunder sedanThe only solution is to make visitors do compute. Compressing files for the archive to access other files would be perfect for this.
Cross verify hashes to prevent cheating.
Ez.
- ignoramous - 64305 sekunder sedan
- UltraSane - 65260 sekunder sedanWhy not put it in S3 with downloader pays?
- lousken - 65007 sekunder sedanAI companies should pay billions to wayback machine for access
- Onavo - 66678 sekunder sedanWhy not just offer a paid endpoint for the crawlers? It's not like the demand is going to go away anytime soon.
It serves nobody except CloudFlare and hardware companies when one side set up blockers and the other side spend money putting VPN SDKs in consumer TVs.
I am also curious how the (Russian?) paywall bypass mirror archive.is is doing given that they are probably subject to similar amounts of traffic.
- mrhcon - 2858 sekunder sedan[flagged]
- halfblood_walks - 49662 sekunder sedan[flagged]
- sehw - 29799 sekunder sedan[dead]
- josefritzishere - 55285 sekunder sedan[dead]
- unkeen - 64729 sekunder sedan[flagged]
- xyst - 64649 sekunder sedan[flagged]
- swingandamiss - 65869 sekunder sedan[flagged]
- msephton - 62085 sekunder sedanI've been getting this error a lot. Asking users to email them with details of their OS, browser, IP address is just crazy. Their support is supposedly already swamped and they are asking for more!? Changes made by IA shouldn't become my responsibility.
Nördnytt! 🤓