8/24/2026 at 3:34:49 PM
Long live Anna’s Archive. I stand on the shoulders of the Internet, Wikipedia, Anna’s Archive, Z-Library, LibGen, YouTube, Hacker News, Reddit, and Sci-Hub.I deeply admire the people who are obsessed with their passions and strive to build things that will lay the foundations for others.
by num42
8/24/2026 at 3:44:24 PM
Why don't Google, OpenAI, Anthropic, Facebook & Co defend Anna's Archive publicly?Coming out would be a bold move for them.
by sam_lowry_
8/24/2026 at 3:51:56 PM
Anthropic paying $1.5 billion in fines for downloading Anna's Archive established a moat. They want it to be illegal to pirate books: they can afford the penalties and continue doing it. Just like they want it to be illegal to run local ML inference.by peri-cl
8/24/2026 at 5:13:54 PM
>Anthropic paying $1.5 billion in fines for downloading Anna's Archive established a moat. They want it to be illegal to pirate books: they can afford the penalties and continue doing it.This seems like a "heads I win, tails you lose" type of argument. If Anthropic was pro-piracy I can imagine everyone getting mad that they're flouting law and want to "steal from artists" or whatever.
>and continue doing it
Source? AFAIK they were caught and stopped. That's why there was the recent story about how they were destroying old books to scan them.
by gruez
8/24/2026 at 5:31:20 PM
> From the start, Anthropic “ha[d] many places from which” it could have purchased books, but it preferred to steal them to avoid “legal/practice/business slog,” as cofounder and chief executive officer Dario Amodei put it (see Opp. Exh. 27).https://cdn.arstechnica.net/wp-content/uploads/2025/06/Bartz...
Sounds pretty pro piracy to me.
by bix6
8/25/2026 at 3:16:07 AM
Yes there are people who are both for and against copyright.by austhrow743
8/24/2026 at 11:50:19 PM
They stopped because they already had everything in their model anyway.by wolvoleo
8/24/2026 at 5:28:41 PM
[flagged]by sam_lowry_
8/24/2026 at 5:51:28 PM
> This is an unethical as a company may behave, short of killing people.This is hysterical. No, format shifting old unwanted books is not unethical. The books still exist, in an internal digital library. If copyright law were to change to allow sharing orphan works some day, Anthropic could share them. But under current law, the books are preserved digitally and used for transformative uses that all Claude users benefit from.
by yonran
8/24/2026 at 5:42:30 PM
did not follow all the details, but my understanding is that some form of copyright law nudges in the direction of destroy after scan?by empiricus
8/24/2026 at 8:52:35 PM
> they can afford the penalties and continue doing it.I thought they could've bought just a single copy of each book and use the content to train their models. In that case, it falls into the fair use doctrine and they wouldn't need to pay the fine. And that will be way less expensive than the $1.5B price tag.
by hintymad
8/24/2026 at 9:01:32 PM
That's what they're doing now, when they are established.But when it was a proof of concept, they were using pirated data.
Just like Spotify did.
by Gud
8/24/2026 at 5:03:27 PM
> Just like they want it to be illegal to run local ML inference.Citation?
by vaylian
8/24/2026 at 5:13:00 PM
https://news.ycombinator.com/item?id=49076057 ("Our position on open-weights models (anthropic.com)", 1812 comments)by peri-cl
8/24/2026 at 8:00:23 PM
Over the past decade I've noticed on HN the following order of frequency in choice of words, most common to least:1. Citation
2. Source
3. Reference
Long ago in a career based on original research, I/we ONLY used "reference."
by bookofjoe
8/24/2026 at 8:23:42 PM
While it is definitely over a decade at this point (over two in fact), some of this likely comes from the term [citation needed], that originated on Wikipedia, as a cynical backhanded response to unsourced claims. It has become a catch-all. Language and how it evolves is a pretty interesting subject.by tygon
8/24/2026 at 9:26:17 PM
You're right. Has to be of Wikipedia use origin. Thanks!by bookofjoe
8/24/2026 at 3:58:24 PM
In theory none of them actually got the right to train on illegally downloaded books. Anthropic was simply punished for doing it once.One wonders if they're still doing it.
by spwa4
8/24/2026 at 4:32:39 PM
OpenAI plainly admitted that it is impossible not to do so in a House of Lords inquiry. So, presumably there is no way around it to train models. There is just not enough non-copyrighted data out there.by ungut
8/24/2026 at 4:50:52 PM
You mean this one https://committees.parliament.uk/writtenevidence/126981/pdf/ where they write "it would be impossible to train today’s leading AI models without using copyrighted materials"? That doesn't mean they have to download those materials illegally. For a billion dollars, you can easily buy one legal copy of each book in Anna's Archive and still have some cash left over to run a whole-of-internet scraping operation.by yorwba
8/24/2026 at 8:27:29 PM
I wonder if anyone has run the numbers on what the actual cost, both in cash and logistical headache, contacting so many copyright holders would be. That seems like quite the feat to calculate.by tygon
8/24/2026 at 8:42:21 PM
There's an established network of intermediaries that can supply a large variety of books for a few dollars apiece, so no need to contact copyright holders directly.by yorwba
8/24/2026 at 9:15:44 PM
This is very true. As someone with quite the experience with materials published under Penguin, Scholastic, etc. you effectively have a "dictionary attack" on the matter, rather than true "brute force," but that still leaves quite a list to compile to send to each and is easier for larger titles than smaller ones. I wonder how that leads to a bias in what materials get used for training. You are not getting many local self-published books this way.It is almost like we need a "for use for training" agreement across the board. This would not fix the current issues (at least without substantial work), but going forward would allow for creators (or publishers/rights holders) such as this to designate a work as crawl-able for AI. A robots.txt just for Claude.
by tygon
8/24/2026 at 11:28:00 PM
Doesn't really matter. The incentive structure to steal clearly exists, so why would they even go through the trouble?by ungut
8/24/2026 at 11:25:05 PM
Pretty easy to assertain that they don't acquire them legally due to the plethora of evidence and court cases against them. No copyright holder would be sueing them if they knew they sold the works in the first place.I always wonder why y'all feel the need for these impressive mental gymnastics. You can use the models /and/ think they are trained unethically. Living through the ambiguity without abandoning your ideals completely is a valuable skill these days.
by ungut
8/24/2026 at 5:11:17 PM
I'm pretty sure we would know if they did that. And we don't.Plus this is not legal in the EU (and Canada, and ... let's just say the entire rest of the world, and accept that I'll be wrong for one or two smaller countries). Doesn't that matter? Or is only Mistral disallowed from training on copyrighted materials? Je veux ma chaton fat, goddamit!
by spwa4
8/24/2026 at 6:00:21 PM
AI companies legally acquiring books have indeed been in the news: https://news.ycombinator.com/item?id=49330742And where are you getting the idea that Mistral doesn't train on copyrighted data? There's not a lot of code written by people who've been dead for more than 70 years, but somehow Mistral has been able to release coding models anyway.
by yorwba
8/24/2026 at 8:17:41 PM
But they have been training on copyrighted data since GPT-2 at least. 2019, and that's when it came out, so before that of course.by spwa4
8/24/2026 at 8:54:38 PM
GPT-2 was trained using data scraped from the web (https://cdn.openai.com/better-language-models/language_model... section 2.1), i.e. copyrighted data provided free of charge to anyone with an internet connection.by yorwba
8/24/2026 at 5:08:26 PM
What tosh. It's copywrited material, they pay to access it like everyone else.by deadbunny
8/24/2026 at 8:11:37 PM
I thought the outcome of that was basically it's legal to train on books, but they acquired the books in the wrong way. If they went out and bought copies of them and trained it would have been fineby IncreasePosts
8/24/2026 at 4:25:27 PM
Of course they are. They have just put on their Swiss Banker suit now and have all sorts of deflection techniques in place such that, of course, "the money has the stamps that says its clean" (when it it really blood money hidden behind a pretty wall).by outside1234
8/24/2026 at 3:45:42 PM
No gain, all liability. Easier to cut them a check for access to training data and say nothing. Unless legal discovery was performed, the outside world would never know, and the payment records would roll off corporate records through a record retention schedule eventually. Could obfuscate it as a contractor consulting fee ("knowledge management subject matter expert") if you wanted to get tricky, depending on the risk appetite of whomever would receive the funds.(not legal advice!)
by toomuchtodo
8/24/2026 at 3:51:53 PM
There’s zero liability in a company stating publicly that they support Anna’s Archive. Zero. Free speech protections cover much more egregious statements than that.by p-e-w
8/24/2026 at 8:55:10 PM
Your freedom of speech is your opponents' lawyers' wet dream. Your publicized support for a known piracy operation will not look very good in the court when you get sued by copyright holders.by pibaker
8/24/2026 at 3:52:38 PM
I disagree. Anyone with even a hint of standing will sue, and keep suing. As someone who has to work with corporate counsel often, do not say anything you don't have to say. Only say what is absolutely necessary. Free speech protects you from your government. It does not shield you from civil suits, and the US is extremely litigious.by toomuchtodo
8/24/2026 at 4:26:01 PM
Maybe they are worried about claims of contributory infringement?by criddell
8/24/2026 at 9:24:01 PM
Those are all law-abiding organizations, which AA is not.INB4: "Here is one time one of those organizations broke the law". Don't go there, absolute lowest level of conversation.
by TZubiri
8/24/2026 at 7:15:35 PM
I'm pretty sure[0] they're all using shadow libraries, and saying things in favor of them would increase their liability.Furthermore, every pirate wants to be an admiral. None of the big tech companies are actually in favor of any amount of copyright reform. They never have been. There is a huge gulf between "personally benefitting from copyright theft" and "actually wants to legalize the theft". Anthropic still believes they deserve to be paid for their models, they just have this delusion in their head that doing a bunch of computation on stolen data is equivalent to actual human creativity.
[0] OpenAI, Anthropic and Facebook have been shown in court to be using shadow libraries, I don't know about Google.
by kmeisthax
8/24/2026 at 4:45:25 PM
You missed Gigapedia (library.nu [1]) , which preceded most of the others.Although I partially applaud what Anna's Archive is doing, I like their model the least (they want to CHARGE to download while doing Copyright infringement?).
In any case, I strongly believe that Anna's Archive is the wrong approach, as it has a single point of failure. We have been doing massive P2P sharing for more than 26 years; we have the algorithms for fully distributed file sharing and databases. Why are we still depending on HTTP/DNS based interfaces that are easily taken down by people wanting to limit knowledge?
I'm glad and thankful that the people behind Anna's Archive dedicate their time maintaining the huge base of human knowledge (Encyclopedia Galactica Asimov would say), but we (the people) should make it really distributed, really infallible and accessible (no, downloading 10TB torrent files doesn't make sense, except for archiving purposes).
We should have something like Popcorn Time but for knowledge.
by xtracto
8/24/2026 at 4:54:28 PM
> Although I partially applaud what Anna's Archive is doing, I like their model the least (they want to CHARGE to download while doing Copyright infringement?).This is bollocks. AA gives users a means of paying to enjoy faster speeds as a means of contributing to costs, but the downloads are free to anyone who doesn't want to pay, and very often quick enough.
by squidbeak
8/24/2026 at 6:16:17 PM
Whilst it pays for the service (and, in that respect, may be a necessary evil), it's morally questionable (at minimum) to charge for things that, by law, aren't yours in the first place.by tentacleuno
8/24/2026 at 6:43:04 PM
Less morally questionable than claiming to users they are "buying" access to media that can be revoked at any point in the future with no recompense, of course referring to Sony and Amazon.by dessimus
8/24/2026 at 6:56:47 PM
> Why are we still depending on HTTP/DNS based interfaces that are easily taken down by people wanting to limit knowledge?Those are arguably the hardest protocols to block on the open Internet without causing major issues for all other sites, forcing those trying to take them down to play "whack-a-mole". If they were to create a new "AATP" for distributing data, it would make it trivial to block on every ISPs firewalls.
by dessimus
8/24/2026 at 6:47:57 PM
BBSes were the first, of course. In particular, Libgen, Sci-Hub and others can be traced back through several generations of libraries to the SU.BOOKS FidoNet echo conference created in the early 90's.by orbital-decay
8/24/2026 at 11:55:48 PM
Yes and I remember a thriving community of IRC DCC file servers. I used them back in the day to read books on my palm pilot before wink existed.by wolvoleo
8/24/2026 at 5:55:45 PM
I wouldn’t include YouTube given how only Google is allowed to index itby mips_avatar