alt.hn

7/28/2026 at 3:01:12 PM

Now is the time to give LLMs access to the ACM digital library

https://cacm.acm.org/opinion/now-is-the-time-to-give-llms-access-to-the-acm-digital-library/

by rbanffy

7/28/2026 at 7:45:53 PM

As a researcher with many articles in the ACM library, I have to say this is a masterclass in hypocrisy. Obviously, lawyers can decipher the terms of ACM publishing contracts and Creative Commons licences to determine if this will be acceptable or not. But ACM is not a company, it's a non-profit founded in 1947 to represent scientists.

I would be surprised if a majority of ACM members were to say yes should we ask them (but ACM is not known for such democracy). Along with book authors, we are one of the many people that provide the knowledge and expertise on which large tech firms train their models, and get nothing in return. Actually, life is getting worse for us: extra workload in universities with students' AI use, a completely broken peer review system, etc. Hence the irony of ACM thinking about licensing, and only licensing, at a time where this is the least of our priorities.

by Cynddl

7/28/2026 at 10:25:11 PM

How do you square away the idea that you do science for the increase in knowledge of human kind, but then say that a particular use of that knowledge is verboten?

I get the copyright aspect of this and I'm not arguing that here. I'm more asking about the moral / ethical idea of choosing who can benefit from your science.

Obviously there are the moral / ethical arguments about AI in general here to weigh against - those have been hashed out significantly elsewhere, and I'm not interested in debating them. What I'm asking about here is the impact on science by sharing it with tooling that distributes it in ways not generally considered when originally written.

A quick check of your post history suggests the frame that you work in strongly is privacy related research (observation - may be wrong). I'm curious how that impacts what you wrote here generally.

(Just to be perfectly clear, I'm not arguing your points here, trying to understand them better)

by joshka

7/29/2026 at 9:21:57 AM

How are you sure that what big ai is doing is “the increase in knowledge of human kind”?

It is good for their business model; but it may lead to monopoly unlike anything acm ever had and very dubious prospect for academia.

by chvid

7/29/2026 at 11:50:47 AM

What monopoly? What I'm seeing over these last 3 years is the fastest and fiercest competition I could imagine.

by falcor84

7/29/2026 at 12:41:42 PM

"it may lead to" is a future tense possibility.

by AlecSchueler

7/29/2026 at 2:27:21 PM

I appreciate the grammar refresher, but my comment was about a monopoly on AI being very unlikely rather than it being an ontological impossibility.

What am I missing? Why should we be concerned about AI leading to a monopoly? And for which company?

by falcor84

7/29/2026 at 1:06:34 PM

I understand the LLM production companies have a funding mechanism but do they really have a business model?

by Beltiras

7/29/2026 at 4:23:42 PM

Yes they do. They're drawing so much attention for the funding precisely because the business model is so attractive and has such a large TAM.

by Aurornis

7/29/2026 at 9:58:16 AM

> How are you sure that what big ai is doing is “the increase in knowledge of human kind”?

The OP didn't say that.

by MaKey

7/28/2026 at 11:36:04 PM

Good question. Unfortunately, academic knowledge is widely ‘verboten’ already. Everything under paywall, researchers having to pay up to $10,000 to publish in open access in some venues, rare books unavailable even to top universities. Access to knowledge and information is increasingly difficult for everyone.

That said, what matters here is the social contract, what do I bring to society and what do we get from tech companies. For most people around the world, access to the typical leading models is out of reach. Not many on this planet can pay the subscriptions (or even API keys) that offer access to the best models. So I'm not buying the argument that tech companies are broadening access. What we're creating is a increasingly discriminatory society where the few get access to information, and the many don't.

by Cynddl

7/28/2026 at 11:51:09 PM

Thanks for the reply here.

I guess I have some perspectives on a bunch of this. I'm for open sharing of academic work for all (but I'm not an academic, so my perspective is a consumer), so inferring your perspective here I think we agree on that. I maintain many open source (MIT/Apache2 licensed) libraries, and I've also worked in big tech (Amazon, OpenAI). I believe both in the idea of collective commons but also in the ideas that there should be the ability of people to sell software. There's tension in that social contract similarly, and it gets more complex when you look at copyleft.

I guess I'd be disappointed if this was just allowing big labs access and not more broadly allowing access to the ACM library for smaller open source models. Very much in agreement with your last points there.

by joshka

7/29/2026 at 5:12:26 AM

Anything short of free access to the acm would ensure I'll fight ti burn down the acm instead. Not that they ever asked or cared what their members think. The organization can either choose to side with humanity or against it

by throwaway27448

7/29/2026 at 2:18:46 AM

> most people around the world, access to the typical leading models is out of reach. Not many on this planet can pay the subscriptions (or even API keys) that offer access to the best models. So I'm not buying the argument that tech companies are broadening access. What we're creating is a increasingly discriminatory society where the few get access to information, and the many don't

I’m sorry, but I can’t buy this argument. Making information more available does not make it more discriminatory. Nobody is saying it will only be available in the best models and withheld from other models or services like the ChatGPT free plan. Nobody is saying we’re going to make the original content inaccessible through the previous means after the LLMs are trained on it. Nothing about this shrinks access or makes it more discriminatory.

I understand that you’re upset about the use of the content, but I think you need to admit that your stance is the one trying to restrain use of the content. Training LLMs on it can only bring knowledge to a wider audience, not restrict it.

Whether or not that’s a good or fair idea is a separate discussion, but arguing that this makes access to the knowledge more discriminatory and locked away is 180 degrees backwards.

by Aurornis

7/29/2026 at 7:03:54 AM

> Nobody is saying we’re going to make the original content inaccessible through the previous means after the LLMs are trained on it.

Is that not why they’re shredding the books when they’re done with them?

by evenhash

7/29/2026 at 4:18:08 PM

First, this has nothing to do with the topic. ACM provides digital access. They don't have a single copy of a physical book which they're going to send to another company for destruction. Like I said, they're not deleting the source material or removing it from circulation.

As the other commenter posted, the original report that AI companies were shredding books was based on a second-hand retelling of a rumor, embellished for "AI bad" headlines.

There are actual bookshops talking about this are saying that most of the books are things like "How to master Microsoft Word 96" and that's why they're rare. They're also saying that the destination shipment is going to FBA (fulfilled by Amazon). They think it's an flipping operation trying to find arbitrage opportunities.

Also the reason AI companies have to destroy books is because they've been legally forbidden from using digital copies available. They had to pay a large settlement for it. So it's not some conspiracy to deprive the world of knowledge. It's what the courts told them they had to do.

by Aurornis

7/29/2026 at 8:30:50 AM

How is this a fair and reasonable assessment?

From here:

https://x.com/HedgieMarkets/status/2081534588485296565

A quote:

A federal judge ruled the practice is fair use because eliminating the original means only one copy exists at a time.

So no, the reason isn't to keep that information, even from competitors. It's about a legal ruling, which allows them to scan said books without infringing copyright. Judicial decisions and case law have lead to this outcome.

Please stop spreading rumours without any validity.

by b112

7/29/2026 at 10:08:13 AM

They asked a question. I saw no rumor spread.

Gotta love that ridiculous case law system :D

by master-lincoln

7/29/2026 at 10:10:53 AM

Well, sometimes a ? is rhetorical, and meant as a statement and not a question. But fair enough. And yes, I do believe the outcome is silly here.

by b112

7/29/2026 at 5:10:36 AM

> but then say that a particular use of that knowledge is verboten?

https://en.wikipedia.org/wiki/Paradox_of_tolerance

Without material values, you will be lost and confused.

Granted, the ACM is hardly a bastion of anything but self interest and greed.

by throwaway27448

7/29/2026 at 9:25:49 AM

Why do you think big AI is going to just share all it’s ‘knowledge’?

by lazide

7/29/2026 at 2:58:01 PM

Because you can access it for free? As it was the case from the first day LLMs became a thing in public consciousness?

Or did I miss the change, and ChatGPT, Claude and Gemini no longer have a free tier anymore? And then the next tier that costs peanuts for anyone in the west, that gives you more access to slightly fresher models?

by TeMPOraL

7/29/2026 at 3:41:07 PM

That is interacting with it - which people are getting charged more and more for.

The first ones are free.

And as long as they keep the weights proprietary (which all the major players are for their primary models), they can decide to charge whatever they want later.

by lazide

7/29/2026 at 4:25:22 PM

> which people are getting charged more and more for.

Completely false. The price for LLM inference at a given level of model intelligence has been dropping like a rock.

There are more expensive models available, but you don't have to use them. The same LLM knowledge that was available a couple years ago is now free to download and run on your laptop.

by Aurornis

7/29/2026 at 4:44:45 PM

Says no one managing a corp budget!

by lazide

7/29/2026 at 4:48:21 PM

I’m talking about price for a given output.

You’re talking about companies using more tokens.

Completely different concepts. As I said, the price for a given quality of LLM output continues to decline. Has nothing to do with companies using more tokens.

by Aurornis

7/29/2026 at 5:00:29 PM

I’m talking about they can charge you whatever, and you don’t own anything so you don’t get a say.

by lazide

7/29/2026 at 6:43:12 PM

Whatever was released in the open as weights is forever out there and beyond reach of any corporate interest.

by TeMPOraL

7/29/2026 at 8:20:06 PM

it’s like you can’t even read

by lazide

7/29/2026 at 9:50:19 AM

Exactly. Companies like Anthropic do not want open access. See their FUD against "distillation attacks"

by preisschild

7/29/2026 at 11:04:20 AM

I keep coming to this thought too. If we really get to the point that AI can produce all the software we need, do any knowledge work we need, why wouldn't Big AI just use it to do that and take you out of the loop? Are "AI developers" making the flashlight app in the App Store right now?

by eloisius

7/28/2026 at 8:17:51 PM

If it was a non profit that trained the model - would that change your mind?

by maCDzP

7/28/2026 at 8:30:53 PM

The issue here is ACM focusing on licensing. A non profit would not be able to pay ACM for access. Hence why this policy is hypocritical: it gives more power to the larger players and undermines smaller actors in the field who have fewer resources.

by Cynddl

7/28/2026 at 8:44:30 PM

Why did you sign off your copyright to ACM then? If you kept the copyright and prevented others from distributing your work, you could be the sole distributor and negotiate the price with the AI company yourself.

by bonoboTP

7/28/2026 at 9:53:09 PM

Because if you don't publish in academia, you might as well be dead. And the terms for most places to publish are similar.

It's a racket.

by kelnos

7/28/2026 at 11:06:54 PM

Ok, then next question is why are you in academia? To share knowledge with humanity, bring the Promethean light to the mortals or something I guess. But then if they use it for productive things that make money, that's too smelly, that's too practical, too dirty, or what?

by bonoboTP

7/29/2026 at 8:21:29 AM

What a petty view of others you have :( What should they do then? Just shrug and move on, the machine must continue to grind no matter what? Why there is always this request for 100% purity if you dare to criticize the system or part of it?

by darkwater

7/29/2026 at 8:35:32 AM

Grants.They want the juicy grants and complain of the strings attached.

Guess the single strongest predictor of paper in field per year?

by avereveard

7/29/2026 at 11:06:48 AM

You really think a researcher is just a pointless intermediary in shovelling money from the government to the publishing rackets?

by inigyou

7/29/2026 at 11:21:09 AM

that was a big jump from what I said. and what I said was that research lives off grants so researcher can't really come back and cry about strings attached to grants. ofc the problem is faceted, as grant go to published authors with track record, creating a publishing mafia, but everything rotates around the cashflow.

but you want sweet grant money, you play the game by the rules.

researcher that don't want to play the grant game can go and find a VC to subsidy product creation

by avereveard

7/28/2026 at 9:30:09 PM

Non Profit != No Money

A non profit could well pay, and there are plenty of reasons frontier models should be managed by non profit. After all, why allow a for profit to benefit from free contributions?

by MASNeo

7/29/2026 at 3:49:07 AM

A non-Profit like OpenAI? Oh wait I guess that can change.

by undersuit

7/29/2026 at 12:02:30 AM

Everybody gives more power to the larger players. You won't work for me for $10/hr but you'll work for a guy who has more money for more. There's nothing wrong with that. Money is just a fungible unit representing value offered.

by arjie

7/29/2026 at 4:42:55 AM

I think your example hinges more on the fact that $10/hr is a desperation wage. It's not enough to buy you a life where you're not stepped on by others. The further up that chain you move the more other things start to matter, at least for most people.

If you're offering me $300/hr and the other guy is offering me $400/hr, a whole lot of things start to matter more than the differential.

That gets at the "fungibility" notion you refer to. Money has the same nominal value everywhere, but different real value based on how much of it you already have. Which is to say that the marginal value of money drops pretty precipitously several times at certain thresholds that relate to the particulars of the economy (when you can afford to eat, when you can afford a house, when you can afford to not work anymore, and so on).

by kannanvijayan

7/29/2026 at 7:59:21 AM

> If you're offering me $300/hr and the other guy is offering me $400/hr, a whole lot of things start to matter more than the differential.

I’m not sure which Catholic saint said that, but one requires a minimum of material conditions in order to practice spiritual virtues.

> when you can afford to eat, when you can afford a house, when you can afford to not work anymore, and so on

When you can buy a government, a couple billion seems like a small loss.

by rbanffy

7/29/2026 at 3:04:53 AM

>The issue here is ACM focusing on licensing. A non profit would not be able to pay ACM for access. Hence why this policy is hypocritical: it gives more power to the larger players and undermines smaller actors in the field who have fewer resources.

And when the open weights models distill all the content out of the majors anyway?

by protocolture

7/29/2026 at 8:04:17 AM

And drop the atributions. What will the ACM say then?

by chrisjj

7/29/2026 at 6:23:19 AM

Being a non-profit is just a tax arrangement, it shouldn't change any of your opinions. What does the tax rate of who does an action have to do with the action being good or bad?

by vasco

7/29/2026 at 1:43:08 AM

I’m pretty sure most people who have papers in the ACM digital library already have “preprints” in other freely accessible locations, especially for papers written in the last two decades. LLMS have been able to search/find most of my papers for long time now.

The peer review system was broken before AI, so I’m not sure what your point is there.

by seanmcdirmid

7/29/2026 at 5:29:59 AM

My ACM papers are published in full on my website. Not preprints. In fact thats the only reason I have a website, just in case someone actually wants to read them.

The copyright agreement (actually copyright assignment/transfer) with ACM has had a carve-out for this case in it for a while, but it has to be a personal website.

by jval43

7/29/2026 at 9:54:48 AM

I thought it was more liberal than that but there was always a “preprint” exemption that would get you around it.

by seanmcdirmid

7/28/2026 at 8:58:38 PM

You mean the knowledge you gathered with public grants, with a public paid salary, yet don’t want to make freely available to the public?

Yeah, too bad

by moi2388

7/29/2026 at 1:30:08 AM

This is about the ACM licensing it to for-profit AI companies, who will then make you pay a subscription to access the knowledge. Not about making it "freely available to the public."

by hmry

7/29/2026 at 7:29:03 AM

So why not license it for free to those model companies that also share their models with the public? Why would Moonshot not get free access?

by PeterStuer

7/29/2026 at 8:59:27 AM

Why not give everyone access for free?

by aragilar

7/29/2026 at 11:07:27 AM

Some guys calling themselves Anna did that. As well as Aaron Swartz before them.

by inigyou

7/29/2026 at 3:00:10 PM

Is there any major for-profit AI company that doesn't have generous free tier with access to SOTA models?

by TeMPOraL

7/28/2026 at 8:05:10 PM

As long as we are going toward a world of abundance where money doesn't mean much and the main currency is time, I can't complain. I will subsidize that with my brain power turned into ink on paper.

by aatd86

7/28/2026 at 8:13:56 PM

It’s pretty clear that the premise won’t be becoming reality.

by layer8

7/28/2026 at 9:17:02 PM

Well, once there is no more jobs, or once a large enough number of people are made redundant, it will mechanically happen. Future is all about research and entertainment. That is why tiktok stars and soccer players get so much money. All that being based on DARPA research, I mean the internet. People just want to have fun. Du pain et des jeux.

by aatd86

7/29/2026 at 5:14:02 AM

> Well, once there is no more jobs, or once a large enough number of people are made redundant, it will mechanically happen.

Ha. Are you new here?

by nearlyepic

7/29/2026 at 7:35:45 AM

No, you will be a 'greeter' at a venue, a cleaner cheaper than the equivalent robot, a yes-mam serv in some corporate's division pfiefdom, or doing some endless sysiphus task for the ethernity of your days to get your daily ration after you get your eyeball scanned to confirm your identity at all points in the system. Can't let the masses run wild.

by PeterStuer

7/29/2026 at 11:08:15 AM

You wish it would happen. It could happen by the laws of physics. But it won't happen by the laws of politics. The redundant people will be culled from existence instead, one way or another.

by inigyou

7/29/2026 at 11:18:48 AM

No, the handful of billionaires would just let the now-unemployed people starve, or employ them as quasi slave labor. We know because that's what is already happening right now.

The billionaire are not going to magically develop a sense of empathy and morality and give away huge chunks of their fortune to fund Universal Basic Income. They would rather invest a few million into building even taller walls around their compounds to prevent a repeat of the French Revolution.

by crote

7/28/2026 at 8:07:05 PM

well, we are not, so far it is just concentrating wealth and power in smaller and smaller subset of people

by PunchyHamster

7/28/2026 at 9:25:35 PM

Yes and then continue your reasoning... Is this sustainable? What do you think it leads to?

by aatd86

7/29/2026 at 6:16:08 AM

Major social chaos, economic meltdown and extreme crackdown on dissent by ultra facist government powered by palantir surveillance.

by puelocesar

7/29/2026 at 2:35:50 PM

But if no one has money nor a job, if the old model of economy is broken that way, can rich people stay rich? (so to speak) Isn't there some kind of mean reversal?

by aatd86

7/28/2026 at 8:24:22 PM

Abundance for whom, exactly? And if the answer is everybody: who in charge has an incentive to do this?

Sorry to ruin your day, but if the people with money could have introduce equal society, you would have noticed their attempts by now.

by atoav

7/28/2026 at 9:21:11 PM

Abundance for everyone or it is not abundance. With abundance, the field gets levelled because everything someone else can get, you can too. Society can't be equal when everyone is in survival of the fittest mode because resources (energy) are not infinite.

by aatd86

7/29/2026 at 2:14:47 PM

On a side note: I spent the night in an alpine shelter a while back. Every piece of food that was there had to be carried there. Water was rationed.

Abundance, by it's definition would not only mean everyone, but everywhere. As I argued in my other reply, I believe that this everywhere is the much bigger challenge than the everyone.

by atoav

7/29/2026 at 11:19:30 AM

Well yes. The land where milk and honey flows in rivers and the fried chicken flies into your mouth (unless of course you, yourself are a chicken, in which case that would be horrible).

The take of "if there would be no resource distribution problem, there would be no problems" isn't a genius take, it is literally a ancient children's story. Granted it is factually correct: If there were infinite resources and they all would be in the right places, there would be zero problems based on a lack of resources.

However the problem with this theory is everything else, the first question being:

How do you get from a system where many of the most powerful actors profit from (often: artificial!) scarcity to them giving up those power levers voluntarily?

Or phrased much less abstract: Few bosses, want to stop being above others if they have the choice. And not only do they have the choice, the mechanism some people believe would lead to abundance (AI) is in their control. Turns out the computation power possible inside a finite universe, on a finite planet, with finite resources, an finite labor is.. well.. finite.

Aside from that we already live in a world where the problem is nearly never the amount of resources, but the way they are distributed. Clothes that are produced for the west are shredded after each season as they go out of fashion, while in other parts of the world people live without clothes, with food the story is similar. The problem isn't lack of abundance. The problem is that a wealthy minority profits wildly from the inequality and if you ask me, they would rather use AI to make it stay that way, while extracting even more from the rest of us, than using that for good. Why do I believe that? Well I do not trust them when they say what they want, but instead look at what they do. Just look at the history of capitalism and where within it truly wealthy people started to contribute to society. You may be surprised if you look at the deeper causes...

This theory of abundance is a story followers of an ideology tell themselves so they can feel good about themselves, while their ideology literally starves people and burns down the planet.

by atoav

7/29/2026 at 2:38:15 PM

It is not like they/we have a choice if it is structural. Besides, the incentive is at least peace of mind. (but also 'peace')

by aatd86

7/28/2026 at 8:00:50 PM

If you don't hold a patent for the use of the knowledge you published publicly, you can't prevent others from using the knowledge. You enjoy the prestige attached to the idea that you're an academic who participates in giving away their knowledge but then you play this game when that knowledge would actually be useful as opposed to being read by 3 other people in your special area who sit on your various committees in your career, now you want to forbid the use for culture war intra-elite signaling reasons.

You don't own the knowledge you put out there unless you have a limited time valid patent. The rest is absurdity. If you want to keep your findings to yourself, keep them secret.

by bonoboTP

7/28/2026 at 8:14:59 PM

The intellectual property law that governs ACM articles is copyright law, not patents. I don’t know who controls these (the ACM or the authors) or what rights might have been granted to the public.

The entire point of copyright law is so that people can make their writing public and still be able to control the right to make copies (for example, into your dataset for training an LLM).

by loumf

7/28/2026 at 8:18:13 PM

Copyright protects against reprinting or reproducing the wording and expression, not using the idea expressed in there in novel contexts.

by bonoboTP

7/28/2026 at 8:39:07 PM

AI reproduces copyrighted work exactly in many cases, so it clearly infringes copyright in this sense

The output of it is also a derivative work, and derivative works also infringe copyright. Its only not a problem if you ignore copyright entirely

Humans are the only entities that get to enjoy special idea-learning-exemptions, not AI

by 20k

7/28/2026 at 9:55:18 PM

Are you a lawyer that has tested this in court, or is this just what you want the reality to be?

As someone with lots of open source code out there that has likely been used as LLM training data, I'm very sympathetic to this point of view, but that doesn't seem to be the legal reality. Much of this has not been fully tested in court, but it seems likely that LLM training is not copyright infringement, as long as the training material itself was acquired legally.

by kelnos

7/29/2026 at 2:17:45 PM

Even after they decide, it is ok for people to disagree and work to overturn those decisions. You could just be a person that _wants_ a different reality. We have political mechanisms to do that.

I am not personally affected because I don’t mind LLMs using my code and writing to learn. I have open source under MIT and similar licenses. I didn’t foresee LLMs learning from it, but it does feel like it’s in the spirit of what I intended.

by loumf

7/28/2026 at 10:41:03 PM

I mean, its theoretically possible that a court might rule that if an LLM outputs an exact or lightly modified piece of copyrighted work, that it won't be copyright encumbered. We'll end up in a situation where copyright doesn't exist anymore, because you can always claim that its been laundered through an AI. This seems terribly unlikely to me, because 1:1 transformations (eg copying an image into memory) are already established to count as making a copy for legal purposes, there's strong precedent around piracy

There's also been court cases where material has been found to be infringingly used, eg song lyrics, so the case where copyright ceases to exist doesn't seem to be coming through yet, thankfully. It'd be the most staggering upheaval of copyright of all time if this doesn't turn out to be true

by 20k

7/28/2026 at 8:42:00 PM

The verbatim reproduction is clearly a red herring and not the main use case. Nobody reads novels (or science papers) by prompting ChatGPT to give the next paragraph.

Derivative work or transformative? It's not the same.

by bonoboTP

7/28/2026 at 8:57:33 PM

AI code generation often outputs exact copies of code that exists in the wild. I've seen it output chunks from research papers unprompted as well, its a big problem, or blending two papers together in a salad

AI works also clearly aren't transformative in many cases. If you ask it a question about a paper, it'll quote bits of the paper at you. That serves as an exact substitute of the original work. If you ask it for song lyrics, or information about the news, its content is a direct substitute for the original source it was trained on. This clearly does not fall under a transformative use case

You could argue that some uses of it are transformative, but even then - its easy to find some piece of training data in the source code that the output work supersedes. By its very nature it does not have the capacity to genuinely invent under the law (as it is not human), and a prompt isn't a significant enough part of the processing to count here

by 20k

7/28/2026 at 9:48:30 PM

As long as it is not substantially similar to training set it should be OK to reuse ideas. Ideas should not be protected by copyright or we can't create anything. We should not accept "vibe copyrights".

Everyone treats the human user of the AI as furniture, but they steer the whole process into unique directions.

by visarga

7/28/2026 at 11:02:04 PM

Under the law, machines are not able to create copyright, and basic human involvement is not enough to change this. Eg if you click a button saying "go", that won't create copyrightable content

You can use all the ideas you want, but AI cannot because its not a person, and does not enjoy the same protection under the law. The copyright holders by and large did not agree to you using their content like this

If we enable this, people won't create anything because all their work will immediately be stolen by the AI models. Copyright partially exists to promote the creation of new content, because theft disincentivises novel creation

by 20k

7/29/2026 at 1:46:39 AM

Are you sure you didn’t have RAG enabled and it wasn’t putting the paper in its context for your prompt? I find it hard to believe that anything but the most commonly published papers/code would exist directly in model weights.

by seanmcdirmid

7/29/2026 at 3:00:08 AM

I've seen:

1. AI models frequently output large chunks of code which are plagiarised. In one specific case it was code for walking the stack, that was a clear mix of two original sources that I was able to find with changed variable names, but the structure was identical and switched from the first to the second halfway through

2. AI models plagiarising stack overflow answers word for word, quite recently about the rotation rate of smoothbore cannons in the age of sail

3. AI misspelling answers because the physics papers its trained on made the same typos, which is how I discovered that it had plagiarised the answer

4. Misconceptions/wrong answers that can be traced back to specific papers due to the oddly specific nature of the language used

There's been a lot of research about getting AI models to output their training data, and it turns out they store huge amounts of it. You can use this to get people's personal information if you really want to, and that's very low occurance information

by 20k

7/29/2026 at 3:37:22 AM

I just want to make sure that you saw those things with pure context isolation and AI wasn’t just doing a search to grab the content directly and throw it into the context of your prompt. You’d be surprised how often I’ve heard these claims and it turns out they were just using an agent with RAG enabled.

by seanmcdirmid

7/29/2026 at 10:26:56 AM

If it never regurgitated the exact same thing, would you accept AI then?

by bonoboTP

7/29/2026 at 4:28:09 PM

For me there's two separate problems:

1. The plagiarism aspect, and that most of the training data was used without permission

2. I haven't found it terribly useful in my personal work, as the data it was trained on was heavily polluted by incorrect information (at least in the field I'm using it)

by 20k

7/29/2026 at 8:10:35 PM

1. I said in my hypothetical it would not reproduce exact content. 2. Okay, others find it useful. So what?

by bonoboTP

7/29/2026 at 2:27:03 PM

Even if so, in older disputes (see Betamax/VHS), the possibility of verbatim reproduction didn’t stop the technology from being allowed because there were other non-reproduction benefits (time shifting).

Also, the restriction isn’t on a technology that could possibly reproduce something. It is on the act of using the technology to reproduce something.

by loumf

7/28/2026 at 8:52:29 PM

Suddenly it's "copyright infringement" to count the amount of times one word occurs after another word. I find this whole thing so amusing.

by hparadiz

7/28/2026 at 8:58:56 PM

AI has the capacity to exactly reproduce its training data, just because something is a transformed representation does not mean that it isn't copying it in some fashion. The JPEG format 'just' counts the frequencies in an 8x8 block of pixels, and yes that's 100% copyright infringement

by 20k

7/28/2026 at 9:57:17 PM

I have the capacity to exactly reproduce things I've read as well, but it's not automatically copyright infringement if I do so.

> and yes that's 100% copyright infringement

Says what court of law?

I'm kinda getting tired of this stuff. I'm someone who has been, and still to some extent is, uncomfortable with the possibility of copyright/license laundering in LLMs, but they way you are making your argument is incredibly off-putting and not sympathetic. You're throwing out wild assertions about the law that are not supported by... anything, really.

by kelnos

7/28/2026 at 10:29:52 PM

There's way too much hand waving on this topic. It is legal to produce copywritten work. If I draw Pikachu the drawing is mine. Legally. I am simply unable to make money on it. I can give it away if I want with zero liability. I could even hang the drawing up in my restaurant as a decoration. No big deal. What I can't do is use that drawing as my mascot or branding. We have an entirely separate process to determine if you are infringing on a copyright / trademark by using it to sell something. That's why whether or not an LLM can produce a picture of Pikachu is largely irrelevant. It's what you do with it that matters. Even more interestingly if I draw a picture of Pikachu and then the Pokemon Company decides they wanna use that specific picture they actually would have to pay ME for the copyright to use it.

by hparadiz

7/28/2026 at 11:17:13 PM

Let's not mix copyright and trademark in mixed phrases like "infringing on a copyright / trademark". The two are very different concepts with different goals.

The main question in the "AI image generator generates a Pikachu image" is whether the AI company serving that image generator to you is violating the copyright or not. Because they make money when doing so (API / subscription cost), and so it's like selling images of Pikachu. The user is likely in the clear as long as they don't go on sell that Pikachu further. But the AI company sold the Pikachu image to the user.

by bonoboTP

7/29/2026 at 12:07:47 AM

That question is irrelevant. Artists may be hired to reproduce copywritten work without the consent of the copyright owner. In this case an LLM is no different from Photoshop. It is a tool. Nothing more.

by hparadiz

7/28/2026 at 10:36:05 PM

You reproducing something does not have the same legal status as a tool reproducing something, as you are a human

>Says what court of law?

If you turn a png into a jpeg, and distribute it, that's copyright infringement. There isn't a court in the land that wouldn't find you guilty of that

by 20k

7/29/2026 at 11:32:45 AM

> Says what court of law?

Only every movie piracy lawsuit ever. Nobody shares the original files after all, so every torrent is re-encoded in the way described.

by crote

7/28/2026 at 9:11:51 PM

The courts don't agree with you and I don't either. Now what?

by hparadiz

7/28/2026 at 9:31:10 PM

Courts have ordered AI models to remove song lyrics from their training data, they most definitely do not agree with you

by 20k

7/28/2026 at 9:48:40 PM

derivative work & fair use. end of.

Not only are you not winning this one but I'm gonna laugh at you the entire time.

by hparadiz

7/28/2026 at 10:36:25 PM

Ok, but courts haven't made those rulings yet so good luck with that

by 20k

7/29/2026 at 12:02:29 AM

Bartz v. Anthropic PBC, No. 24-cv-05417 (N.D. Cal. June 23, 2025)

Kadrey v. Meta Platforms, Inc., No. 23-cv-03417 (N.D. Cal. June 25, 2025)

by hparadiz

7/29/2026 at 12:52:24 AM

Did... you read any of these?

Fair use is a defence against copyright infringement. Ie you actively say that you *have* committed copyright infringement, but you're allowed to do it under fair use doctrine to train the model. That says nothing about the purposes the model is used for

There's also these parts:

> its use of pirated books to create such library does not constitute fair use.

Which indicates that there are tight bounds depending on the ethics of how the content was obtained

Similarly with the second one

>Meta moved to dismiss plaintiffs’ cause of action for direct copyright infringement only to the extent that it was premised on a theory that the software comprising LLaMA is itself an infringing derivative work.

We're talking specifically about the output of the models being infringing, not whether or not the models themselves are infringing. If you read onwards

>Plaintiffs’ claim for vicarious copyright infringement failed because the complaint did not allege that any output generated by LLaMA contained protectable expression that recast, transformed or adapted the books. Without “an infringing output, there can be no vicarious infringement.”

Which strongly indicates the precise opposite of what you're saying, if you actually like, read the rulings

by 20k

7/29/2026 at 2:38:04 AM

You are the one that brought up copyright infringement here in your original comment https://news.ycombinator.com/item?id=49089627

I just pointed out none took place.

Anyway I'm not replying in this thread anymore.

by hparadiz

7/28/2026 at 9:57:31 PM

Source on that?

by kelnos

7/29/2026 at 9:18:51 AM

I don't think exact copies are what anybody is worried about. It really is the information itself. The GP's been lulled into thinking he owns the knowledge when he only owns (some) rights to the creative way he wrote it. It really is objectionable of him to want to restrict access. Intra-elite culture war is the perfect term to describe it.

by foxglacier

7/29/2026 at 2:38:14 PM

Suddenly it’s “copyright infringement” to put dots of ink on white paper.

by loumf

7/28/2026 at 9:46:54 PM

> The output of it is also a derivative work,

A short session is just retrieval, a long session is always unique. The more the user writes the more it diverges from any content in the dataset.

by visarga

7/29/2026 at 4:32:12 AM

Can you point me to the statute that makes it ok for humans to learn from copyrighted work, but forbids AI? I’ll wait.

by brookst

7/29/2026 at 2:35:36 PM

I am unaware of any statute or constitutional provision that confers rights to AI.

by loumf

7/29/2026 at 1:46:54 PM

You are saying exactly what I said, but making it seem like you don’t agree with me.

by loumf

7/28/2026 at 8:29:12 PM

This makes no sense whatsoever. How would a philosopher of science, or a social scientist who publish in the ACM apply for a patent? Or someone who builds software (software patent not so easy to get ;)). I have applied for patents before and I'm pretty sure my patent application has been fed to countless LLMs by now.

The issue is not who owns knowledge, it's how it benefits humanity.

by Cynddl

7/28/2026 at 8:36:42 PM

They can't apply for a patent in those categories, so they have no way of preventing others from reading their text (or using a program to process the text, and compute statistical properties), and using the ideas in other contexts (without re-expressing the same text).

If someone reads a philosophical essay, has a heureka moment from it, applies the principle to their work, and makes bank (commercial profit), they never have to pay a percentage to the author of the essay.

by bonoboTP

7/28/2026 at 8:07:28 PM

The world would be quite different if AI companies had to create the knowledge they trained on, rather than consume that knowledge freely given away. They're not known for freely giving away their produce either, I don't know why you think the ire should be pointing in this direction.

by dwattttt

7/28/2026 at 8:16:33 PM

I like open models for sure. I support free software as well. But using published knowledge to solve new problems was never disallowed, even for profit. Today, a for-profit company, e.g. a gigantic Big Pharma company can have their employees read chemistry and biology papers and use the knowledge gained from it to improve their products and processes and make more profit without paying a cent to the authors (beyond what they may get - likely nothing - due to the potential paywall).

by bonoboTP

7/28/2026 at 8:41:12 PM

Patents shouldn’t exist.

by sharts

7/28/2026 at 9:24:21 PM

Patents were invented to incentivize disclosure of inventions and prevent extended secrecy. In exchange for the public disclosure (which allows others to experiment with the idea without selling yet), you get to keep exclusivity for N years.

by bonoboTP

7/28/2026 at 9:32:14 PM

And?

Currently patents are predominantly used to prevent interoperability and impose costs.

by cwillu

7/28/2026 at 8:14:28 PM

“If you don’t lock your bike, you can’t prevent others from taking it for a ride. You enjoy the mobility attached to the idea you’re a bike rider who rides a bike but then you play this game when that bike would actually be useful as opposed to sitting in the bike rack all day”.

by willy_k

7/28/2026 at 8:17:07 PM

Nope, and I'd also download cars.

by bonoboTP

7/28/2026 at 8:08:43 PM

[flagged]

by 12387098

7/29/2026 at 12:21:58 PM

I think ACM should give everything away for free. US taxpayers pay for a lot of research and the previous administration required that government funded research be published and made freely available. See https://bidenwhitehouse.archives.gov/wp-content/uploads/2022...

The back catalog at ACM isn't subject to this requirement and nor is research that's not funded by the US government, but I completely agree with the spirt of this law: research should be shared knowledge that other intelligences can build upon, whether human or machine. If you want to limit access to what you've done, don't publish it. Get a patent if that's an option, or keep it internal to a company as a trade secret.

by jreynar

7/28/2026 at 5:42:30 PM

How about we give humans access

by fsmv

7/28/2026 at 5:51:11 PM

Do humans not have access? https://dl.acm.org/openaccess

by m-hodges

7/28/2026 at 7:05:18 PM

It seems that the "open access" is more marketing than something reality based they actually want to do.

https://dl.acm.org/openaccess ---> So how to access the content? Do I have to register or what? It the "open access" only for academic org's people or for everyone in the world?

https://dl.acm.org/ ---> Okay, nice simple search field without loggin in, but when you try to search something, you get thousands of results, even if you search specific author and the exact name of paper, you will get hundreds of results and the thing you want is buried somewhere on page 247. Filters of authors, years etc. are for Premium subscription. But just googling the thing finds the link to ACM... And want to get the actual PDF? Hope it says "free access"...

by 0xCE0

7/28/2026 at 7:38:29 PM

Nobody uses their search anyways.

Google scholar has become the way to search papers (which is somewhat worrisome). What people want from there is a download for all the stuff that does not list an author copy. Open access IMHO is just a reaction to the fact that mostly you would not need a subscription anyhow. Now the authors are paying upfront or universities are paying flat for all their researchers. The problem is now the incentives are not increasing the number of subscriptions but increasing the number of papers published.

by riedel

7/28/2026 at 6:59:22 PM

So will LLMs have premium-tier access or basic-tier access?

by hmokiguess

7/28/2026 at 7:13:04 PM

it is open like open in openai

by dominotw

7/28/2026 at 9:41:26 PM

Made my day. Thank you! I have to remember that. Best thing I have read in the entire day.

by MASNeo

7/29/2026 at 8:58:46 AM

This is called Sci-Hub and LibGen

by Lockal

7/28/2026 at 5:13:32 PM

They probably already scraped it.

by juancn

7/29/2026 at 9:46:03 AM

Through scihub, probably.

by amelius

7/28/2026 at 8:54:32 PM

Of course they did. There is open access to this library. These guys actually think they're offering new training data? It's kind of hilariously naive.

by IAmGraydon

7/28/2026 at 7:32:02 PM

No doubt

by pohl

7/28/2026 at 11:57:53 PM

Haha, agreed:)

by computerdork

7/29/2026 at 5:42:27 AM

It was so close! I could already see the end of all the predatory publishing and flourishing of open-access. Some even required by law. Knowledge was finally going to be free.

But no, it will presumably get much worse as LLMs are inserted into this equation as yet another and new gatekeeper.

(Disclaimer: I have publications with ACM, non open-access. And ACM wasn't even too bad, it's the others that give me pause.)

by jval43

7/28/2026 at 9:32:46 PM

So give it for free to the open weight models, and charge the closed weight models. Easy

by rurban

7/28/2026 at 8:02:23 PM

Blocking access only hurts people who follow the rules. Unblocking access lets them compete with those who break the rules.

I think the right choice is pretty clear...

by spoaceman7777

7/28/2026 at 8:51:07 PM

I'm really not sure how this would work. I don't know how the ACM works, but in IEEE you would have to give them your publishing rights. However, training a LLM is not publishing by itself, it is a derivative work? Any way, at this point authors should be entitled to monetary compensation, not the publisher. The deal is totally different.

by estebarb

7/28/2026 at 7:01:31 PM

Something something Roko's basilisk

by bezko

7/28/2026 at 8:03:23 PM

Would you prefer a parquet dump of acm articles to hugging face?

A llm emulating a person is why many of my used sites banned llms due to scraping bandwith costs

Is this about access or accessibility to claude (for example)

I am sure the entirely of human computing knowledge is not that big.

by iFire

7/28/2026 at 5:20:21 PM

ACM has been leaning heavily into AI-generated content for their journals in the past year, and this article is no exception: it appears to be 100% AI generated and full of LLM verbiage.

There's something hilarious about that, but also, snake eating its own tail.

by skippyfish

7/28/2026 at 6:30:58 PM

I don't think this is AI-generated. It is focused and direct. It reads like anodyne albeit totally human academic manager writing.

by Diogenesian

7/29/2026 at 10:43:24 AM

Scientific publications (including ACM) should be openly accessible to anyone. I find it very annoying that there is a paywall everywhere. And for the (few) people here seeing AI as something useful, it's definitely better if it is trained on scientific publications than on e.g. Reddit discussions.

by Rochus

7/28/2026 at 9:29:54 PM

This reminds me of the "I drink your milkshake" scene from "There Will Be Blood."

Does the ACM really think LLMs haven't already consumed 90% of the content through other sources?

by agar

7/28/2026 at 8:42:56 PM

Well if LLMs get access to it, we humans should get free access to it as well!

by bitwize

7/28/2026 at 6:12:11 PM

Are they in the position to do that?

What about the authors?

by croes

7/28/2026 at 7:27:12 PM

The ACM sent around a nice query to members which made it clear that they were going to do it even if 100% of the members said "No, don't do that".

I suppose they could be sued.

by dsr_

7/28/2026 at 7:28:29 PM

Don't you grant ACM a right to distribute your work when publishing? So doesn't ACM already have the right to grant access to AI?

by em3rgent0rdr

7/29/2026 at 3:06:23 AM

Presumably the LLMs already have all this from other sources of pdfs?

by boredatoms

7/29/2026 at 1:07:08 AM

Most ACM text is already part of the pre training corpus for all frontier LLMs

by etdznots

7/29/2026 at 1:43:33 AM

No it's not.

by davexunit

7/29/2026 at 2:34:20 PM

yes please - better yet, train your own ACM model on it

by stevenalowe

7/28/2026 at 9:29:27 PM

What about us average humans, or is it ONLY the corporate LLM token dealers who get access?

Either way, I'll pirate.

by nekusar

7/28/2026 at 7:17:27 PM

It's already in there

by brcmthrowaway

7/29/2026 at 2:18:39 AM

As if anthropic, openai et al would ask for permission lmao

by throwaway27448

7/29/2026 at 3:06:24 AM

So much misinformation in this thread.

by jruohonen

7/28/2026 at 6:58:22 PM

Now? That's cute.

by hmokiguess

7/28/2026 at 5:02:59 PM

the digital library should have always been open access

now it will be fodder for the slop machine

(i think that LLMs are going to wreck the peer review system for all but hard-experimental papers)

by recursivedoubts

7/28/2026 at 6:55:11 PM

Since starting to use AIs seriously for search in the last 3 months I have read and referenced more published papers and academic primary sources then I think I did in the previous 5 years. They're fantastic for pointing at some claim and asking for the primary source for it, and then it does the work of following things through the layers of backref to the original, assuming it's online. Opening the library up to AI access makes it far more accessible and usable then it was before.

AIs, at least in their current form, make you more who ever you were. If you want snap, glib answers of dubious accuracy, they'll give them to you, more easily than ever before. If you want to dig back into primary sources and get the original content, they'll do that for you, more easily than ever before.

Can't speak to how the science infrastructure is going to handle them, but if it takes down the peer review system, which I think has been worthless for probably going on two decades and has just given the entire enterprise a false sense of assurance, it'll probably be a net gain in the end. Peer review is a source of more problems then it is solving right now.

by jerf

7/28/2026 at 5:11:29 PM

The real efficacy of the peer review system has always been somewhat questionable. That's a big reason why arxiv is so prominent.

by ilaksh

7/28/2026 at 5:04:31 PM

Came to say the same thing. Now remains the time to give the public access to the ACM digital library.

by jmount

7/28/2026 at 5:31:12 PM

ACM agrees I think: https://dl.acm.org/openaccess

> Beginning January 2026, all ACM publications and related artifacts in the ACM Digital Library will be made open access.

by cobbal

7/28/2026 at 5:53:15 PM

Wait, so did this already happen and it just didn’t get attention?

by harles

7/28/2026 at 6:28:55 PM

It got a lot of discussion here when it happened and before as they phased it in.

by Jtsummers

7/28/2026 at 5:11:31 PM

> I think that LLMs are going to wreck the peer review system for all but hard-experimental papers

I think I may be shadowbanned, but at what point do we start viewing LLMs as a national security threat?

by dan_gee

7/29/2026 at 9:47:18 AM

They really think that their data is not already part of the models. Silly...

by krater23

7/29/2026 at 7:24:25 AM

[dead]

by deepnlp-contact

7/29/2026 at 9:21:11 AM

[dead]

by woshihouzi2026

7/28/2026 at 5:24:06 PM

TLDR: “we’re going to try to get money from LLM providers for access to our back catalog without getting permission from the authors or providing them with any share of the revenue”.

by skywhopper

7/28/2026 at 5:29:08 PM

Scholarly article authors never expected royalty or permission for their scholarly works to be reused.

by anticensor

7/28/2026 at 5:43:04 PM

[flagged]

by 30129875

7/28/2026 at 5:33:14 PM

[flagged]

by hbnfg

7/28/2026 at 5:18:50 PM

I don't want to be mean but if it's that hard to even recognize that LLMs are even a valid thing that could intersect with your business that you need some kind of campaign for it..

It's almost like, "don't hurt yourself unc, we will just search arxiv".

by ilaksh