Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

In fairness, I don't think Meta would have (had) any trouble paying the fair price of every book they downloaded (the price of exactly one copy) if that had been possible to do at scale.


Paying the price of one copy does not imply that you can use it for training, right ?


(note; not a lawyer) It depends on if a model is a derivate work from it's source material or not. If yes, then all copyright protections come into force. If not, then the author can't rely on copyright to protect themselves.

My instinct/gut says that an AI model is a derivative work from the training data (in that it quite literally takes training data to produce a new creative output, with the "human addition" being the selection of training data to use), but there's not really clear judgements on it either way for the time being, which leaves room to argue.

The actual methodology used ("isn't an LLM like a computer reading a book for yourself?") is an irrelevant distraction in this regard. Computers aren't people and don't get that sort of protection; they're ultimately tools programmed to do things by humans and as a rule we hold humans responsible when those tools do something bad/wrong. "Computer says no" works on the small scale, but in cases like this, it's not really an adequate defense.


Or rather, that is how it should be; I think the uncomfortable truth here is that we need Congress to make laws to clarify the situation in the favor of society, and Congress does not seem willing to do that.


Doesn't synthetic data complicate this reasoning? If I train a model on synthetic data, which is not protected by copyright, I am free to do as I please. It won't even regurgitate the originals, it will learn the abstractions not memorize the exact expression, because it doesn't see it.

But it's not just supervised training. Maybe a model trained on reasoning traces and RLHF is not a mere derivative of the training set. All recent models are being trained on self generated data produced with reward or preference models.

When a model trains on a piece of text it won't derive gradients from the parts it knows, it will only absorb the novel parts. So what it takes from each example depends on the ordering of training examples. It is a process of diffing between model and text, could be seen as a form of analysis not simple memorization.

Even if it is infringement to train on protected works, the model size is 100x up to 1000x smaller than the training set, it has no space to memorize it.

The larger the training set, the less impact any one work has. It is de minimis use, paradoxically, the more you take the less you imitate.

That should matter when estimating damages.


Understood, Was there any conclusion to the past copy right cases that have been filed against open ai / anthropic ?


All still pending as far as I'm aware. The only concluded lawsuit is that LAION isn't responsible for how AI companies use it's dataset and that merely providing a tagged image index isn't in and of itself copyright infringement (and that lawsuit was ruled in Germany, not the US.)


That's what I do all the time, when I buy a copy of a book and read it.


Since there is no such thing as training rights, they would have a reasonable claim.


I think it is more reasonable for content owners to say what can and cannot be done with their data. After all, content is what make AI possible, and content owners could easily start their own LLM if they wanted to since a lot of it is open source now.


You're taking an "everything not permitted is forbidden" approach, which contradicts the common law principle of residual freedom.

This would automatically outlaw any new use of information (eg music sampling) by default.

If all novel uses were banned from the outset, cultural progress would suffer immeasurably.


I don't think cultural progress will suffer from copyright holders preventing AI from using their content.

What I think will suffer more is the bank accounts of AI corporations.


So to be clear, you're arguing this one specific use (machine learning) should be knocked on the head? And not all novel uses?

Because "content owners to say what can and cannot be done with their data" is quite broad.


No, that's not what I said.

If we want to use data owned by others and make money with it, we can do two things:

(1) just grab the data

(2) ask the content owners

I think what is fair is closer to (2) than to (1). Especially since the data was originally intended for human consumption. What you call "training" is what another person might call "mechanized processing", and would not fall within fair use of the data.


I'm honestly at a loss here. I can't figure out what your position is.

> If we want to use data owned by others and make money with it [...] ask the content owners

So is it "no commercial use without permission" you're arguing for?

> mechanized processing

Or are you arguing that training should fall under the existing mechanical license provisions for songs? I don't think you are, because those licenses are compulsory, and you seem to want an element of choice for the copyright holder.

Ok, put the chatbots aside for the moment. If [brand new use] for a book is invented, and I buy a copy of that book and want to do [that new thing] with it, should the copyright holder of that book be able to block me?


they are not content "owners" though. they have a a copyright that regulates who can copy and distribute that data. they don't have a say how that content is used when acquired legally as long as you activity doesn't constitute a distribution.


That is not reasonable, should a child heed the restrictions placed on the 1st grade math book later in life, when they become PhD?


LLMs aren't people


LLMs aren't the ones making the decision to use the copyrighted information as training data, and it's that decision that is at issue here.


No, but people are the one's training the models.


>I think it is more reasonable for content owners to say what can and cannot be done with their data.

They lose that right as soon as they sell it to other people.

No, you can't sell a book to someone and then sue anyone who reads the book, upside down.

That would be ridiculous. If you don't want someone reading your book upside down, or training on it, then don't sell books.


You assume that "training" and human learning are similar things.

This is a bit like saying that taking a holiday picture of someone, and putting a surveillance camera on the street are the same thing.

I think many books actually prohibit the storage into an information retrieval system and AI can be considered a form of that.


> You assume that "training" and human learning are similar things.

No I don't. Because a human is choosing to enact the training regardless.

Just like if a human held a book up to a rock. It would be ridiculous that an author could ban a human from "training" a rock from a book. Its their book, and they can show it to a rock if they want!


If you buy a DVD and show it at work, then that's also ok, because it is your DVD and you can do with it whatever you want?

Turns out, nope, that's not ok.


> show it at work

No that distribution.

So my point stands. If you sell someone a book, you can't just put arbitrary restrictions on it. You cannot ban someone from training a model on it, nor ban you ban someone freak reading it upside down, or showing it to a rock.

You tried to claim that basically any restriction can be put on how something is used. Thats simply not true. Distribution is a specific carve out that has regulations on it.

But someone absolutely could train a model, as well as read it upside, or show it to a rock.


That is still an open question.


They would have a better defense if they had escrowed that money and/or reasonably tried to buy.


Indeed. Although there is the case of owners of rights invited to come forward to receive their due, if it wasn't possible to contact them before. You probably need a proof that you made an effort though.

It's also true that anyone can go to a public library and read all the contents for free- the point is they can't further distribute them except in a highly processed form (i.e. they can distribute original products influenced by what they have read). Here the issue is the scale of both the "reading" part and of the "producing original work" part.


If anyone could do it at scale, it would be Meta.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: