Hacker Newsnew | past | comments | ask | show | jobs | submit | cladopa's commentslogin

AI is not the problem here, but the difference in deep beliefs of the company founders.

When I had partners in my business I studied them first deeply like I was going to marry them because I WAS going to marry them!

You should agree on your values or else the company will go nowhere. One founder will raw one way and the other the opposite making the boat to stay in the same place.

I have a software company and the opposite problem: I am obsessive about testing and meta programming. We write way more testing code than code, we automate everything with meta programming.

It was clear several times that we tested and automated way more than necessary. We automated things that only used once for example(although we used later multiple times on other projects).

So we had not the usual problems than companies have like bloat or security vulnerabilities but we had the problem of having software that was was super expensive because we did double or triple the code than the competition with tests that developers hated(they are boring to write). Our software took also way more time to market.

Companies love to buy the cheapest they can today and tomorrow we will see, so that put us in a difficult position. It was specially hard at first, but then this principles gave us an edge(because most companies would compromise early and not reap the benefits later).

Had I a business partner with a different mindset and beliefs, it would have been impossible to me not to change course of the boat in hard times.

BTW we use LLMs with great success. It let's us test even more spending way less than before.


Thanks for your input. I find testing (especially FE, where a single change can ripple through the whole chain) is actually a great way to use LLMs

It's sad because I've known this guy for 20 years and since the beginning he had a tendency towards "work smarter, not harder", a good motto at the time you're 20 but I often saw it crippling him in the long run, while having short-term gains.

Long-term, funny enough, he said he wanted back into this project because he was very unsatisfied with "easy money" (we both got into SEO at the right time and made a ton of easy money) and wants something stimulating for his brain.

In the meantime I had kids, worked my ass off to get the health and fitness level I want (he also wants it... alas nothing changes).

I'm kind of quite committed to this and we're maybe 30% away from finishing, but I almost debate sometimes if I should just do the FE vibe-coded and let him see that it won't be as good as I can manually do it (FE is hard in the end). Just get it out of the door, keep the customers that we have and let it kind of die down.


I don't know the specific case of Switzerland but I worked in a University in Spain when I was student with servers and stuff and our serious software was never Microsoft's. It was from big Mainframes like IBM or Sun Microsystems servers. We ported that software to Linux and it took very little time, saving us a lot of money with small machines.

That was more than 20 years ago.

Half of our laboratory PC rooms used Linux or BSD and we had licences like Matlab for Linux early on. The other half used Windows and yes, we could get very good license prices from Microsoft precisely because we had so many Linux machines. It was almost free.

You could even say that Microsoft spent money on us trying to buy the University sales gatekeepers with things like free expensive travel to States' conferences, free laptops and stuff.

We were probably not a normal University as we were a technical University and regularly organised conferences inviting people like Stallman.(Our technical University also managed the servers of the Humanities side of the University).

It was not that hard. Today, with so many programs running in Webassembly, it should be much easier.

>Is it just political symbolic migration because they have anything better to do with tax payers money?

Usually when you replace privative software with open source you spend 10x or so less in licenses. The problem being that the open source software is not as capable as the privative.

I did not understood why instead of spending 10 times less with open source, you spend half and pay for the necessary capabilities improving the open source software.

I understood it later. In the past it was a scheme lots of politicians used for getting negotiation power against Microsoft. Digital sovereignty was very down the priorities list compared to getting a good price.

Now, after the Ukraine war, and Trump actively using the US dependence against Europe, digital sovereignty is a very real issue, even critical.


So you believe that writers are not making the necessary effort and write using LLMs, so you then use an automatic LLM to filter it as a reader(because you don't care as a reader and don't want to make the effort manually).

So this way there will be a lot of false positives like with school assignments.

I think a better solution would be to have a network of people that you trust manually read and label texts instead, so this way no machines are used and you don't need to read a text that 100 of your trusted friend/trusted readers flagged as artifitial.

By the way, this man does not care about AI slop. He cares about AI usage. In the same way there is very good code created with the help of LLMs, albeit a minority like Linus Towards says, there will be very good writing created with the help of LLMs.


The [content-based] trust layer of the internet is what we have deferred to this point, and what needs to be created (which is easier said than done, of course).

Personally I think a web-of-trust Keybase-style solution could work... But buy-in is difficult... you'd need some sort of "seed" strategy to make the app useful alack of 100 (a huge number) trusted friends.


It is not a big deal. Since the invention of the printing press any important book has been duplicated by thousands, tens of thousands or even million of units.

Just taking one of those and "destroying them"(it is not destroyed, a digital copy with the ability of doing millions of copies is stored somewhere) is not problematic for Humanity.

By the way, I always search for second hand books. Most of the books there are garbage. Most people clean their shelves with the books they don't care about, but preserve the ones that are good. If they are young people that inherited a house and don't care about books, they pick and sell the good ones, giving away the bad books.

If you go to a recycling centre, the garbage to quality ratio is over 100 or more. That is, for every 100 books that are garbage there is one good quality book. It is very rare to find a jewel there.


> a digital copy with the ability of doing millions of copies is stored somewhere

somewhere were we can't access it. as the article states: “permanently locking human knowledge inside private corporate servers”

the issue isn't that the physical copy is gone, it's that they are preventing people from making digital copies that are actually accessible by destroying the physical copies.

> If you go to a recycling centre, the garbage to quality ratio is over 100 or more.

archivists keep everything, because we don't know right now what will be important 100 years from now.


> somewhere were we can't access it

By any metric imaginable, it's making the information more accessible, not less. First, it's taking a single copy of a 10k physical print and it's making it digital. Is it "locked"? Yes, by copyright laws, if you don't like that lobby to have them changed. But it's _closer_ to being widely available, not farther.

Plus having the info part of a LLM makes it immediately available to literally billions.

I happen to actually actively shop second hand bookstores, so I am potentially affected by this - as opposed to most people complaining because they don't like the idea. And I still absolutely support it.


>Plus having the info part of a LLM makes it immediately available to literally billions.

Help me understand how. Not only are these LLMs expressly prohibited from specifically regurgitating copyright works if the users asks them to, but they habitually hallucinate or paraphrase things wrong.

If they won't regurgitate the copyrighted text verbatim, and are known to be confidently incorrect and hallucinatory, I'm struggling to see how these texts are "immediately available to literally billions".

And I ask this as someone who has had LLMs give me incorrect assertions about the contents of books.


Depends upon what you want.

For example, knowledge about how the book smells when you open it, something that reader do talk about enough I don't think this should be a strawman, is lost. But, that is about the experience of reading the book, not the knowledge of the book.

The exact text? Yeah, I think that is largely lost as well. This is a summary. And for rarer books, it will be a particularly bad summary. The basics of the book are being captured in a space that the right question that retrieve it, but worse than a sparksnote and any well read reader will tell you all the sorts of things a sparknotes already loses compared to reading the book directly.

But, that little bit of data is a bit more data than existed before, and future LLMs should get better at giving the information. So, in that sense, the knowledge is better being spread compared to copyright where the book stays in a warehouse until it is disposed of. If it was between this and a sparksnote of the book being made, the sparksnote is far better, but between this and the book simply being disposed of, then the LLM is better but far, far from great.

That's a lot of assumptions that goes into the judgment, which is probably why different people reach different conclusions. One person is imagine the alternate fate of this book being slowly rotting in a landfill, the other resting on a bookshelf where it is read at least once fully and then flipped through time to time, and neither are wrong.


> But, that little bit of data is a bit more data than existed before,

No, it’s not. Undigitized data is still data. Wording, style, nuance, type, artwork, binding, metadata, attributions, citability… all of these things are permanently lost, which is a fucking tragedy because they don’t have to be. Even if you’re not willing to take the time to scan every page in a cradle scanner, as rare books should be, (and don’t tell me they don’t have the money to get a handful of library interns to do this,) you can disbind the books and store them as they did in the Caselaw Access Project at Harvard Law. They removed the pages from the binding, scanned them on a high speed conveyor belt scanner which yielded full color 600 DPI jp2 images, placed the pages back in the binding like a folio that could be re-bound if needed, vacuum sealed them, and stored them in a salt mine. It’s not like it was slow, either — we did 40k in 18 months and we did take the time to scan the rare ones with a cradle scanner. And we did it all in less than open AI probably spends in a day on inference.

> So, in that sense, the knowledge is better being spread compared to copyright where the book stays in a warehouse until it is disposed of.

That’s a false dichotomy. Libraries exist for this exact reason, and their not already having a copy does not make “you snooze you lose” a morally acceptable strategy.

I’ve been pretty cool on the direction of SV for the past decade at least, but I am absolutely gobsmacked by the unbridled hubris of these companies over the past 5 years.


>all of these things are permanently lost

A small fraction of them is saved in the model. Far more is saved in the digitized copy as long as they keep it which they have plenty of incentives to do so (future training of newer models).

That's more than what happens if that book was burned or sent to a landfill, but less than if the book is giving a loving home.

>They removed the pages from the binding, scanned them on a high speed conveyor belt scanner which yielded full color 600 DPI jp2 images, placed the pages back in the binding like a folio that could be re-bound if needed, vacuum sealed them, and stored them in a salt mine.

My understanding is that this simply isn't legally allowed for these books. The original must be destroyed for the digital copy to not be copyright infringement.

>That’s a false dichotomy.

I pointed out there is a spread of possible outcomes and that different people are considering different outcomes and the comparison of if this is good or bad depends upon which outcome one considers. I even mention that both outcomes are sometimes right. That's about as far from a false dichotomy as I can see it.


> Undigitized data is still data. Wording, style, nuance, type, artwork, binding, metadata, attributions, citability… all of these things are permanently lost

I understand what you mean but... "permanently lost" sounds dramatic. When I trow away old pictures, old drawings or pieces made by my son at school, they are lost as well. Not that I do that often, but it begs the question, should every 'ip' made by humans be preserved?


It sounds dramatic because it’s dramatic. You’re combining two things — whether something is, in fact, completely lost, and if it is something that should be kept. Something that should not be kept is still completely lost if it’s destroyed. It’s difficult to imagine they’d digitize it if it was worthless.


> Is it "locked"? Yes, by copyright laws, if you don't like that lobby to have them changed.

I don't if that is true: a lot of old books might still have copy or other rights associated to them, likely owned by author and/or publisher, directly or inherited, but often those who have the rights do not have digital or physical copies at hand anymore (some old books are, well, really old). Does Anthropic make sure to track down, contact and then share the digital copy they make with those who have rights on the work? If not, they are not making it in any way easier to re-print the books, while making their supply more scarce (they destroy existing embodiments).


> Is it "locked"? Yes, by copyright laws, if you don't like that lobby to have them changed. Or be useful and scan book for Anna archive and other shadow libraries.

Citizen's lobbying against megacorp is a mirage.

>But it's _closer_ to being widely available, not farther.

By what metric ? The copy is now guarded by a company instead of being on the second hand market.


> Is it "locked"? Yes, by copyright laws,

> Plus having the info part of a LLM makes it immediately available to literally billions.

Isn't this a contradiction? I mean, maybe you can invoke a fair use policy if an LLM spits out some text from the scanned book, but then you are _not_ "making it available to literally billions".


You really drank all the Koolaid they had to offer.

Destroying the last copy of a book is OK because you shop for second hand books, and because some private company hold the last digital copy and have no incentive to make sure it survives. Damn..


> permanently locking human knowledge inside private corporate servers

History tells us that very few "permanent" situations are truly permanent.

Provided a set of information has value (which in this case it clearly does) then the overwhelming likelihood is that, eventually, through some method or other, the information will become public.


> History tells us that very few "permanent" situations are truly permanent.

If you destroy the only copy of a physical artifact, the situation is as permanent as it can get.


I mean books get destroyed all the time, really they are a major pain in the ass to keep together, especially as they age. Paper loves to crumble. Insects think they are tasty. Floods and fire love destroying them too.

So physical books are rather non-permanent themselves.


Exactly. Calling anything on an SSD permanent is criminal.


I wonder if Anthropic could rent or sell access to their collection to the internet archive? It would probably be a good PR move (which they probably need right now), but I'm not sure what type of legal shenanigans they would need to do in order to not violate copyright law.


> any important book has been duplicated by thousands, tens of thousands or even million of units.

Books in the former USSR display their print runs on the last page. "Important book" is a vague and arbitrary term, but rhere are works in whole fields (e.g. history, archaeology, linguistics, ethography) that any scholar would consider key references, and as few as 100 copies were printed.

The shadow libraries have made a lot available to the whole world. It would suck if private corporations scan and shred remaining copies of these before the shadow libraries can get a scan.


Our little project endeavours to digitise books from the erstwhile Soviet state which were published in many languages

We have, over the last 15 years, acquired/borrowed from libraries, and scanned a couple of thousand books on all topics of interest. All of them are out of print. Some of the physical copies were have, especially in indicate languages might be some of the few surviving ones

https://mirtitles.org/

https://archive.org/details/@mir-titles


I think you may be greatly underestimating the long tail. Several times a year I read sources that reference older books that I can't find online. When I am able to locate them, they often cost at least several hundred dollars, sometimes into the 10s of thousands.


Do you think that you are somehow special and unique in needing these books? If the books cost in the 10s of thousands, then there's obviously value to other people. I've seen no evidence that Amazon or others are buying $10k books to scan into their corpus. All evidence I've seen is that they're scanning cheap books with no current value and no clear use to people today.

I have volunteered with a library, and probably threw hundreds of books over a couple week engagement from a university library into a shredder at the direction of a professional, academic librarian. Libraries are constantly culling books, the EXACT category books we're talking about here (old, never-read). This is happening at a much larger scale, so I would recommend railing against university librarians in addition to the AI juggernauts.


I'm just stating that not everyone is reading the most popular million books. And, there are many, many millions of books in the long tail.

A few years ago, after running into this issue several times, I looked into the economics of it to see if there was a opportunity to republish digitally. I found the sales data for some, the ones that had value were selling on average for around $50-150 dollars. Because the typefaces weren't modern, OCR wasn't scalable. Because only 20-40 were transacted each year, it wasn't worth the time.

What kind of books are they? In my case mostly historical documents by some relatively unimportant person who was highly important for a very, very niche subject. They still contain valuable, irreplacable information.

An analogy: imagine you discover a really cool video game from 20 years ago. You love it and want to find the developer's previous work. But, you find out that the company was purchase by another company which was purchase by another, etc, etc. Sure, maybe the game still exists in some digital vault. But, more than likely it's gone because old things only seem have value these days if some influencer broadcasts it.

It's interesting that libraries are purging them. Several times, I've found that the only available copies were at some random university rare book collection in middle of no where, and basically impossible to access unless one fly in - which isn't worth it. They are often donated by a benefactor and stuck there - which is why they're never read. It's not that the content isn't valuable.


I think I get your point, and I’m sure there are ideas that will disappear forever.

The books we were disposing of were so totally inane that I couldn’t tell you what they were, but they weren’t taken out in like 30-40 years. Northeast US state R1 research university. Maybe we got rid of stuff that would come in handy!


> I've seen no evidence that Amazon or others are buying $10k books to scan into their corpus.

Have you seen evidence that they’re buying only readily available books that are plentiful on the market?


I have books, but they are just objects. They're nice objects, but just objects.

Fetishizing books isn't going to help.

In fact, most of those "rare" books don't sell because nobody wants them. The AI companies are making them even more rare, so the booksellers should be thankful.


Exactly, also all those physical books would anyways get molded, eaten by moths or just naturally decay. It's not like the AI companies are obliterating every copy of every single book.


We should actually thank the AI companies for helping us get rid of our trash!

Thank you anthropic! Thank you OpenAI! Thank you Google!


Somehow I don't think they're looking for the books that have been copied over and over.


Why do you think that? Do you have evidence, or is this just a hunch? Why would Amazon waste money on uber rare books when there are thousands and thousands of not-so-rare books that could serve the exact same purpose?


For one they've already used the entire library genesis. Anything not in there is going to be obscure in some capacity.


Whenever my wife wants to visit antique stores, I always look for old books. I have found several 100+ year old gems.


>Since the invention of the printing press any important book has been duplicated by thousands, tens of thousands or even million of units.

No, every popular book has been duplicated thousands of times. This is not the same thing as important. They're orthogonal. When an important book is popular, it is safe. When it is not, it is in danger. Only fools assume that important books are recognized often enough to become popular.


> any important book has been duplicated by thousands, tens of thousands or even million of units.

It is common for academic books to have publication runs in the low three digits.

You may argue these books are not important. But how do we know if we fail to preserve it?


Is there evidence that these mythical low-print-run books are being purchased by Amazon and friends for destructive scanning?

I simply don't see why railing against the AI giants about this without also contributing significant effort into protecting these low-print-run books from getting jettisoned by university libraries makes any sense.


> Is there evidence that these mythical low-print-run books are being purchased by Amazon and friends for destructive scanning?

Is there evidence that they aren’t? All of the common works are already on libgen, if these books are worthless repetition they will contribute little to the training corpus, labs want high quality interesting texts and they have unlimited budget to spend on it. Paying $300 for something rare with millions of tokens of interesting and original text for training is definitely fucking worth it, spending $1,200 to get all the copies and block your competitors from getting it is most definitely worth it!

> I simply don't see why railing against the AI giants about this without also contributing significant effort into protecting these low-print-run books from getting jettisoned by university libraries makes any sense.

Do you think it would be OK to rail against for example, a genocide if I wasn’t substantially contributing to some effort to stop it? This sophistic (and uninteresting) bit of rhetoric boils down to: if youre not trying to fix it yourself, dont complain!

(reminder that one of the basic principles of a democratic society is that each person has some concern outside of their personal affairs)


we're done - straight to genocide, later buddy.


With all due respect, I would say it wouldn’t hurt not to downplay this.


Sure but are we going to do anything sensible about it or is it just a convenient culture war vector?

This is happening because copyright means you can't scan these without destroying them as a format conversion.

No one's felt compelled to try and fix that so we can do this sort of digital archival and preservation, and copyright allows works to be frozen and undistributable because a claim might exist for decades without any actual use (I.e. the number of games which get stuck in legal limbo).

If the only desire is to sling mud at AI companies but not try and improve the legal situation, then it's worse then useless because there's no intent to stop it - in fact stopping it would remove a useful outrage tool.

The idea in the title here is point 1 of the blind leading the blind: it's illegal to scan and store these books without destroying them in many jurisdictions.


Beware that the notion of "quality" is entirely different for AI companies: they don't seek entertainment, but sentences in a language to train an LLM.


I don't know why you think all these irrelevant remarks somehow support your core thesis.


Honestly, it’s easier to find a good movie than a good book, because books are way cheaper to publish. These days, the quality of pretty much all kinds of content has become a problem...


>It is not a big deal

have you ever held and read an old book?


Read more about this. Depends on the definition of the "big deal" but from what I can understand the problem is that they buy rare things - which exist in just several copies - and they tend to buy _all_ copies.


I was also skeptical of this claim but managed to find someone who explains it:

https://downtownbrown.substack.com/p/five-fallacies-ai-and-d...

It's not that individual companies buy all copies of a given book, but that there's more than one book scanning company, and they aren't sharing the scans with each other. The result: books that were rare but nevertheless easy to find for purchase (thanks to the internet) are now vanishing off of the market, becoming de facto no longer accessible to the public.


Good article, but I still feel unsatisfied because even it cannot find an example of a book that’s actually been lost because of the destructive scanning frenzy (it only lists books that hypothetically could be lost because there’s not many physical copies available for sale online.).

If anyone has an example, I’d love to hear it.


Since we don't know what was actually purchased and what was actually destroyed, how do you expect us to furnish an example? This would require the destroyers to admit it, and beyond that it would require all of them to admit it since more than one of them might have been responsible for the extinction of one text. Seeing as they were already keeping this operation under wraps, I don't see that happening. "possibly extinct because no copies available online" is probably the best we can do.

The distributed nature of the problem and the utter lack of transparency are huge factors here too.


These companies are buying from small book stores; surely these collector types would know if they lost any one-of-a-kinds? Or at least someone would be keeping track of extremely rare books disappearing (especially now since this matter has been public for weeks).

Regardless I feel like they gotta figure this out for optics reasons. “So and so books are lost forever to Anthropic’s servers” is much more outrageous than “Anthropic is destroying a bunch of books that have other copies” imo.


Most old books that are rare and unpreserved are so because their value is marginal, so nobody has bothered to collect and preserve them.

But where did you hear that they’re buying “all copies”? And to what end?


This is a classic mistake. We have no way of estimating the future value of a given book. It's perceived current value (a large part of which is simply obscurity) may be low. But it's future value - to historians, ethnographers, to researchers seeking a specific fact or example of language use or a hundred other things - is literally inestimable.

To take a crude example in a different medium - new york in the 90s - widely documented right? Yet, if you want to find high definition video of street life in a given burrough on a given day or year, you're faced with an enormously difficult task. There were some HD test videos done in Manhattan in the late 90s (which have been posted to Hackernews before), but there's no equivalent for the other burroughs. Your best best would be finding original negative out takes or location scouting footage from feature films, a very hard task. That's only 30 years ago. Outside of the focal points of the worlds attention - English language, rich countries, places in the news, contemporaneous sources for 'non notable' events (lifestyle, how people spoke dressed etc) is surprisingly poorly preserved.

Hopefully you can infer how this tracks to the written world and primary sources for language, technical manuals etc etc.


>Hopefully you can infer how this tracks to the written world and primary sources for language, technical manuals etc etc.

I honestly can't and I think you can't either or you would have used an example with books/printed media rather than film, an entirely different ballgame.


OK... I'm going to assume good faith even though your wording makes it somewhat unlikely.

Similar textual examples would be any text containing actual language as it's spoken in a given place or time. Or any factual textbook detailing the buildings present in a given location. Or any text book detailing a now defunct construction process. Or any text book (generally small run) detailing a niche interest, now missing ecosystem or the state of a particular political situation at a given time. Essentially all textual primary sources for events which are not currently considered important - but which we have no way of estimating the future importance of. One can continue to create countless counterfactual examples in this vein. My overall point is we cannot know what may be useful or even essential in the future, and knowledge should be preserved under the assumption that it is likely to be.

Historians frequently refer to this paradox - how everyday aspects of life are frequently not explicitly documented, since they're so obvious to the communities or communities of expertise that observe and carry them out. So it's actually incredibly important to preserve what seems like ephemera.

Hell we couldn't have AI training at all if we lacked the corpus of existing written literature - but there was no way any author could have anticipated this future utility more than a couple of decades ago.


I think there are two issues here:

1. If someone is acquiring books in bulk for bargain-bin prices and shredding them, they're books whose physical copies have essentially no market value and which, absent this buyer, were overwhelmingly headed for pulping or landfill anyway. Millions of books are destroyed every day.

Could one of these worthless looking books turn out to contain information historians care about in a 100 year? Sure. But that doesn't create an obligation for someone to pay to warehouse every extant copy forever. Physical Preservation has costs: space, cataloguing, handling, transportation etc. Archives and libraries have always had to make choices for this reason.

2. I'm not arguing that preservation has no value. The question is whether destroying a physical copy after digitizing it is a serious loss when talking about mass-produced printed material.

Was this the last surviving copy ? Is the information unavavilable in libraries, archives, other editions, scans, citations, contemporary works etc ? If not, nothing has been lost except one physical instance of a reproducible object.

Your film analogy worked a lot better because old film footage is often unique primary source material. A camera recording of a random brooklyn street in 1993 may literally be the only recording of those people, storefronts and circumstances. The nth printe dcopy of a technical manual is not analogous to that.


How would film be any different in any capacity whatsoever?

Believe it or not, what matters here is the message and access to the message, not the medium.


The medium obviously matters. One was designed to be cheap and mass produced. Almost anything published has copies in the hundreds or thousands at least and the chances you are concerning yourself with the 'last copy' of some valuable information you can't get elsewhere is minuscle. Film most certainly was not like that.


Assuming these texts exist in large quantities seems to definitionally contradict the term *rare* book. Plenty of books have small print runs or have reached out of print status.

One could obviously reproduce films too.

Ultimately, you don't know how big N is, and that's what matters. Part of the problem is there is no transparency around this. The population of all books != the population of a specific book or set of books. Assuming N is large when you have basically zero information on specifics is foolish. These books might be just as rare as a film for which only one source exists you don't know.


>Assuming these texts exist in large quantities seems to definitionally contradict the term rare book. Plenty of books have small print runs or have reached out of print status.

They're books that are no longer in print. Doesn't mean there aren't a lot of copies around.

>Ultimately, you don't know how big N is, and that's what matters. Part of the problem is there is no transparency around this. The population of all books != the population of a specific book or set of books. Assuming N is large when you have basically zero information on specifics is foolish. These books might be just as rare as a film for which only one source exists you don't know.

All of this is frankly irrelevant. Books that you buy in bulk at barging bin prices are books that have essentially no market value and were going to the pulp or landfill anyway. Nobody has an obligation to spend money to preserve every extant copy of every book forever. Millions of books are destroyed everyday. Even libraries and archives make these choices. I was simply pointing out the film analogy worked better, but it doesn't change anything.


Again, you're assuming a lot about this situation that we don't actually know to be true. For example, that these purchases are limited to "bargain bin" books that have essentially no market value. These may be cases we know about, that doesn't mean these are the only purchase happening. I would hypothesize that if we know about a few specific purchases there are probably many more we do not know about/haven't traced back to these companies.

You realize "market value" isn't the only kind of value too, right? Even books that are completely obsolete can have inherent historical value insofar as they enrich our understanding of the past and how human beings used to live. Contemporary market value isn't everything. There are plenty of works of scholarly interest that wouldn't command nearly requisite value to a NYT bestseller when you adjust for scale. That doesn't mean these books aren't worth keeping or aren't intellectually important.

Also, these books presumably even have some perceived value, otherwise these companies wouldn't think they were relevant for training, and wouldn't be spending money to purchase and scan them in the first place. The whole premise is that these books are valuable enough such that training on them will make the models valuable and make paying customers line up for tokens. This is an inherent imputation of value to these works on the part of these companies. You can't say, simultaneously "these books are worthless" yet somehow, at the same time, "they are worth something for producing artificial intelligence". Your assumptions don't make any sense.

> They're books that are no longer in print. Doesn't mean there aren't a lot of copies around.

Sure, but there also could be limited copies around. We don't know. We won't know unless these companies are more transparent about exactly what they are buying and scanning. It's that simple. If it's foolish to assume there might be limited copies, it's equally foolish to assume they have plentiful copies. We simply don't know.

> Nobody has an obligation to spend money to preserve every extant copy of every book forever.

I don't think anyone is claiming that. People are bothered because they (in my mind, rightly) recognize that old books are a historical record of human accomplishment and have some amount of cultural value. Letting a handful of companies further plunder the collective output of humanity and now, potentially keep it locked on their servers in perpetuity is something that should make you upset if you have any inkling of interest in the history of humanity or human accomplishment and if you have any sliver of curiosity about these things. If all you care about is the present and getting better LLMs, sure, I guess I understand why you don't care, but then I also think you are very unwise and myopic.

> Millions of books are destroyed everyday.

So what? Millions of crimes are committed every day, that doesn't mean we shouldn't be bothered by the crimes we hear about. Billions of pollutants are spewed into the environment every day. That doesn't mean we shouldn't try to combat pollution. That's a pathetic attempt at justification. Also, there are clear distinctions between processes like libraries removing works they can no longer store and companies scooping up books for training and personal gain.


Here we have another entry in the long list of "things described on the Internet that never happened".


Where did you see they tend to buy all the copies? This comment is the first time I've heard of this.


I'd also be curious about the provenance of that statement. Why would they buy all copies? What would be the purpose of scanning multiple copies?


I could see buying multiple copies being useful to mitigate problems, like damage.

But buying all the copies is categorically different and I cannot imagine why they would attempt that, except perhaps to prevent their competition from getting a hold of the same text?


It’s just another lie of the type these threads tend to be filled with nowadays.

Of course they aren’t buying “all copies”, and that wouldn’t even be possible in most cases since such books are usually flea market/attic material and most copies aren’t for sale (or even catalogued) to begin with.

I’d be interested to learn who comes up with such lies though. Is it really just random people venting their frustration, or some kind of organized astroturfing operation?


It really does seem like an organized operation, given that it's so prevalent and they all seem to be in lockstep with their specious claims.


I wouldn't be surprised to find out that this outrage is fueled by the same actors as the AI data center water outrage. Whoever they are


I think these type of lies usually come about as a result of a game of broken telephone and things get exaggerated


what is your source?


First of all, it IS destroyed and it is a big deal. Hardcover copies of books especially 1st - 2nd edition ones (even with mistakes) are rarer than digital scans.

Maybe the Bodleian Library at Oxford University should give all their rare books to AI companies to scan and destroy them since it is not a "big deal" anyway.

Except that when they did do a pilot with OpenAI to scan these rare books, [0] they did NOT destroy them. I wonder why?

[0] https://www.bodleian.ox.ac.uk/services/research-partnerships...


Why wonder? The answer is abundantly clear if you follow the news. If a book is copyrighted under U.S. law, scanning and destroying counts as a format conversion which qualifies it as fair use, so there is no need to negotiate with copyright holders. See Judge William Alsup’s decision. If Anthropic did not destroy the books after scanning it would have not won the lawsuit, and scanning would be illegal. If a book is already out of copyright then of course they do not have to destroy it afterwards.


Why is it a big deal?

Do you think there's no difference between the books curated at Oxford University and the crap that Amazon is buying?


Why would amazon buy "crap"? Surely they want their model to succeed and they want to train it on valuable input, no? They have more than enough resources to determine whether or not the books are worth buying. They have been in book selling for a long time.


I have released some of my projects as Open Source, I also have a company with privative software.

Claude and AI partners have taken all what they could from the Open Source projects without giving credit or respecting the licenses. They have increased the traffic on websites in an absolute disrespectful way increasing the hosting cost in inefficient and ridiculous ways.

They have taken all the important books and not asked permission from the authors.

Fair enough. Fair use.

Of course I would create a competitor software to Claude or any others if I could. Using Claude(and others) of course.

I am not paying you 200 dollars/month for you to tell me that I could not create code that competes with you. If you try to go to court in Europe with this you will lose.

It is just the same fair use you proclaim for taking the data from others.


They would likely lose in the US too. But until that happens, they can keep living in their fantasy world...


It is not non sense. I am spaniard myself and I did not understood what he was talking about until I saw the correction(is this a club name or something?).

If I want to talk about "Seattle" and use "Siadol" a lot of people are not going to understand.


"Saragossa" is an old and well-established spelling of that city's name in English and this is an English-speaking forum


A name is what something is called. For people it's rude to call them by a name they didn't choose but for non-persons that doesn't matter, if everybody else wants to call it "Siadol" then that is its name.


Every language has his own name for tons of places. In Spanish it's Zaragoza and in English "Saragossa". Ditto with Londres/London.


At least you learned something new today


As someone who has helped drug addicts in an organisation called "Proyecto Hombre" in Spain you probably don't understand how hard is to help someone in this situation and that sometimes you CAN'T.

It is very typical that when someone gets on drugs even the father of the person stops trying to help after several failed trials. Only the mothers stay. And it is so painful to watch. The kind of abuses mothers tolerate from their children.

Substances destroys the typical regulating behaviour of the brain. They literally can't see the harm they do to others. You usually can't help them until they hit rock bottom and open a window for you to intervene.

For helping you need knowledge and training and "trying to help" without it could do more harm than good.

About the moral obligation to help the drug addicts, as "social justice", I would call this "injustice". Here is the typical situation: Someone takes drugs in a social group because is "cool". In the same group there is someone that decides not to take them. He is humiliated by the group and it takes a lot of work. One of the members of the group gets addicted, do terrible things and asks for help.

It is usual that the one that did not take drugs, or the brother in the family ask the question "why are you helping this person that decided to go this way herself?" "Why are you expending like 10x more resources in this person that did everything wrong than in me when did everything well (studied instead of earning easy money selling drugs, that had to withstand the pressure and humiliations by the group).

Drugs are super expensive and addicts will do anything to get their dose, supporting the addicts usually takes all the resources from a family, neglecting the resources to other members of the family. If the State provides the support, the same thing happens. Should the State spend exorbitant resources that the rest of the society has to pay taking the resources from everyone that needs them?


Looks great technically. I am a language nerd as I have worked in multiple countries like China or Austria or the US and I am from Spain.

If you want to sell your app you will need to explain it for humans, like this guy is doing: https://www.youtube.com/watch?v=ROEJe-nBmqQ

You will need to get better at public speaking and telling stories. I recommend going to your closest Toastmasters club.


Yes, you're right, public speaking is not really my thing. I'm a tech nerd. I should probably rely on partnerships with other people to help it grow. But the first thing I wanted to find out was if the app is actually useful to anyone other than myself. And the response from this community is encouraging. I also thought about letting content producers publish their content/courses as collections here with restricted access. The app already supports that. It is also possible to realign the transcript to their "gold" transcript. They would have instant Anki cards in multiple languages, and the shadowing practice for their subscribers. I would have users, win-win :) But again I have no idea how to do marketing.


Frankly who cares about dildos when your personal freedom and private property of you and your family is at stake.You can buy uncountable dildo models in a shop.

I recently was in Venezuela, I have been in Cuba. I am a native spaniard. There you have a group of people that took control of the weapons in the country and uses it to basically enslave the rest of the country.

When the people in power have automatic weapons and you don't there is basically nothing you can do to defend yourself from the abuses of power.

That is a real thing the people in power have wet dreams and would love to do in any country, including the US.


USA seems to have a lot of automatic guns, but still succumbs to authoritarism and masked goons terrorising population.


Ummm, well, seems in historical perspective that quite a few instances of rebels using improvised molotov cocktails against tyrannical governments. When it's a country that the US disapproves of, seems they support this activity as "freedom fighters". But when this behavior occurs in protests inside the homeland, obviously the authorities come down hard.


> Frankly who cares about [sex stuff] when ...

Don't ask me. I think the focus on sex things is utterly insane. However, even on Hacker News many people definitely seem extremely concerned at least about people under 18 having access to Internet pornography, which to me, clearly isn't even remotely close to the biggest problem adolescents are facing.

In America we're worried about 3D printed ghost guns. Why? Ostensibly it's because ghost guns are showing up more at crime scenes, since 3D printing is just so accessible these days. How many gun deaths do 3D printed guns currently account for? As far as I know, something on the order of magnitude of around 0.01%, at least in America. That number is probably mostly small because we never actually really did anything about regular gun violence.

This is not a new pattern. Nuclear energy has very few fatalities in its track record. Shockingly few when you consider its reputation. Depending on how you count it, it is pretty much a rounding error by any reasonable measure. Meanwhile, we send at least somewhere around 10,000 bodies to the morgue every year as a result of pollution from natural gas power plants. How many from nuclear? Probably less than one on average. Certainly nowhere near 10,000 no matter how you shake it, twist it or bend it. Even the dumbest and least reputable studies couldn't force the number to half as high, which is pretty funny to me. Is nuclear fission still the future? Given the density and reliability of energy production from fission it's hard to entirely count it out even in a mostly solar + battery future... yet here we are, with politicians arguing about the safety of ~0 deaths per year energy production method vs >10,000 deaths per year energy production method.

Does it matter that politicians are so interested in regulating sex when bigger issues are at stake? Well, yeah. It's no less alarming even if it seems trivial. Not everyone has to pick every battle, but to me free expression is my bugbear so I give a fuck when it's in the mouths of politicians. The apparatus for suppressing expression never stops where it starts. Never. It wouldn't be right if it did, but the point of trying to come up with something that sounds like an obvious net good is that it can make people more forgiving to implement a terrible idea and set an awful precedent.

Which, of course, is why it's pretty likely in the near future you will need a driver's license to reply to this comment.

Somehow, it started with Internet pornography.


As an engineer myself working for them 20 years ago, we were certainly not well paid like the article said. Quite the contrary: I was still on University(had not finished the final project) and had to do most of the hard technical work myself for someone else to just overview the results and sign. My salary was miserable.

Once I had finished I could earn 3 to 4 times more on several places.

They were also extremely creative taking foreign systems, studying the patent and modifying it to pay zero to the creators of the patents. This was done with things like the aluminium beams for electricity delivery that I think was developed by Italians, or the tunnelling machines that had all the pieces replicated inhouse.


Sounds like Spain all right.

The article also makes a big deal out of country-level factors like the system of autonomous communities, governance, in-house expertise etc.

But all of those should apply to Malaga as well, which also built a metro in the 2000s. But that one became a city-wide joke for always being supposed to open "this year" and that continued for at least 5 years...

There was definitively none of the cheap or fast involved in that project, a relatively limited line to make travel to the airport more convenient which still couldn't deliver. Today it actually operates, but I think the rest of the network (it was supposed to be a "proper" metro system and not just isolated lines) is still vaporware. Haven't lived in Malaga in many years, though.


> in-house expertise

This is under-appreciated.

> I believe that the U.S. suffers from a distinct lack of state capacity. We’ve outsourced many of our core government functions to nonprofits and consultants, resulting in cost bloat and the waste of taxpayer money. We’ve farmed out environmental regulation to the courts and to private citizens, resulting in paralysis for industry and infrastructure alike. And we’ve left ourselves critically vulnerable to threats like pandemics and — most importantly — war.

> It’s time for us to bring back the bureaucrats.

* https://www.noahpinion.blog/p/america-needs-a-bigger-better-...

Even if you do outsource some level of tasks, you still need in-house folks who know something so you don't get fleeced.


> * https://www.noahpinion.blog/p/america-needs-a-bigger-better-...

Woah, this is really a “worst person you known just made a great point” moment for me.


Curious, what do you dislike about him? I find many of his takes to be pretty reasonable, though he gets temperamental and snarky sometimes.


Too many strongly held and unnuanced opinion about topics he has no qualification for.

Pundits like him are among the cancers of democracy.


> the tunnelling machines that had all the pieces replicated inhouse

Wait... did you build your own tunnel boring machines? Or just spare parts for them?


If you have the capability to manufacture heavy machinery, building it by copying an existing design is not rocket science. They're not crazy complicated and often built or adapted to a dig specific tunnel anyways.

Now, operating a tunnel-boring machine, that's a different beast, but you'd have to do that either way. Probably should get outside help if your engineers and scientists haven't planned and dug a tunnel in their life.


Don't students usually have miserable salaries? I would expect any engineer's salary to go up a lot after getting some experience and graduating, so it's hard for me to tell how much this data point generalizes to the whole project.

As for the patent side, I kinda give them kudos for that.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: