Artificial intelligence labs are in a new arms race to buy up millions of rare books, slicing them open, scanning the pages and pulping the remains — sparking concerns that the last remaining copies of out-of-print texts are being destroyed on an industrial scale.

ISBNdb notes that “print books from the pre-LLM era are structurally guaranteed to be free of this contamination”.

“Millions of the most valuable books have never been digitised. They exist only in physical form, scattered across library shelves, used bookstores, and out-of-print catalogues. We get them to you at scale.”

  • unwarlikeExtortion@lemmy.ml
    link
    fedilink
    English
    arrow-up
    15
    ·
    1 day ago

    To be honest, I’d be MUCH less against this practice if this shredding at the very least included saving the original scans at at least 300 dpi for everyone to access, freely.

    Ideally, they wouldn’t cut them by the spine, but even giving the despined ones to an archive/library would be just above the unpermissible line.

    Despining, scanning and public access archiving is borderline permissible.

    • JackbyDev@programming.dev
      link
      fedilink
      English
      arrow-up
      9
      ·
      23 hours ago

      Also, if every AI company does this, then it’s more books gone. If they cooperated they’d only need to do it once. Good scans of all books in the public domain would be so useful for us all as society.

      • Armok_the_bunny@lemmy.world
        link
        fedilink
        English
        arrow-up
        2
        ·
        7 hours ago

        While that would be nice, unfortunately it would almost certainly run into the very same legal issues that are also causing them to pulp the books after scanning them.

          • Armok_the_bunny@lemmy.world
            link
            fedilink
            English
            arrow-up
            2
            ·
            3 hours ago

            IIRC destruction of the physical copy to prevent resale while still possessing a copy is required for book scanning to be free use, or otherwise not encounter problems with copyright law. Proceeding to share/distribute those scans, or especially trying to charge money for them, would encounter those exact same copyright barriers.

            • JackbyDev@programming.dev
              link
              fedilink
              English
              arrow-up
              2
              ·
              3 hours ago

              Oh, maybe. Even then, for things in the public domain this shouldn’t matter. I know that’s probably not the majority of what they’re scanning though.

              • Armok_the_bunny@lemmy.world
                link
                fedilink
                English
                arrow-up
                1
                ·
                18 minutes ago

                I would assume they aren’t bothering to scan it if it’s in the public domain. Much easier to just grab a copy off archive.org or any other website hosting such things, and that’s assuming the entirety of the public domain isn’t already in the training set.