Authors using a new tool to search a list of 183,000 books used to train AI are furious to find their works on the list.

  • kromem@lemmy.world
    link
    fedilink
    English
    arrow-up
    5
    arrow-down
    1
    ·
    1 year ago

    Generally they probably bought the books they read though.

    If George RR Martin torrented Tolkien, wouldn’t he be infringing on the copyright no matter how he subsequently incorporated it into future output?

    I completely agree that the training as infringement argument is ludicrous.

    But OpenAI exposed themselves to IP infringement by sailing the high seas in how they obtained the works in the first place.

    I hate that the world we live in is one where so much data is gated behind paywalls, but the law is what it is, and if the government was going to come down hard on Aaron Swartz for trying to bypass paywalls for massive amounts of written text, it’s not exactly fair if there’s a double standard for OpenAI doing the same thing in an even more closed fashion.

    But yes, the degree of entitled focus on the premise of training an AI as equivalent of infringing is weird as heck to see from authors drawing quite clearly from earlier works in their own output.

    • st0v@lemmy.zip
      link
      fedilink
      English
      arrow-up
      3
      ·
      1 year ago

      I have to assume that openAI also paid for the books. if yes then i consider it the same as me reciting passages from memory or coming up with derivative text.

      if no, then by all means, go after them and any model trainer for the cost of one book.

      Asking an LLM to recite an entire novel isn’t even vaguely a thing yet.

      • kromem@lemmy.world
        link
        fedilink
        English
        arrow-up
        3
        ·
        edit-2
        1 year ago

        Well, here’s straight from one of the suits against them:

        “The OpenAI Books2 dataset can be estimated to contain about 294,000 titles. The only ‘internet-based books corpora’ that have ever offered that much material are notorious ‘shadow library’ websites like Library Genesis (aka LibGen), Z-Library (aka B-ok), Sci-Hub, and Bibliotik. The books aggregated by these websites have also been available in bulk via torrent systems.”

        I’m not even sure how they would have logistically gone about purchasing 294,000 books in bulk in digital form to be fed into training. Using the existing collections seems much more likely, but I suppose we’ll see what turns up in litigation.

        Also, the penalty for downloading copyrighted material if willful infringement is up to $250,000 per work. So it’s quite a bit more than the cost of one book on the line…

    • Omniraptor@lemm.ee
      link
      fedilink
      English
      arrow-up
      2
      ·
      edit-2
      1 year ago

      God that Aaron/jstor thing makes me see red every time. Swartz was scraping jstor to publish it for the benefit of everyone, openai is doing it to make billions of dollars. Don’t forget who the bad guys are (and donate to sci-hub)