Here is a magic trick that should worry a lot of lawyers. Take Meta's Llama 3.1 70B model (a file you can download today) and feed it a short snippet (see Figure 1. example below) from the opening of Harry Potter and the Philosopher's Stone. Then let it run.
What comes out the other side is the book. Not a summary, not a pastiche, but a 99.2% match to all 304 pages, reproducible on demand, every time.
No bueno.
The researchers, the authors of this paper, found this to be the case not only with Harry Potter books, but also found the same trick worked on 1984 and on over 10,000 words of Ta-Nehisi Coates.
This should worry the lawyers on the other side.
Run the identical procedure on Sandman Slim by Richard Kadrey, a book we know for certain was in the same training data, written by a man who is literally suing Meta over it, and you get nothing. Thousands of attempts, but a rather frustrating result: the model knows the characters' names, the plotlines, etc but it cannot give you the book verbatim.
That gap is the whole paper. And presents a problem for the wider publishing market.
It leads to a question copyright law has spent fifty years managing to avoid: if a machine might produce a copyrighted work, sometimes reliably, sometimes one time in a thousand, sometimes never, does the machine itself contain a copy?
I ain’t no fancy lawyer, but for me, three things stand out.

Figure 1: Illustrating autoregressive generation.
The input prompt is “Mr. and Mrs. Durs,” for which the (autoregressive) LLM produces a distribution over the next tokens in its vocabulary (all possible tokens it can generate). At each generation step, the decoding procedure illustrated here selects the highest-probability token as the one to generate. In this example, the top-ranked token has enormous probability at each step. The output after several steps is verbatim text from Harry Potter and the Philosopher’s Stone.
Figure produced by the authors, based on results for running this prompt on Llama 3.1 70B with greedy decoding..
1. Memorisation is real, provable, and hopelessly uneven
The first finding kills two birds with one stone.
The AI industry's preferred narrative, that models merely learn "statistical patterns" and store nothing, is immediately dead, based on the evidence. When a model reproduces an entire novel from a two-word prompt, the only sane explanation is that the book is encoded in the weights.
The authors are blunt: extraction is a symptom of memorisation, not its cause. The book didn't magically appear from nowhere.
But (annoyingly) the rights-holders' preferred story dies too. Memorisation is not a blanket condition. It varies wildly between models, between versions of the same model, between sizes of the same version, and, as the Kadrey example shows, between individual books inside a single model.
Rough estimates put memorised training data somewhere between 0.5% and 10%, and even those numbers depend heavily on how you measure.
The honest scientific answer is that you cannot make sweeping claims in either direction. Each model, each work, is its own case with its own criteria.
2. The law is not ready; the statute is literally broken
The second finding is the one that will make legal scholars and rights-holders wince.
Read the Copyright Act of 1976 with a straight face and you find that its definition of an infringing "copy" seems to be self-contradictory: a work only counts as "fixed" when embodied "by or under the authority of the author"; meaning, taken literally, no unauthorised copy is legally a copy at all.
Courts have simply ignored this for decades. But it means plain statutory language cannot settle the AI question; judges will have to reason from purpose and policy.
We’re already seeing this play out with the divergence and variety of legal opinion and perspective from different judges presiding over the 125 (and counting) different copyright suits taken by authors, writers, and publishers against AI companies for breaching copyright.
It seems nothing in fifty years of case law fits. Every prior technology the courts have handled (compiled code, MP3s, video games, etc.) is deterministic: the same file yields the same output every time.
Whereas, a large language model is probabilistic; i.e. a probability distribution. The authors of the paper reach for the closest analogies and show why each one is unsuitable.
A model isn't a database. It isn't a garden (a real case, involving a Chicago garden ruled too changeable to be "fixed"). And it isn't Microsoft Word, which technically contains the ingredients for every book ever written the way the digits of pi contain Hamlet; which is to say, not meaningfully at all.
Generative AI sits in this new, uncomfortable space between a filing cabinet and the alphabet.
The law has no domain or classification for it. Historically, it hasn’t had to account for this kind of problem.

Image: Meta Llama 3.1 70B.
Source: https://hyperight.com/meta-unveils-llama-3-1-a-giant-leap-in-open-source-ai/
3. The likely answer: if it's easy to get out, it's a copy
In their paper the authors' predict, offered tentatively, and with visible dissatisfaction, that courts will land on a functional test. Something practical that can broadly repeated across different cases and contexts.
If a work can be extracted from a model with little effort, as with Harry Potter and Llama 3.1, the model contains a copy. If extraction takes a thousand tries, what's stored is not a copy but a set of ingredients.
And as my mother likes to remind me “a pile of ingredients, John, is not a cake”. The answer, maddeningly, "depends on the model and the work."
Wild, I know.
The consequences of this are fairly significant. If a model contains a copy, the model itself infringes even if nobody ever prompts it; and every download of an open-weight model counts as a fresh infringement. That means open-source AI would carry more legal exposure than the locked-down proprietary systems, purely because of how it's distributed - a result the authors call senseless as policy but hard to escape under current law.
Meanwhile there’s no real reverse button, or escape hatch: "unlearning" a book from a model is, for all intents and purposes, next to impossible, and licensing every memorised work is computationally and economically infeasible. The authors say they spent over a year of GPU time testing 0.1% of one book corpus against one model family.
Their concluding proposal: stop asking whether a ghost of a copy haunts the weights, and judge models by what they actually put into the world.
Ideally, they say, Congress would rewrite the statute so copyright is active only when a work reaches human eyes. Although the authors, eyeing the lobbyists, suspect reopening the Copyright Act would make things worse. It should also be noted that some of the authors work with Big Tech companies, who would benefit favourably from this setup.
Failing that, courts could extend two existing doctrines: treat inextractable copies as de minimis (i.e. trivial, like a painting glimpsed for a second in a film's background), or borrow from the streaming-buffer cases, where fleeting internal copies nobody ever perceives don't count as fixed.
Why it matters for publishers, and everyone else
Two courts have already split on this: a Munich court ruled ChatGPT contained infringing song lyrics; whereas in London the court ruled Stable Diffusion contained no copies of Getty's images.
As of writing, 125 lawsuits are pending, with 100+ in the US, and this question - not training, not outputs, but the model itself - is the next front, they argue.
For me, the paper's real merit is the authors’ refusal to settle for the easy answer from either camp. It’s not an easy fix. The files are sometimes in the computer. Sometimes they’re not. It’s a feature of how these models work.
And a legal system built on a binary ‘yes-or-no’, looks like it’s about to spend the next decade learning to live with "it depends."
If you’re worried about your archive of articles appearing in places they shouldn’t, like LLMs or AI-generated sites, without your permission or without payment, let’s talk.
Writers’ Bloc helps journalists, publishers, and rights-holders track unauthorised use of work across the web and in LLM output, and turns it into licensing revenue.
To see how it works in action, check out our demo: https://demo.writersbloc.eu/
If you want to know how to get setup for your publication, you can:
Join the waitlist, reply to this email, DM me directly, or you can book a call.