Perplexity Says You Can't Copyright Facts, and Retrieval Is the Real Legal Frontier
Asked about a publisher’s claim that commercial operators must pay to use its material, a Perplexity spokesperson gave a four-word answer in substance: facts aren’t copyrightable. It’s a real defence, it’s been good law for a very long time, and it may be the wrong argument for what these systems actually do.
The copyright fight so far has been about training. Did the model ingest protected work to learn from it, and is that fair use. That question is enormous, unresolved and slow, and it’s being litigated in dozens of cases.
Retrieval is a different question and it’s arriving faster.
Training and retrieval are not the same act
Training happens once, in advance, and produces a set of weights. Whatever a model absorbed from an article is diffused across billions of parameters, and the argument for transformation is at least coherent.
Retrieval happens at query time. The system fetches the current article, reads it, and generates an answer from it, right then, in response to a specific question. There’s no diffusion and no learning. There’s a copy being made, used commercially, and discarded.
That’s a much harder set of facts for the AI side. Encyclopedia Britannica and Merriam-Webster made exactly this distinction in their suit against OpenAI, pairing the mass-copying claim with a retrieval claim and alleging that outputs contain full or partial verbatim reproductions of their material. They also raised trademark, arguing that fabricated content attributed to their brands damages them.
That trademark angle is underrated. It doesn’t require resolving anything about fair use. It only requires showing that a system put words in your mouth.
Where the facts defence holds and where it breaks
Facts genuinely aren’t protected. The population of a city, the date of an election, the score of a match. Nobody owns those, and a system that reports them owes nothing to whoever published them first.
The defence starts failing when the thing retrieved isn’t a fact but a piece of work. The selection and arrangement of material, the framing, the analysis, the specific phrasing, the judgement about what matters in a story. That’s expression, and expression is protected regardless of how factual the subject is.
An answer engine summarising a reported investigation isn’t extracting neutral facts from the atmosphere. It’s extracting the output of months of work that only exists because someone did it, and reproducing the shape of it.
The commercial version of the argument
There’s a separate claim underneath, and it may matter more than either doctrine. Britannica’s complaint frames the harm as cannibalised traffic: summaries that satisfy the user so completely that nobody visits the source.
That’s a market substitution argument, and market effect is one of the factors courts actually weigh in fair use. It’s also the factor where publishers now have the strongest evidence, because the traffic data is unambiguous and public. Search referrals collapsed. The timing lines up with the deployment of AI answers. The correlation is about as clean as anything in this industry ever gets.
What to watch next
Retrieval cases will move faster than training cases because the facts are simpler. There’s a fetch, a copy, an output and a timestamp. No need to reconstruct what happened inside a training run two years ago.
Expect the first meaningful rulings to concern real-time retrieval rather than training data, and expect them to reshape product design more than they reshape licensing. A system that can’t retrieve freely has to license its sources or get worse at current events.
Both outcomes suit publishers. Neither is guaranteed.