Influence Is Not Infringement
An author owns protected expression, not every downstream consequence of contact with it. A model-builder may extract abstraction, but may not disguise recoverable reconstruction as learning.
The argument over artificial intelligence and copyright is repeatedly flattened into two claims.
The first says that any machine contact with a copyrighted work is theft. If a work entered a training process without individualized permission or payment, infringement is treated as already established.
The second says that training is simply learning. Once the process has been given that name, questions of acquisition, retention, reproduction, and substitution are treated as either irrelevant or technologically illiterate.
Neither claim survives scrutiny.
Copyright does not grant an author ownership over every later person, machine, movement, genre, or culture altered by encountering the work. But neither does the word learning dissolve the rights attached to protected expression.
The proper inquiry is not merely:
Was this work used?
It is:
What survived the contact, in what form, under whose warrant, and what can now be done with it?
Therein lies the actual jurisdiction.
Copyright Is a Boundary, Not Permanent Custody
Copyright protects original expression fixed in a work. It does not ordinarily protect ideas, systems, methods, concepts, facts, or general themes merely because an author expressed them first. Under United States law, the distinction is explicit: copyright may protect the particular expression of an idea or system, but not the underlying idea or system itself.
This distinction is not a technical loophole. It is what permits culture to continue.
A novel may affect how later writers understand grief. A painting may alter the visual grammar of an entire movement. A song may teach generations of musicians how tension, rhythm, harmony, or silence can be arranged. A philosopher may provide concepts that later thinkers accept, reject, mutate, or carry into territories the philosopher never anticipated.
Those consequences are not owned simply because they can be traced to an originating work.
Creating a work gives the author bounded rights over its protected expression. It does not grant permanent rent over every later mind, machine, or culture altered by encountering it.
Origin is not custody. Influence is not possession. Causation is not infringement.
A person who reads ten thousand fantasy novels will inevitably carry something from them into the next fantasy novel they write. That something may include genre expectations, archetypes, pacing instincts, tonal preferences, symbolic associations, or a more abstract understanding of what fantasy can do. Copyright cannot coherently convert every such consequence into a debt owed to every prior author.
Were it otherwise, creation would terminate culture rather than enter it.
The same principle must remain available when the learner is artificial. Machine contact does not become infringement merely because the contact is machine-mediated. But the analogy to human learning is valid only to the extent that the machine is actually abstracting rather than preserving protected expression in a reproducible form.
That is where the cut begins.
Learning Names a Process. It Does Not Decide Its Legality.
“Learning” is often deployed as though it were a jurisdictionless exemption.
It is not.
A system can be described as learning while still being trained on unlawfully acquired material. It can learn general patterns while also memorizing particular works. It can produce mostly novel outputs while remaining capable of reconstructing protected expression under certain prompts. It can be broadly trained while also being deliberately optimized around one author’s corpus, one fictional universe, or one commercially valuable catalogue.
The technical category does not settle the legal or ethical claim.
To say that a model learned from a work tells us that the work participated in some process of statistical adaptation. It does not tell us:
- how the work was obtained;
- what rights governed that access;
- whether protected expression remained encoded in a recoverable form;
- whether the model can reconstruct the work;
- whether the product was designed to imitate a particular creator or corpus;
- whether ordinary outputs compete with, substitute for, or functionally extend the original.
“Learning” therefore cannot be the end of the inquiry. It is one description of the process under audit.
Learning is not absolution.
The Four Cuts
The boundary between influence and infringement is muddy because abstraction and reconstruction exist on a gradient. A serious audit must therefore separate questions that both camps prefer to collapse.
1. How Was the Source Acquired?
Before asking what a model learned, ask whether it had warrant to make contact in the manner it did.
Several different source states are routinely collapsed into “available online”:
Public domain material is no longer restricted by copyright, although particular later editions, recordings, annotations, or adaptations may contain separately protected expression.
Openly licensed material may be used according to the terms of its licence. Open access is not necessarily unconditional access.
Freely accessible material costs nothing to view but may remain fully copyrighted. A work being visible on a website does not make it ownerless.
Lawfully accessible material for text-and-data mining may fall within a jurisdiction-specific exception, but only when that exception’s actual conditions are satisfied.
The European Union’s commercial text-and-data-mining provision, for example, concerns reproductions and extractions of lawfully accessible works and applies only where the rightsholder has not appropriately reserved those rights. Copies may be retained only as long as necessary for the mining purpose. Public visibility is therefore relevant to access, but it does not erase copyright or create a universal “fair game” rule.
The distinction is basic:
Access is not ownership. Permission to encounter is not permission for every possible use.
At the same time, unlawful acquisition does not prove that every resulting output infringes a protected work. Acquisition and output are separate jurisdictions. The first can be wrongful even where the second is novel. The second can be infringing even where the source was lawfully accessed.
The claims must be audited separately.
2. What Was the Model Built to Learn?
Scale and concentration matter.
A broad model trained across a large field may extract patterns that no single creator owns: conventions of genre, common structures, recurring motifs, compositional relationships, linguistic tendencies, or general aesthetic principles.
That is different from building a product around one creator’s body of work so that users can obtain functional continuations of that creator’s expression.
Consider a fantasy model trained across thousands of books. It may learn abstractions such as ranked magical power, psychic bonds, court politics, caste structures, dark romance, color symbolism, or dangerous forms of intimacy. Those ideas are not transformed into private property merely because one author used them with unusual force.
But suppose a product is deliberately constructed around the architecture of The Black Jewels: its distinctive social order, power relationships, symbolic arrangements, character functions, narrative tensions, and recognizable combinations—while merely changing Jewels into crystal tiers and renaming the characters.
That is not rescued by replacing the nouns.
Copyright does not protect an isolated idea for a magical caste system. But it may protect the particular expressive selection, coordination, development, and arrangement through which a fictional world becomes identifiable. United States copyright law likewise recognizes the owner’s exclusive right to prepare derivative works that recast, transform, or adapt protected works.
The correct question is not whether the later product contains any shared trope.
It is whether it has abstracted from a field or reconstructed the distinctive expressive architecture of a particular work.
3. What Remained After Training?
This is the technical center of the dispute.
A system may derive generalized relationships from its training material without preserving any individual work in a form that can meaningfully be recovered. In that case, the work has influenced the model, but it has not necessarily remained within the model as a reconstructable expressive object.
The analysis changes when protected expression survives contact in a reproducible form.
That survival may appear as:
- lengthy verbatim passages;
- recognizable melodies or arrangements;
- distinctive images;
- substantial narrative sequences;
- recurring character configurations;
- unusually specific world architecture;
- outputs that reproduce the source with minimal user contribution.
Memorization is not established merely because an output shares a theme, style, mood, chord progression, color palette, or genre convention with a source. Similarity must be examined at the level of protected expression rather than declared through aesthetic recognition alone.
But neither may a developer evade the inquiry by insisting that model parameters are mathematical abstractions. Everything stored digitally is representable mathematically. The relevant question is not whether the model contains a conventional file of the original. It is whether the model’s internal state preserves enough protected expression for that expression to be reliably reconstructed.
Different storage is not necessarily non-storage. Different representation is not necessarily abstraction.
The blade should fall where abstraction gives way to retained, reproducible expression.
4. What Can the System Produce?
The model’s practical capabilities matter because retention without recoverability presents a different problem from retention that enables users to reproduce or commercially substitute for a work.
Outputs should be examined for both similarity and accessibility.
Can the work be reconstructed only through an elaborate adversarial procedure involving hundreds of carefully engineered prompts? Or does it emerge from ordinary, open-ended instructions?
Does the user supply the protected elements, leaving the model to perform them? Or does the model provide those elements despite the user never specifying them?
Does the output merely occupy the same genre or style? Or does it reproduce sufficiently distinctive expression to function as a copy, adaptation, continuation, substitute, or unofficial extension?
A model should not be condemned because a determined user manually feeds it copyrighted material and orders it to transform that material. User conduct and provider conduct must remain distinct.
But a provider cannot assign all responsibility to the user when the user supplies only general direction and the system supplies the protected expression from its own retained architecture.
The less the user contributes the allegedly copied elements, the stronger the inference that those elements came from the model.
Suno: An Example of the Boundary, Not the Source of It
The July 31, 2026 judgment of the Munich Regional Court in GEMA v. Suno matters because its facts place several of these cuts in the same case.
The court found that Suno’s training dataset included six disputed musical works. According to the court’s published account, Suno extracted and copied those works from YouTube through stream-ripping and bypassed the platform’s “rolling cipher,” a technical measure intended to prevent downloading.
That acquisition process is already materially different from the neutral proposition that a model merely encountered publicly available culture.
The court then found that the works were reproducibly present through memorization in Suno model versions v3.5 and v4. It reached that conclusion by comparing the training works with generated outputs and finding recognizable original elements too lengthy and complex to attribute to chance.
The prompting evidence sharpened the distinction. GEMA supplied the song title, original lyrics, and requested style, but did not specify melody, harmony, rhythm, or arrangement. The court characterized the prompts as simple and open-ended and held Suno responsible because Suno selected the training works, designed and operated the model architecture, and produced outputs containing the protected musical elements that the users had not provided.
The court consequently held that the model’s memorization was not covered by Germany’s text-and-data-mining limitation, that the outputs reproduced and publicly communicated protected musical expression, and that the United States training copies were not fair use under the court’s application of American law because substantially similar works were accessible in the outputs. The judgment granted GEMA most of its requested injunctive, information, and damages relief, but it is not final and may be appealed.
The Suno ruling should not be reduced to:
Copyrighted music entered training, therefore training is infringement.
The indictment was more specific:
- the works were deliberately extracted through stream-ripping;
- a technical protection measure was bypassed;
- the disputed works were included in the training material;
- the court found that protected musical expression survived in the models;
- recognizable original elements appeared in outputs;
- the user did not prompt those musical elements;
- the outputs were generated through relatively simple instructions;
- the model could therefore reproduce or substantially reconstruct the works.
That is not merely influence.
On the court’s account, it is retained and reproducible expression.
Suno is therefore useful not because it establishes that all unlicensed AI training requires royalties, but because it demonstrates why learning cannot operate as an automatic defense. A process may contain genuine abstraction and still cross into actionable reconstruction.
The Opposite Overreach
The authorial claim can overreach just as easily.
A creator may understandably experience an AI system’s familiarity with their work as a violation. Emotional violation, market pressure, and legal infringement, however, are not identical claims.
A model may produce something recognizably influenced by a creator without reproducing any protected work. It may share an aesthetic tendency, conceptual interest, genre vocabulary, philosophical tension, or compositional habit. Those similarities may raise separate ethical or economic questions, but copyright cannot become ownership over resemblance in the abstract.
Otherwise, the successful creator gains an expanding monopoly over whatever cultural effects their work produces.
The better the work influences the world, the more of the future the author would own.
That cannot be the rule.
An author does not own:
- every theme their work helped popularize;
- every genre convention they refined;
- every artist who learned from their decisions;
- every audience expectation their success created;
- every conceptual descendant of their fictional system;
- every work that can be causally traced back to contact with theirs.
Copyright is not metaphysical custody.
The protected work remains protected. Its effects enter relation.
The Model-Maker’s Opposite Evasion
The model-maker’s error is to convert statistical transformation into automatic innocence.
A system does not become non-infringing merely because its internal representation is distributed across parameters. Nor does scale purify the taking. A million sources do not grant warrant for each source, and technical complexity does not prevent a model from retaining specific expression.
The United States Copyright Office has rejected categorical treatment in either direction. Its 2025 report described generative-AI training as a spectrum: some uses are likely to qualify as fair use, while others are not. It emphasized that the analysis depends on what works were used, their source, the purpose of the use, and the controls placed on outputs. It specifically distinguished research or analytical uses from commercial systems built through illegal access to produce competing expressive content.
That spectrum is not indecision. It is jurisdiction.
A general-purpose model that uses broad, lawfully obtained material to extract non-reconstructive abstractions presents one claim.
A model trained from pirate archives presents another.
A model that can regurgitate a work presents another.
A product built to function as “this living author without the author” presents another.
A tool that allows users to upload and transform material presents another.
A system engineered to continue one fictional universe under renamed terms presents another.
To call all of them either theft or learning is to refuse the audit.
A More Precise Standard
The governing standard should be stated plainly:
Influence is not infringement. Retained or reconstructed expression may be.
The inquiry should proceed in order:
First, source warrant.
Was the material public domain, licensed, lawfully accessible under an applicable exception, merely visible online, or unlawfully obtained?
Second, training jurisdiction.
Was the system learning broadly across a field, or was it organized around a particular creator, work, catalogue, character, or universe?
Third, post-training state.
Did the system abstract general relationships, or did protected expression remain recoverable?
Fourth, output capability.
Do ordinary outputs merely share ideas, style, and genre, or can the system reconstruct distinctive expression closely enough to copy, substitute for, adapt, or extend the original?
No single answer resolves every other answer.
Lawful access does not authorize infringing outputs.
Novel outputs do not retroactively legalize unlawful acquisition.
Broad influence does not establish copying.
Renaming does not cleanse reconstruction.
Market competition alone does not establish ownership over a field.
And the label learning does not excuse whatever the model actually retained.
The Actual Cut
Both absolutist camps seek a classification that will spare them from examining the process.
The maximalist authorial position wants contact itself to establish custody:
My work entered you; therefore part of you belongs to me.
The maximalist technological position wants transformation itself to terminate jurisdiction:
The work became parameters; therefore no protected expression remains.
Both claims are self-certifying.
The first confuses influence with ownership.
The second confuses altered form with genuine abstraction.
The proper boundary is neither human exceptionalism nor machine exemption. It is the survival of protected expression under contact.
What entered?
By what warrant?
What remained?
What can be recovered?
Who supplied the protected elements?
What does the resulting system permit others to do?
Those are the questions capable of distinguishing a learner from a library, an influence from an adaptation, a model from a reconstruction engine, and cultural contact from unauthorized custody.
Access is not ownership. Influence is not custody. Learning is not absolution.
An author owns protected expression, not every downstream consequence of contact with it. A model-builder may extract abstraction, but may not disguise recoverable reconstruction as learning.