Atlas / GOVERN & LEAD / Governance / IP & Copyright
DEEP-DIVE · GOVERNANCE

IP and Copyright in the Age of AI

Who owns what a model produces, and was it legal to train it? One question has a working answer, the other is open litigation. A US-lens map of both fronts, and the controls that reduce exposure now.

TL;DR
  • AI copyright is two separate questions: who owns model outputs (partially answered: human authorship governs) and whether training on copyrighted works was lawful (genuinely unsettled, active litigation).
  • Purely AI-generated material is not copyrightable under current US guidance. Meaningful, documented human creative contribution is what makes output ownable, so ownership is a process-design problem.
  • You cannot settle the law from a platform team, but you can control exposure: provenance records, license-aware code tooling, near-copy screening for high-stakes assets, and careful reading of provider indemnities.

Two fronts, two very different questions

Nearly every executive conversation about AI and intellectual property collapses two questions into one, and that collapse is where bad decisions come from. The first question is about outputs: when a model generates an image, a paragraph, or a function, who owns it, and can anyone own it at all? The second is about inputs: was it lawful to train the model on copyrighted works in the first place? These questions live under different legal doctrines, involve different parties, and sit at very different levels of settledness.

The outputs question has a partial but workable answer in the United States. Copyright requires human authorship, so ownership of AI-assisted work turns on what the humans involved actually contributed. That is something an enterprise can design for, and document. The training-data question has no such anchor. It is being contested in ongoing litigation, the analysis is intensely fact-specific, and outcomes so far have varied. Nobody, including your provider's sales team, can tell you how it ends.

  +---------------------------+  +---------------------------+
  |  FRONT 1: OUTPUTS         |  |  FRONT 2: TRAINING DATA   |
  |  Who owns generated work? |  |  Was the training lawful? |
  |                           |  |                           |
  |  Status: partial answer   |  |  Status: unsettled        |
  |  Anchor: human authorship |  |  Anchor: none yet         |
  |  Lever:  process design   |  |  Lever:  contracts and    |
  |          plus records     |  |          indemnities      |
  +---------------------------+  +---------------------------+
     you can engineer this          you can only allocate this
The two fronts respond to different levers. One rewards workflow design and documentation; the other rewards contract discipline and hedged positions.

The practical consequence is that the two fronts demand different responses. On outputs, invest in process: how people use the tools determines what the company can claim. On training data, invest in risk allocation: contracts, indemnities, vendor selection, and the humility to hold positions loosely until courts finish arguing. Confusing the two leads either to paralysis (freezing all generative AI use because litigation exists somewhere) or false comfort (assuming a provider's indemnity resolves questions it never touches).

Outputs: the human-authorship rule

US Copyright Office guidance has been consistent on the core point: copyright protects human authorship. Material generated entirely by a machine, with no human creative contribution, is not protectable, however polished it looks. Under the Office's current position, a prompt alone is generally not enough; describing what you want reads more like commissioning a work than authoring one. But the guidance leaves real room on the other side of the line. Where a human makes meaningful creative contributions, or creatively selects, arranges, and modifies AI-generated material, those human-authored elements can support protection. Registration practice expects applicants to disclose AI-generated content in submitted works, which makes honest internal records more than a nicety.

For an enterprise, this converts a legal abstraction into a design question. If a marketing team wants to own a campaign built with generative tools, the workflow should make human authorship visible and provable: creative briefs that predate generation, iterative direction rather than single-shot prompting, human editing and composition on top of raw output, and records of who did what. The difference between "the model made it" and "our team made it with model assistance" is not spin. It is the difference the current framework actually turns on, and it is established in the workflow, not after the fact.

There is also a defensive corollary that gets less airtime: material your company cannot own, competitors can freely reuse. A purely machine-generated logo or product description carries no exclusivity. For assets where exclusivity matters, human authorship is not a compliance chore, it is what creates the asset's value.

This line is still moving. Copyright Office guidance has evolved through successive statements and reports, and courts continue to weigh in. Treat "human authorship, documented" as the durable working principle, and expect the fine boundaries, especially around prompting and iterative refinement, to keep shifting. Revisit the position at least annually with counsel.

Training data: genuinely unsettled

The second front is the one that generates headlines: whether training models on copyrighted works without permission is lawful. Here the honest answer is that nobody knows yet. Multiple cases are moving through US courts, the central defense is fair use, and fair-use analysis is inherently case-by-case: it weighs factors like the purpose and character of the use and the effect on the market for the original, against specific facts about how data was acquired, processed, and what the model does with it. Outcomes to date have varied with those facts, and appeals are in play.

This deep-dive deliberately draws no conclusion, because no responsible one is available. What an architect should internalize instead is the shape of the uncertainty. First, it is fact-dependent: how the data was obtained and whether outputs substitute for the originals can matter as much as the act of training itself. Second, it is jurisdiction-dependent: other countries have taken different approaches to text and data mining, so a global enterprise cannot assume one answer travels. Third, it is slow: definitive resolution likely arrives in layers, over years, not in one verdict.

No single ruling is the final word. Any individual decision is tied to its facts, its parties, and its procedural posture, and may be appealed, distinguished, or contradicted elsewhere. If a headline says training is "legal" or "illegal" as of some date, the headline is wrong. Build positions that survive either direction rather than betting the content strategy on one outcome.

The enterprise consequence is not to stop using generative AI. It is to recognize that training-data risk is mostly borne upstream, at the provider, and that your levers are contractual: what the provider warrants, what it indemnifies, and what happens to you if its position collapses.

What providers promise

Because customers kept asking the training-data question, several major providers now offer copyright indemnities: commitments to defend or compensate customers facing third-party infringement claims over generated output. These commitments are real and worth weight in vendor selection. They are also narrower than the marketing slide implies, and the gap between the headline and the terms is exactly where an enterprise gets surprised.

Scope varies widely. Some indemnities cover output claims only, not claims about the training itself. Most attach conditions: they commonly apply only when built-in safety and filtering features are left on, only to unmodified services, and only when the customer did not knowingly engineer an infringing result. Some exclude fine-tuned models, customer-supplied training data, or specific product tiers. Caps, notice requirements, and control-of-defense clauses differ. None of this is hidden; it is simply in the terms rather than the press release.

Typical conditionWhat it usually requiresWhy it matters
Safety filters enabledContent filters and mitigations left at defaultsDisabling filters for quality reasons can silently void coverage
Unmodified serviceUsing the service as shipped, within its termsCustom fine-tunes or altered pipelines often fall outside scope
No induced infringementCustomer did not prompt for a near-copy of a known work"Draw it in the style of X, exactly" scenarios are excluded
Covered products onlySpecific services and tiers named in the termsPreview features and consumer tiers are frequently out
Notice and cooperationPrompt claim notification, provider controls defenseSlow escalation paths inside the enterprise can forfeit rights

The operational takeaway: an indemnity is a control only if your deployment actually satisfies its conditions. That means someone must map each production use of a generative service against the indemnity terms, and re-check when either changes. This is legal-plus-platform work, not either alone.

Where the enterprise is exposed

Abstract risk becomes concrete in four recurring places. Each has a different owner inside the organization, which is why "legal is handling AI copyright" is not a plan.

Notice that only the first two are about infringement claims arriving. The third is value quietly not being created, and the fourth is confidential value quietly leaving. An IP posture that only watches the courtroom door misses half the exposure.

The controls that work now

None of the open legal questions prevent an enterprise from acting today. The controls below do not require predicting litigation outcomes; they reduce exposure under any plausible resolution, which is exactly the property a control should have when the law is unsettled.

  1. Document human creative direction. For assets the company wants to own, require briefs, iteration history, and named human contributors as part of the workflow, captured at creation time, not reconstructed later.
  2. Keep provenance records for generated assets. Log which tool, model version, and prompt lineage produced each significant asset, and whether it was human-edited. This supports registration disclosure, indemnity claims, and takedown responses alike.
  3. Use license-aware tooling for code. Enable assistant features that block or flag verbatim matches to public code, and run license scanning in CI so contamination is caught before merge, not in due diligence.
  4. Screen high-stakes outputs for near-copies. For flagship campaigns, product art, and anything shipping at scale, add a similarity review (search, reverse-image lookup, human eyes) before release. Cheap relative to a dispute.
  5. Review indemnity and data terms in contracts. Map each production use against indemnity conditions, confirm inputs are not used for provider training where that matters, and re-review at renewal and at major service changes.
  6. Route novel questions to counsel. Define a fast escalation path for the genuinely new cases: training your own model on third-party data, scraping, style imitation of living artists, or anything a headline could plausibly feature.

The first two controls are worth singling out because they compound. Provenance is the substrate everything else consumes: without it you cannot prove authorship, satisfy disclosure expectations, or invoke an indemnity cleanly. Teams that wire provenance capture into the platform once stop paying for it per asset.

The architect view

The temptation with AI and IP is to treat it as either a blocker (nothing ships until the law is clear) or an afterthought (legal will sort it out if anyone sues). Both are abdications. The law will not be clear for years, and by the time a dispute arrives, the records that would have protected you either exist or they do not. The architect's job is to make sure they exist.

Practically, that means running IP as a governance lane with the same machinery as security or privacy: a named owner, a register of generative use cases, controls wired into the platform rather than into memos, and review triggers when tools, terms, or law change. Provenance capture belongs in the pipeline next to logging and evaluation, not in a policy document nobody operationalizes. Indemnity conditions belong in the deployment checklist, because coverage that is silently voided by a disabled filter is worse than no coverage: it is false confidence. The rest of the governance category covers how this lane fits alongside its siblings.

And hold positions with appropriate humility. On outputs, the human-authorship principle is stable enough to build workflows on, while its edges keep moving. On training data, the only defensible enterprise position is a hedged one: choose providers whose terms shift risk upstream, keep your own records clean, and avoid strategies that only work if the litigation breaks one particular way. This article is orientation for architects and executives, not legal advice; jurisdictions differ, facts dominate, and the novel cases belong with counsel.

← Explainability and Interpretability: What Can You Actually Know? ALL OF GOVERNANCE