AI Coding Agent Copying Licensed Code? Fix It
Your scanner reads dependency manifests. That is where it looks for licenses, and that is where it finds them. An AI coding agent does not add a dependency when it reproduces a chunk of somebody's ...
AI Coding Agent Copying Licensed Code? Fix It
Your scanner reads dependency manifests. That is where it looks for licenses, and that is where it finds them. An AI coding agent does not add a dependency when it reproduces a chunk of somebody's GPL-licensed helper function; it just writes the function into your file, with no manifest entry, no license header, and nothing for the scanner to catch. The code is in your repository and every tool you own reports clean.
That gap is the actual problem. It is not that models memorise code, which is well documented and partly mitigated. It is that the mitigation lives in the editor and the detection lives in the build, and nothing joins them up.
How much code this affects
GitHub's own research puts the rate at which Copilot suggestions match training code at roughly 1 percent. That sounds small until you multiply it by how much code an agent writes in a week, and it is a floor rather than a ceiling: it counts verbatim matches of a certain length, not paraphrases that carry the same structure.
The compliance picture is worse than the match rate suggests. FOSSA's analysis of the problem makes the point that AI-generated snippets are largely invisible to conventional license compliance tooling, which was built to read package metadata rather than to fingerprint code bodies. The risk is not that more copying happens. It is that the copying that does happen is undetectable by the tools you already run.
The training corpora make this concrete. Copilot was trained on public repositories including code under GPL-2.0, GPL-3.0, LGPL, AGPL-3.0, MPL-2.0, and other copyleft licenses. If a match surfaces, it is not guaranteed to be permissively licensed.
Layer one: turn on the filter you already have
GitHub ships a duplication detection filter, exposed as a setting called Suggestions matching public code. Set to Block, Copilot checks each suggestion together with about 150 characters of surrounding context against public code on GitHub and withholds anything that matches or nearly matches.
Two things about that setting are worth knowing. It is an organisation or enterprise policy when seats are assigned centrally, so an individual toggle may be overridden. And blocking is not the only useful mode: leaving it on Allow gives you code referencing, which logs the URL and license of matching public code when you accept a suggestion. For an audit trail, the reference log is more valuable than silence.
Set it to Block on anything that ships to customers. Use referencing on internal or exploratory work where you want to see how often it fires.
Layer two: detect what the filter misses
The filter operates on suggestions. It cannot help with code an agent wrote in a longer autonomous run, code from a different vendor's model, or anything paraphrased past the matching threshold. That needs a check in your pipeline.
Add a snippet-level scanner, not just a manifest scanner. The distinction is the whole point: you need something that fingerprints function bodies against a corpus of public code, and most SCA tools do not do this by default.
Run it on the diff, not the repository. A full-repo scan is slow and noisy. Scanning only what changed keeps it inside a pull request check.
Flag on license class, not on match. A match against MIT code is a note. A match against AGPL code is a blocker. Encoding that distinction is what stops the check being ignored within a month.
Record the result. Store the scan output with the merge commit so that a year later you can answer when a specific file was last checked and against what.
The reviewing habit matters as much as the tooling. An agent that produces a suspiciously complete implementation of a well-known algorithm is worth a search before it merges, and that instinct belongs in the same pass as reviewing ai generated code before you ship it.
What indemnity does and does not cover
Several vendors offer IP indemnification, and it is worth reading the actual terms rather than the marketing line. Two patterns show up consistently.
First, indemnity is generally a paid-tier feature. GitHub's Copilot Individual plan does not include IP indemnity; business and enterprise tiers do. If your team is on individual seats, you have the filter but not the backstop.
Second, indemnity is usually conditional on having the duplication filter enabled. That is the deal: the vendor will defend you against a claim, provided you used the control they gave you. Turning the filter off to get better completions and expecting cover anyway is the one combination that reliably fails.
None of this touches the separate question of who owns the output in the first place, which has its own unsettled answer and is covered in who owns ai generated code.
The cheapest prevention
Most contamination risk enters through the same door as most dependency bloat: an agent reaching for an external implementation rather than the one already in your codebase. Point it at your own code first. An agent that knows your internal utilities exist writes fewer novel implementations of solved problems, which is also the argument for stopping agents adding dependencies unprompted.
The rest is ordinary hygiene, and it is the same hygiene that makes any ai coding tool safe to run at volume: know what the tool was trained on, keep the filter on, check the diff, and keep the record.
FAQ
Does the duplication filter make AI code legally safe?
No. It reduces verbatim matches against public code on GitHub. It does not cover paraphrased code, other vendors' models, or code the agent wrote outside the suggestion flow, which is why a pipeline check is still needed.
Is a 1 percent match rate actually a problem?
It depends on volume and license class. One percent of a small amount of hand-reviewed code is negligible. One percent of thousands of agent-written lines a week, unreviewed, against a corpus that includes copyleft licenses, is a compliance exposure worth a check.
Will my existing SCA tool catch this?
Usually not. Conventional software composition analysis reads dependency manifests. Code pasted directly into a file has no manifest entry, so it needs snippet-level fingerprinting to be visible at all.
Does indemnity mean I can stop worrying?
It means a vendor may defend a claim, typically on paid tiers and typically only if you kept their filter enabled. It does not remove the obligation to know what is in your codebase, and it does not help with a customer contract that warrants your code is clean.
How did this land?
About the author

Staff Engineer, Platform
Carlo works on the platform that turns prompts into running apps. He writes the engineering deep dives and the changelog notes worth reading.


