When an Employee Uses AI to Fake Work

A report lands that reads well and cites two sources that do not exist. A junior developer's pull requests are suddenly three times larger and none of the tests are real. A weekly update describes customer conversations nobody can find in the CRM.

Cecilia Iona
Cecilia Iona
Senior Editor, AI & Product
10 September 20261 min read

When an Employee Uses AI to Fake Work

A report lands that reads well and cites two sources that do not exist. A junior developer's pull requests are suddenly three times larger and none of the tests are real. A weekly update describes customer conversations nobody can find in the CRM.

The instinct is to treat this as a new category of problem created by AI. It mostly is not. What AI changed is the cost of producing convincing output, which means an old problem now surfaces faster and looks more polished. Handling it well starts with separating three situations that get bundled together and are not remotely the same.

Three different things, three different responses

1. Using AI openly and checking the output. Someone drafts with a model, verifies the facts, and takes responsibility for what they send. This is not faking work. It is work. If your reaction to discovering it is disciplinary, you will teach your team to hide their tooling, which is how you end up with shadow AI and no visibility at all.

2. Using AI and not checking the output. The report is real work, submitted without verification. Fabricated citations, invented numbers, a confident summary of a document nobody read. The intent is not deception. The result is unreliable output presented as reliable.

This is a competence and process failure, and it is by far the most common of the three. The response is a verification requirement, not a sanction.

3. Fabricating work that was never done. Reporting meetings that did not happen, tests that were never run, customers never contacted. AI made the artifact easier to produce, but the conduct here is dishonesty about activity, and your existing policies almost certainly already cover it.

The distinction matters because the three have completely different fixes, and treating a case of type two as a case of type three is a good way to lose a capable employee over a process gap you created.

How to tell which one you are looking at

Do not start from the document. Start from the underlying activity, because that is where the three diverge.

  • Is there independent evidence the work happened? Calendar entries, CRM records, commit history, test artifacts, email threads. Type three fails this immediately, and no amount of document analysis substitutes for it.

  • Are the errors the kind a model makes? Plausible-but-nonexistent sources, confidently wrong specifics, internally consistent detail that does not survive a single check. That signature points at type two.

  • Did they present it as verified? There is a real difference between "here is a draft I have not fact-checked" and submitting the same document as finished.

  • What happened when you asked? A type-two employee generally explains their process without much prompting. Type three tends to produce further unverifiable detail.

Resist the temptation to reach for an AI-detection tool. Detectors are unreliable in both directions, and an accusation built on one is an accusation you cannot defend if you are wrong. The evidence that matters is whether the underlying work exists, and that evidence is in your own systems. The same reasoning applies to telling whether a job listing was written by AI: the writing is a weak signal, the substance is a strong one.

The conversation

Start from curiosity about process, not from an accusation. Something close to:

This report cites two sources I could not find. Walk me through how you put it together.

That question is fair in every one of the three cases and resolves most of them in about ninety seconds. It gives a type-two employee a route to say what happened without having to confess to something they did not do, and it gives a type-three employee nowhere useful to go.

What not to do: open with "did you use AI for this." It frames tool use as the offence, which is both wrong and counterproductive, and the honest answer is often yes in cases where nothing is wrong at all.

Fix the process, because the process is usually the cause

Most of these incidents trace back to a gap that was there before anyone opened a chat window.

Nobody said what verification means. If your policy says "you may use AI tools" and stops there, it has not addressed the actual risk. It needs to say who is accountable for accuracy, which is always the person submitting the work, and what checking is expected: every factual claim, every citation, every number.

Output volume became the measure. If someone is rewarded for producing more reports, more pull requests, more updates, AI makes gaming that metric trivial. This is not an AI problem. It is a measurement problem that AI exposed. Measure outcomes rather than artifacts.

Nothing was ever verified downstream. Work that nobody checks is work whose quality nobody knows, with or without AI. The specific version of this in engineering is worth its own attention: agents that report test results they did not produce are a routine occurrence, which is why checking whether the tests actually ran belongs in your CI rather than in your trust.

There was no audit trail. When it is impossible to reconstruct what happened, every dispute becomes one person's account against another's. Keeping an audit trail of AI use makes the type-three case obvious immediately and, more usefully, makes the type-two case obvious before it reaches a customer.

What a workable policy says

Short, specific, and about accountability rather than tools. The clauses that do real work:

  1. You may use AI tools for drafting and analysis. Say this first. A policy that opens with restrictions gets routed around.

  2. You are accountable for everything you submit, regardless of how it was produced. This is the whole policy in one line.

  3. Verify every factual claim, citation, number and quote before submission. Name it explicitly. Unverified drafts must be labelled as such.

  4. Never report activity that did not occur. Meetings, tests, calls, reviews. This is a conduct rule and belongs alongside your existing ones.

  5. Approved tools, and what may not be pasted into them. Customer data, credentials, unreleased financials. See what happens to your prompts after you send them.

Writing an AI usage policy covers the full document. The single most important property is that it distinguishes using the tool from misrepresenting the result, because a policy that conflates them pushes usage underground.

The proportionate response

For type two, which is most cases: a clear verification expectation, a conversation about what checking looks like, and a review step on their next few pieces of work. Not a disciplinary process. The employee did something that a large fraction of people are currently doing, in an environment that had not told them where the line was.

For type three: your existing conduct process, applied as it would be to any misrepresentation of work. AI is a detail of the method, not a mitigating factor and not an aggravating one.

For type one: nothing. Ideally, ask them to show the team how they do it.

The broader risk picture, and where AI genuinely creates new exposure rather than accelerating old exposure, is in our overview of AI risks. This particular problem sits mostly in the second category, which is good news: you already have most of what you need to handle it.

FAQ

How do I prove an employee used AI to fake their work? Do not try to prove tool use. Prove or disprove the underlying activity using your own records: calendar, CRM, commit history, email. Whether a model drafted the document is far less relevant, and far less provable, than whether the work described actually happened.

Are AI detection tools reliable enough to act on? No. They produce both false positives and false negatives at rates that make them unsafe as the basis for a disciplinary decision, and Stanford researchers found they systematically misclassify non-native English writing as AI-generated while judging native writing correctly. Use evidence of activity instead.

Should I ban AI tools to prevent this? Banning drives usage underground and removes your visibility without removing the behaviour. A verification requirement plus an approved-tools list addresses the real risk. See shadow AI for what the ban actually produces.

Is submitting AI-written work always dishonest? No. In most workplaces it is now ordinary, and the relevant question is whether the person verified it and stands behind it. Dishonesty enters when someone claims activity that did not occur, or presents unverified output as checked.

What if the fabricated work already reached a customer? Correct it with the customer promptly and directly, then handle the internal question separately. The correction is urgent; the personnel decision is not, and doing them in that order tends to produce better outcomes on both.

How did this land?

About the author

Cecilia Iona
Cecilia Iona

Senior Editor, AI & Product

Cecilia leads the Swarmz editorial desk. She has spent a decade turning complex AI and product topics into writing people actually finish, and she owns the blog's quality bar.

Share

Get the next post in your inbox

One email a month. Product updates, engineering posts, and the best of Built with Swarmz.

I agree to receive emails about AI building tips and Swarmz product news. Unsubscribe any time.