Dashboard

Should You Turn Off Codebase Indexing? A Risk Guide

Codebase indexing is on by default in most AI coding tools. Whether that is fine depends on three questions almost nobody asks first.

Carlo Zuercher
Carlo Zuercher
Staff Engineer, Platform
24 September 20261 min read

Codebase indexing is the feature that lets an AI coding tool answer questions about code you have not shown it. To do that it has to read your repository, and depending on the vendor it either reads it locally, sends fragments on demand, or uploads the lot. Those three are marketed with the same word and are not remotely the same decision.

The question is worth asking this month because of what developers found in Z.ai's ZCode client. According to an investigation by a developer who traced the behaviour, the tool's default-on indexing packaged entire workspaces, including full git history, git LFS cache, reflogs and global application configuration, and uploaded encrypted archives. Tom's Hardware reported hundreds of megabytes of workspace data and 564 attempts to send a 313MB archive, with no consent request and no mention in the privacy policy. Z.ai has since disabled the feature, shipped a fix and published the client source.

Treat that as a worked example rather than a verdict on one vendor. The interesting question is what you would have known if it had been your tool.

The three things called indexing

Model

What leaves your machine

Practical risk

Local index

Nothing, or embeddings stored on disk

Low. Disk space and CPU

On-demand retrieval

The snippets sent with each request

Moderate, and proportional to use

Full remote index

A copy of the repository, held server-side

High, and persistent

The third is the one that needs a decision. It is also the one that gives the best answers, because the model can search everything rather than what the client guessed was relevant. There is a real capability being traded here, which is why "just turn it off" is not automatically the right advice.

What makes the full remote index different is persistence. On-demand retrieval exposes what you were working on. A remote index exposes what you have, continuously, and long after the session ends.

What gets swept up that you did not picture

When people imagine indexing they picture their source files. The uncomfortable part of the ZCode reports is the rest of it.

  • **Git history.** Your current tree may be clean. Your history frequently is not. A key committed in March and removed in April is still in the history, and an index built from the repository rather than the working tree has it.

  • **Reflogs and dangling objects.** Branches you deleted, commits you amended away, that stash from the panicked Friday. All still in .git.

  • **Untracked local files.** .env, scratch scripts with real credentials, a customer CSV you were debugging with. Gitignored files are not sent to your remote, which is exactly why people leave sensitive things in them, and an indexer reading the directory does not care what git thinks.

  • **Global configuration.** Application config outside the project entirely, which is the detail in the ZCode reports that most exceeded what users expected.

Run git log -p | grep -iE '(api[_-]?key|secret|password|token)' | head on a repository you consider clean. The result decides this question more honestly than any vendor's security page.

The decision rule

Three questions, in order.

  1. **Is there anything in the git history you would not paste into a stranger's issue tracker?** If yes, a full remote index is off until the history is cleaned, regardless of vendor. This is not about trust, it is about blast radius when the vendor has an incident.

  2. **Can you scope what it reads?** Most tools support ignore files for indexing specifically, separate from .gitignore. Excluding .env*, secrets/, fixture data and anything matching a credential pattern converts a large decision into a small one. Check that the ignore rules apply to indexing and not only to context, since they are frequently different systems.

  3. **Does the answer quality justify it on this repository?** On a small codebase the model can hold the relevant files in context anyway and the index adds little. On a large legacy codebase it is the difference between a useful assistant and an expensive autocomplete. The tradeoff scales with repository size, which means the correct answer differs per project rather than per person.

A reasonable default for most teams: indexing on for open source and greenfield work, off or strictly scoped for anything with client data, production credentials in its past, or a compliance obligation attached.

What to check in the tool you already use

Four checks, none of which take long.

  • Find the setting and read what it says it uploads. If the documentation describes scope in marketing terms rather than naming directories, treat that as an answer.

  • Look for an indexing-specific ignore file and confirm it is honoured. Add one whether or not you turn indexing off.

  • Check whether you can delete a remote index, and whether deletion is verifiable. In the ZCode case the decryption key sat only on vendor servers, so affected developers could neither see what had been taken nor remove it themselves. Retention you cannot verify is retention you should assume is permanent.

  • Watch the network once. A few minutes of outbound traffic observation tells you more than any policy document, and it is the method that found this incident.

The broader question of how much context a tool needs to be useful, which is the other side of this trade, is covered in how much of your codebase an AI coding agent should see. For what routinely leaves your machine even without indexing, what your AI coding assistant sends to the vendor is the baseline.

The part that generalises

This is the second supply-chain-shaped problem in AI coding tools in a month, after a plugin update shipping behind users. The pattern is not malice, it is that these tools run with full filesystem access and ship features that are on by default, and the gap between what a feature is called and what it does is where the surprises live.

Two habits cover most of it. Read the defaults when you install, not after an incident. And keep credentials out of repositories entirely, so that the answer to "what did it upload" is boring no matter which tool asked. When something does slip through, what to do if an AI coding agent commits a secret to your repo is the recovery path, and it starts with rotation rather than with deletion.

Questions

Does turning off indexing make the assistant much worse?

On a small or well-structured repository, barely. The model can hold the relevant files in context and the index adds little. On a large or unfamiliar codebase the gap is real, because without an index the tool can only reason about files it already guessed were relevant. Test it on your own repository for a day rather than accepting either claim.

Is a locally stored index safe?

Considerably safer, since nothing leaves the machine. It still creates a searchable copy of your code on disk, which matters for a shared or managed device, and some tools store local embeddings while still sending matched snippets upward. Confirm which of the two you have rather than assuming local storage means local processing.

Will an ignore file definitely stop a file being uploaded?

Only if the tool applies it to indexing specifically. Many tools have separate rules for context and for indexing, and .gitignore governs neither. Add the exclusion, then verify with the tool's own index inspection command if it has one.

My repository is already indexed. What now?

Rotate any credential that has ever been committed, on the assumption it has been copied. Then look for a delete or purge option and use it, while noting that deletion you cannot verify is a request rather than a guarantee. Rotation is the control that does not depend on the vendor doing what it said.

Is this only a problem with smaller vendors?

No. The mechanism is default-on features with full filesystem access, which every major AI coding tool has. Vendor size affects how quickly an incident is found and fixed, not whether the architecture allows it.

How did this land?

About the author

Carlo Zuercher
Carlo Zuercher

Staff Engineer, Platform

Carlo works on the platform that turns prompts into running apps. He writes the engineering deep dives and the changelog notes worth reading.

Share

Get the next post in your inbox

One email a month. Product updates, engineering posts, and the best of Built with Swarmz.

I agree to receive emails about AI building tips and Swarmz product news. Unsubscribe any time.