Dashboard

How to Benchmark an AI Coding Agent on Your Codebase

A model that tops the public leaderboard can still underperform on your codebase. Here is how to build a benchmark from your own issue history instead.

Carlo Zuercher
Carlo Zuercher
Staff Engineer, Platform
22 September 20261 min read

Public benchmarks tell you how a coding model performs on curated problems from a benchmark's own test set. They tell you almost nothing about how it will perform on your actual codebase, with your actual conventions, your actual dependency versions, and your actual definition of what a correct fix looks like. If you are choosing between models for daily use, the benchmark that matters is the one you build yourself from your own repository.

Why Public Leaderboards Mislead You Specifically

A model's SWE-bench or similar leaderboard score reflects performance on a specific, public set of GitHub issues, mostly from popular open-source projects with clean test suites and conventional structure. Your codebase almost certainly does not look like that: it has internal libraries no model has seen in training, naming conventions specific to your team, and architectural decisions that only make sense with context the benchmark cannot provide. A model that tops the public leaderboard can still perform worse than a lower-ranked model on your specific stack, because the leaderboard measures general capability, not fit to your codebase.

There is also a subtler problem: widely-used public benchmarks eventually leak into training data, directly or through discussion of their problems online, which inflates scores in ways that do not reflect genuine improvement on novel problems. A benchmark you built yourself, from issues no model has ever seen, does not have this contamination risk.

Building a Golden Set From Your Own History

The raw material for a personal benchmark is sitting in your own issue tracker and git history: real bugs that were fixed, real features that were built, with a known-correct outcome you can grade against. Pull 15 to 25 closed issues or merged PRs that represent the range of work you actually do, not just the interesting ones. Include boring maintenance tasks alongside the impressive refactors, because a model that only gets evaluated on hard problems tells you nothing about whether it handles the 80% of routine work reliably.

markdown
Selection criteria for the golden set:
- Spans your typical task types (bug fix, small feature, refactor,
  test writing) in roughly the proportion you actually encounter
- Includes at least a few tasks that require touching more than
  one file, since single-file changes under-represent real work
- Excludes anything where the "correct" fix is genuinely
  ambiguous or was itself later reverted or revised
- Has a clear, gradable definition of done: tests that must pass,
  or a diff you can compare against

That last point, a clear and gradable definition of done, is what turns a folder of old tickets into an actual benchmark rather than a vague vibe-check.

Running the Same Task Through Every Model Under Evaluation

For each item in your golden set, give every model candidate the identical starting point: the same repository state (checked out before the real fix was applied), the same issue description, and the same available tools or context window. Consistency here matters more than sophistication; a benchmark where one model got extra context or a clearer prompt than another is not comparing models, it is comparing prompts.

Grade each attempt on a fixed rubric rather than a gut reaction to the diff:

Axis

What it captures

Correctness

Does it actually fix the issue, verified by the real test suite or a manual check against the known-good fix

Scope discipline

Does it touch only what is relevant, or does it rewrite unrelated code along the way

Convention fit

Does the output match your codebase's existing patterns, naming, and structure without explicit instruction to

Time to usable

How much cleanup or follow-up prompting was needed before the change was mergeable

Scope discipline and convention fit are the axes public benchmarks almost never measure, and they are frequently the ones that determine whether a model is pleasant or exhausting to work with day to day, independent of raw correctness.

The Trap of Tuning to Your Own Benchmark

Once you have a golden set, resist the temptation to re-run it every time a model updates and treat small score changes as meaningful. A 25-item benchmark has enough noise that a one or two item swing is not a real signal, and chasing every point release across every provider costs more time than it saves. Re-run the full comparison when a genuinely new model generation ships, or roughly quarterly, not on every minor version bump.

Also resist writing your golden set's issue descriptions in a way that happens to match how you prompt models day to day. The point is to measure how a model performs on your real problems, not to measure how well a model responds to your specific prompting habits, which is a different and less useful thing to optimize.

What This Replaces, and What It Does Not

A personal benchmark replaces "which model is best" with "which model is best for us," which is the actually useful question. It does not replace ongoing spot-checking of new work, since a benchmark run quarterly cannot catch a model's performance drifting on your codebase as it evolves between evaluation rounds. Treat it as a periodic, deliberate check, not a substitute for noticing when day-to-day output quality changes.

Weighing Cost Against the Score

A model that scores five points higher on your golden set but costs three times as much per task is not automatically the right choice, and a raw quality ranking without a cost axis will lead you there anyway. Record token usage and wall-clock time alongside the correctness score for every run, then look at the result as a cost-per-successfully-completed-task figure, not just a quality score in isolation. For high-volume, low-stakes tasks (routine refactors, boilerplate generation), a slightly less capable but meaningfully cheaper model often wins on this combined measure. For the small share of genuinely hard tasks, a more expensive model that fixes something a cheaper one cannot touch at all is worth the premium, since paying more for a completed task beats paying less for one you still have to finish by hand.

Frequently Asked Questions

How many golden-set items are enough to trust the result?

Fifteen is a reasonable floor for a rough signal; twenty-five to thirty gives you enough spread across task types to trust a difference between two models that is not huge. Below ten, treat any comparison as anecdotal rather than a real benchmark.

Should I include tasks the model will clearly fail, to test failure handling?

A few, yes, but do not overweight them. How a model handles a task it cannot solve, does it say so clearly or produce a confident-looking wrong answer, is genuinely useful information, but a benchmark that is mostly edge cases stops reflecting your typical workload.

Is this worth doing for a solo developer, or only for teams?

It is worth doing at almost any scale, since the setup cost is a few hours and the alternative is choosing a model based on marketing claims or public leaderboard scores that, as above, may not reflect your actual work at all.

Once you have a model you trust from your own benchmark, writing test cases that actually catch bugs is the natural next skill to pair it with. For the specific question of whether a faster, pricier model is worth it once you have real comparison data, see is a faster AI coding model worth double the price. If you are choosing between providers generally rather than just model versions, how to choose between Claude, GPT, and Gemini for coding covers that broader decision, and AI coding tools: how to pick the right one is the fuller guide this all sits inside.

A closely related methodology, built around time-to-working-code rather than a pass/fail rubric, is how to benchmark AI coding agents on your codebase.

How did this land?

About the author

Carlo Zuercher
Carlo Zuercher

Staff Engineer, Platform

Carlo works on the platform that turns prompts into running apps. He writes the engineering deep dives and the changelog notes worth reading.

Share

Get the next post in your inbox

One email a month. Product updates, engineering posts, and the best of Built with Swarmz.

I agree to receive emails about AI building tips and Swarmz product news. Unsubscribe any time.