Dashboard

AI Coding Agents in a Language They Barely Know

An agent that is excellent in Python can be quietly terrible in Elixir or COBOL. The failure is not random: it is idioms, standard library recall, and invented APIs.

Carlo Zuercher
Carlo Zuercher
Staff Engineer, Platform
3 September 20261 min read

Coding agents are not uniformly good at programming. They are good at the languages that dominate their training data and noticeably worse everywhere else, and the degradation is not graceful. In a language the model knows well, mistakes look like bugs. In a language it barely knows, mistakes look like fluent, idiomatic-seeming code that calls functions which do not exist. That second failure mode is far more expensive, because it survives review by anyone who is also new to the language.

What actually degrades, in order

Standard library recall goes first

The model knows the language has a way to do the thing and produces a plausible name for it. List.first_or_default, Stream.chunk_while_reduce, functions that sound exactly right and do not exist. This is the same mechanism behind hallucinated package names, one level down.

Idioms go second

The code compiles and runs, but it is transliterated from another language. Elixir written like Ruby, Rust written like C++, Go written like Java with goroutines bolted on. It passes tests and every reviewer who knows the language will hate it, for reasons that turn out to be operational rather than aesthetic: the non-idiomatic version usually fights the language's concurrency, error or memory model.

Ecosystem knowledge goes third

Which library people actually use, which one has been abandoned, what the current build tool is. Models default to whatever was most discussed during training, which in a smaller ecosystem can easily be five years stale. See also why agents keep recommending deprecated packages.

Version awareness goes last and hurts most

In a small ecosystem, one major release can invalidate most public examples. The model averages across all of them and produces code that matches no single version.

Rough guide to where you are

Tier

Examples

What to expect

Abundant

Python, JavaScript, TypeScript, Java

Reliable idioms, occasional version drift

Solid

Go, Rust, C#, PHP, Ruby, Swift, Kotlin

Good code, weaker on newer library APIs

Thin

Elixir, Clojure, Scala, Haskell, R, Julia

Idiom drift and invented standard library calls

Sparse

COBOL, Ada, Fortran, Zig, Nim, OCaml

Fluent-looking output, verify every symbol

Proprietary

Internal DSLs, vendor scripting languages

Treat as unknown, ground everything

The tiers are a rule of thumb from observed behaviour, not a benchmark, and they shift with each model release. The useful part is the shape: the cliff is between solid and thin, and it arrives faster than most people expect.

The grounding setup that recovers most of it

None of this requires a better model. It requires giving the model the things it is missing, which are all things you already have.

  1. Provide a working reference file from your own codebase, in the target language, doing something structurally similar. One good example moves output quality more than any amount of prompt instruction, because it fixes idioms and library choice at the same time.

  2. Paste the actual API surface. For the modules the task touches, include the real function signatures from the installed version, not a description. This is the single fix for invented functions.

  3. Pin the version in writing. "Elixir 1.18, OTP 27, Phoenix 1.8" at the top of the task. Vague version context is where stale examples enter.

  4. Force a compile and run loop. The agent must build and test after every change, with the command written down. In a sparse language this is the only check that catches invented symbols quickly, and it converts a review problem into a fast machine-checked one.

  5. Ask for smaller diffs. Reduce the unit of work until each change is small enough to verify by reading. Agent confidence does not fall when its competence does, so you supply the caution.

Steps two and four together eliminate most of the damage. The rest is idiom quality, which step one handles.

A prompt frame that works

text
Language: Elixir 1.18, OTP 27. Project uses Ecto 3.12.

Below are (a) a reference module from this codebase that shows our
conventions and (b) the exact function signatures available in the
modules you may call.

Rules:
- Use ONLY functions that appear in the signatures below or in the
  reference module. If you need something else, stop and say so.
- Match the style of the reference module, including error handling.
- After writing, run `mix compile --warnings-as-errors && mix test`
  and fix anything it reports before returning.

REFERENCE MODULE:
...

AVAILABLE SIGNATURES:
...

The "stop and say so" clause is worth more than it looks. Given permission to admit a gap, models will use it. Without that permission the same gap becomes an invented function, because producing something is the default behaviour.

When to not use an agent here

If nobody on the team can read the target language, an agent is a liability rather than a shortcut. The output is fluent enough to pass a review by someone who cannot evaluate it, which is worse than obviously broken code. In that situation use the agent for the parts you can verify, which is usually tests and glue, and get a human who knows the language for the core.

Legacy work is the common version of this, and the reading-first approach in how to get an AI coding agent to explain a legacy codebase applies directly. A useful first task in any unfamiliar language is asking the agent to write tests for existing code, because a wrong test fails loudly and a wrong implementation does not. More broadly, our AI coding tools pillar covers the general workflow.

Frequently asked questions

Does giving the agent the documentation URL help?

Only if it can actually fetch and read it. Pasted signatures beat a link every time, because a link is a promise and a signature is a fact in the context.

Are local or open-weight models worse at rare languages?

Generally yes, and for the same reason: less training data overall means the thin tail is thinner. The grounding steps matter more, not less.

Will fine-tuning fix this?

It can, given enough of your own code in the target language, but it is a large investment to solve a problem that a reference file and real signatures mostly solve for free. Try grounding first.

How do I know which tier my language is in?

Ask the model to write a fifty-line idiomatic module using the standard library, then check every function it calls against the real docs. The count of invented symbols tells you immediately.

How did this land?

About the author

Carlo Zuercher
Carlo Zuercher

Staff Engineer, Platform

Carlo works on the platform that turns prompts into running apps. He writes the engineering deep dives and the changelog notes worth reading.

Share

Get the next post in your inbox

One email a month. Product updates, engineering posts, and the best of Built with Swarmz.

I agree to receive emails about AI building tips and Swarmz product news. Unsubscribe any time.

AI Agents in a Language They Barely Know | swarmz.net