What Is an OpenAI-Compatible API?
An OpenAI-compatible API copies the request shape, not the behaviour. What compatibility guarantees, the eight places it breaks, and how to test it.
What Is an OpenAI-Compatible API?
An OpenAI-compatible API is one that accepts the same HTTP request shape as OpenAI's Chat Completions endpoint, so code written against the OpenAI SDK works against a different provider by changing two values: the base URL and the API key. It is a promise about the envelope, not about what is inside it.
Almost every model provider ships one now. DeepSeek, Sakana's Fugu, most open-weight hosting services, and every local runtime worth using all expose the same endpoint shape, because it is the closest thing the industry has to a standard.
from openai import OpenAI
client = OpenAI(
api_key=os.environ["OTHER_PROVIDER_KEY"],
base_url="https://api.other-provider.com/v1",
)
resp = client.chat.completions.create(
model="their-model-name",
messages=[{"role": "user", "content": "Hello"}],
)That is the entire value proposition. Two lines, and your integration points somewhere else.
What compatibility actually guarantees
In practice, providers reliably implement a narrow core: the /v1/chat/completions path, the messages array with role and content, the model string, temperature and max_tokens, and a response object with choices[0].message.content in the place you expect it. Streaming over server-sent events is usually there too.
If your application does nothing but send messages and read text back, compatibility holds and swapping providers is genuinely a config change. Most prototypes are in this category, which is why the promise feels solid right up until it does not.
Where it breaks
Compatibility is a spectrum, and nobody publishes where on it they sit. These are the places to look first.
Area | What typically goes wrong |
|---|---|
Tool calling | Supported, but with a different schema dialect, different strictness about JSON Schema, or no parallel tool calls |
Structured output | response_format and JSON schema enforcement are frequently missing or best-effort rather than guaranteed |
Roles | system versus developer role handling differs, and some providers silently merge or drop a system message |
Streaming chunks | The delta shape varies, especially for tool calls, which breaks parsers rather than connections |
Usage fields | Token counts may be absent, differently named, or exclude reasoning tokens you are being billed for |
Sampling params | seed, logprobs, top_k, presence_penalty and stop are commonly accepted and then ignored |
Provider extras | Reasoning effort, thinking budgets and cache controls sit outside the spec, so you lose them on the way out |
Errors | Status codes, error body shape and rate limit headers rarely match, so your retry logic is the first thing to fail |
The failure mode worth internalising: a parameter that is accepted and ignored produces no error and a quietly different result. You will not find that in your logs. You will find it in your output quality a week later.
A swap checklist
Before you treat a provider as a drop-in, run these against their endpoint rather than their marketing page.
Send a request with every parameter you rely on and diff the response against the same call to your current provider. Look for accepted-and-ignored, not just rejected.
Exercise one tool call end to end, including a parallel call if you use them, and read the raw streaming chunks rather than your SDK's parsed view.
Force an error. Bad key, oversized request, rate limit. Confirm your retry and backoff code still recognises what it sees.
Check the
usageblock exists and that its numbers reconcile with the provider's billing dashboard.Ask what happens to the parameters they do not support. Silently dropped is the common answer and the one you need to design around.
Pin the model string and find out their policy for changing what it points at.
That last one matters more than it sounds. A stable model name pointing at a changing model is normal now, and an OpenAI-compatible envelope makes the swap completely invisible to your code.
Compatible is not the same as equivalent
The wire format being identical tells you nothing about prompt behaviour. The same prompt across two compatible endpoints can differ in instruction following, refusal behaviour, formatting habits, and how it handles a long system message. Portability of code is not portability of results, and treating them as the same thing is how a config change turns into a quality regression. How to write prompts that work across AI models covers the part the standard does not.
FAQ
Is OpenAI-compatible an official standard?
No. There is no specification and no certification. It is a de facto convention: providers copy the shape of OpenAI's Chat Completions API because the SDKs and tooling already speak it. Each provider decides how much to implement, and DeepSeek's API documentation is a good example of a provider stating its deviations plainly.
Does it work with local models?
Yes, and this is where it earns its keep. Most local serving runtimes expose an OpenAI-compatible endpoint, so the same application code can point at a hosted API or at something on your own machine. See how to run an AI coding model locally.
Can an orchestration service be OpenAI-compatible?
It can, and several are. Sakana's Fugu presents a pool of coordinated models behind a single compatible endpoint. Your code sees one model. What answers you is a routing decision you do not control.
Should I use the OpenAI SDK for everything then?
For a prototype, yes. For production, wrap it in a thin layer of your own so provider quirks live in one file rather than scattered through your app. It costs an hour and it is what makes a real switch possible later. How to migrate from one AI model to another covers the rest of that move, and the mechanics of how models produce answers covers what sits behind the endpoint.
How did this land?
About the author

Staff Engineer, Platform
Carlo works on the platform that turns prompts into running apps. He writes the engineering deep dives and the changelog notes worth reading.


