OpenAI Decisions API: What the Preview Does
OpenAI announced a Decisions API at DevDay that returns one of your predefined answers instead of free text. Here is what is confirmed, what is reported, and what is still unknown.
OpenAI's Decisions API is a new endpoint, announced at DevDay on 29 September, that does not write free text. You hand it an input and a closed list of possible answers, and it picks one. It runs on GPT-6 Luna and is in limited preview, with broader availability promised "in the coming days", according to The Decoder's DevDay coverage.
What the Decisions API does
OpenAI's developer account described it as a way to define questions and possible answers to classify content, route requests, or choose an agent's next action. The DevDay keynote roundup from Runtime Wire adds that inputs can be text or images.
Some coverage says the endpoint returns the chosen option together with a confidence score, for example opentools.ai's write-up. I could not confirm the response schema against OpenAI documentation this run, so treat the confidence score as reported, not documented.
The speed claim
The Decoder relays OpenAI's figure that the Decisions API makes decisions ten times faster than GPT-6 Luna does through the regular API. The opentools.ai piece puts numbers on it: about 150 milliseconds against 1.6 seconds for a standard Luna call. Both are vendor-presented figures. I found no independent benchmark.
The speed matters because classification sits in front of everything else. If a message has to be labelled before anything useful happens, the label step adds its delay to every request.
Where it fits in an app
A constrained picker is useful wherever your code already branches on a label.
Job | Closed list of answers | What your code does with it |
|---|---|---|
Inbox routing | support, sales, billing, other | Sends the message to the right queue |
Agent step choice | search, answer, ask a human | Runs the matching tool |
Upload check | receipt, contract, photo, other | Starts the matching workflow |
Do not confuse this with model routing, which picks which model answers. Here the thing being routed is the request itself. You can already do the same job with a general model and a careful prompt, as in our guide to prompting AI to triage support tickets. The difference is that a constrained endpoint cannot wander off the list. Screening user content is a common first use, and our guide to adding content moderation to an AI-built app shows where a classifier slots in.
It also pairs naturally with agents. OpenAI's Agents API runs multi-step work, and the "which step next" decision inside such a loop is exactly the kind of small, closed choice the Decisions API targets.
Designing the answer list
A closed list only helps if the list is good. Three habits from classifier work in general, not specific to this API:
Keep the answers mutually exclusive. If a message can honestly be both "billing" and "support", your routing will look random, so either merge them or define a tie-break rule in your code.
Always include an "other" or "unclear" answer. Without one, the picker is forced to choose something wrong for every input you did not anticipate.
Keep the list short. Three to seven answers is easy to audit. Thirty answers usually hides a hierarchy that should be two decisions, not one.
Free-text models fail in a different way: they add a sentence of explanation, invent a label you never offered, or change formatting between calls. A picker that can only return one of your options removes that whole class of bug, which is most of the practical appeal.
What is still unknown
Pricing. None of the sources I read this run listed a price.
Who gets preview access and when general availability lands beyond "the coming days".
Accuracy against simply prompting a general model. No independent comparison has appeared yet.
How it behaves when none of your answers fit, which is where real classifiers usually fail.
What to do with it now
Do not rebuild around a preview. If you classify inputs today, the useful preparation is a labelled test set of 100 to 200 real messages, so any future swap can be judged on your data instead of on launch-day claims. The same DevDay produced plenty of other changes worth tracking, collected in how to keep up with AI news, and the model behind this one has its own entry in our GPT-6.1 Sol launch report.
How did this land?
About the author

Senior Editor, AI & Product
Cecilia leads the Swarmz editorial desk. She has spent a decade turning complex AI and product topics into writing people actually finish, and she owns the blog's quality bar.


