OpenAI's Decisions API, side by side with Jev
At DevDay, OpenAI previewed a Decisions API built on Luna that picks one answer from options you define and accepts images. What OpenAI has shared, the first head-to-head tests, how it differs from Jev, and what is still unannounced.

On this page
At its DevDay keynote on September 29, OpenAI previewed a Decisions API. You give it questions and a fixed set of possible answers, and it returns one of them. It is built to hand your code a usable answer quickly.
It belongs to the same family as TypeSafe’s Jev, which we covered yesterday. OpenAI shipped its own decision model about two weeks after Jev appeared, so developers paid close attention. Every went as far as saying that, of the five DevDay announcements it singled out, this one could turn out to be the biggest.
There is not much to go on yet. OpenAI’s official description is a single paragraph in its DevDay recap, and there is no price. Below is what has been confirmed so far, the first head-to-head tests, how it compares with Jev, and what to check when it opens up.
At a glance
| Item | Details |
|---|---|
| What it does | Picks an answer from questions and options you define |
| Model | Luna, the smallest and most affordable model in OpenAI’s lineup |
| Input | Context as text or images |
| Response time | About 150 ms (OpenAI, as reported by The New Stack) |
| Status | Limited preview, with a broad release “in the coming days” |
| Price | Not announced |
What OpenAI has said
This is how OpenAI’s DevDay recap describes it:
Decisions API enables real-time decision-making by focusing Luna’s intelligence on a specific set of user-defined questions with finite pre-defined answers. Developers supply context using text or images, and get back answers they can use to classify content, route requests, or choose an agent’s next action.
Simon Willison, live-blogging the keynote, noted that it works by giving the Luna model “a predefined set of options to choose from” and responds “in a fraction of a second”. The New Stack quoted OpenAI as putting the response time at 150 milliseconds, against 1.6 seconds for the same decision made with a regular GPT-6 Luna call.
Much is still missing. As of the morning of September 30 (Korean time), OpenAI’s developer docs, API changelog and pricing page had no entry for the Decisions API. An OpenAI spokesperson told The New Stack that the company plans to share more at broad rollout, so this post shows no request format.
One naming note: OpenRouter has an alpha endpoint also called the Decisions API, which serves Jev. It is a different product.
Where it fits
OpenAI names three uses:
- Classifying content: is this message spam, and what topic is it about?
- Routing requests: which team, or which model, should handle this ticket? This is the router in front of your LLMs from our previous post.
- Choosing an agent’s next action: which of the buttons on screen should it click? Every’s example of deciding whether an email needs a reply fits here too.
Image input is the most visible difference from Jev. Jev reads text only, so screens and photos first have to be turned into text by another model. With the Decisions API you could pass a screenshot as context and ask which button to press next. How well that works will only be clear once more people can try it.
The first head-to-head tests
Every published results from testing the preview side by side with Jev.
| Test | Decisions API | Jev |
|---|---|---|
| Replay of computer tasks (text only), correct steps out of 78 | 76 | 73 |
| Typical response time on that test | 230 ms | 500 ms |
| Accuracy classifying conversation threads | Effectively tied | Effectively tied |
| Median response time on that test | 309 ms | 161 ms |
The first two rows come from Every’s Jack Cheng, the last two from Kieran Klaassen of Cora. Each test favored a different side. Every’s summary: the Decisions API did better than Jev on speed and accuracy in a few early tests but trailed on its broader evaluations.
The samples are small and the conditions were not published in detail. Treat the results as a reason to run your own comparison rather than as a verdict.
Side by side with Jev
Both return one of a set of predefined answers and are built for decisions only. With what is public today:
| Item | OpenAI Decisions API | TypeSafe Jev | Source |
|---|---|---|---|
| Status | Limited preview, broad release soon | Early access (waitlist) | OpenAI recap, TypeSafe |
| Input | Text, images | Text only | OpenAI recap, TypeSafe docs |
| What comes back | The chosen answer with a confidence score | The chosen answer, a probability per option, confidence | The New Stack, TypeSafe docs |
| Price | Not announced | $0.042 per million input tokens, output free | The New Stack, TypeSafe docs |
| Response time | About 150 ms (OpenAI) | 70 to 500 ms (TypeSafe) | The New Stack, TypeSafe |
| Option limit | Not announced | 255 per Choice question | The New Stack, TypeSafe docs |
| Where it runs | Not announced | US West Coast | TypeSafe |
The biggest differences are image input and a published price. Teams already on the OpenAI API may like not adding another vendor. Until the price and limits are out, though, nobody can work out what it costs to run.
What to keep in mind
- Price is the deciding factor. Decision calls are small and frequent, so the price decides where they make sense. Every’s view is that if the API is cheap enough for thousands of little decisions, it could be one of DevDay’s biggest releases.
- Option limits and tuning are unknown. The New Stack points out that it is unclear how many candidate answers one request can take and whether developers can tune it on their own data.
- There is no evidence on calibration yet. The value of a decision model is a probability you can trust, and OpenAI has not published a calibration evaluation for this API. Score it on records where you know the right answer before you set thresholds.
- Speed depends on the setup. OpenAI says 150 ms; Every’s two tests saw 230 ms and 309 ms. OpenAI has not said where it runs, so measure from where your users are. When we measured Jev from Korea, the network round trip alone added 0.14 to 0.17 seconds.
- The launch timing differs by source. OpenAI’s recap says “in the coming days”; Every wrote “in a few weeks”.
- Resistance to manipulated context is unknown. Research on Jev showed that a few plausible sentences can flip its decisions. Nobody has yet shown how the Decisions API holds up to the same trick.
And in endue
The principle from our previous post holds whichever decision model you choose: make the cheap, fast decision first and call a large model only when needed. In endue you can apply it by setting a light model as an agent’s default and switching to a stronger one only for the messages that need it. The settings are covered in Choosing a model.
Sources
- OpenAI, DevDay 2026 Recap (September 29, 2026; Decisions API section)
- The New Stack, OpenAI answers TypeSafe’s Jev with a Decision API built on Luna (September 29, 2026)
- Every, Vibe Check: OpenAI DevDay 2026 (September 29, 2026)
- Simon Willison, OpenAI DevDay 2026 live blog (September 29, 2026)
- OpenRouter, Jev documentation
- TypeSafe, Models and Introducing System One Models & Jev


