Decision models
Jev vs OpenAI Decisions API: when a decision model beats an LLM
First published .
Most software does not need an AI to write a paragraph. It needs the AI to pick one option, assign a score, or decide whether something is a match. Jev was built for that narrower job. OpenAI has now announced a Decisions API built on GPT-6 Luna. The new category matters because a bounded decision can be much cheaper and easier to consume than generated text.
Summary
Jev vs OpenAI Decisions API vs regular GPT-6 Luna
The cost example assumes 1,000 decisions with 1,000 total billed input tokens each. The GPT-6 Luna example also assumes 20 output tokens per decision, standard public list pricing, and no caching. OpenAI has not yet published standalone Decisions API pricing.
| Dimension | Jev 1.13 | OpenAI Decisions API | GPT-6 Luna API |
|---|---|---|---|
| Status | Public API | Limited preview | Public API |
| Published price | $0.042 / 1M input tokens; output free | Not published as of 3 Oct 2026 | $0.10 / 1M input; $0.50 / 1M output |
| Illustrative 1,000 decisions | $0.042 at 1,000 billed input tokens each | Unknown | About $0.11 at 1,000 input + 20 output tokens each |
| Input | Text only | Text and images | Text and images |
| Output | Choice, score, or yes/no probability | Finite predefined answers | Generated text or structured output |
| Quality feel | Very strong for bounded binary and choice tasks; weaker on some graded judgments | Promising but early; direct tests so far are mixed by task | Most flexible when the task needs reasoning, extraction, or explanation |
| Best fit | High-volume classification, routing, verification, matching | Bounded multimodal decisions once access and pricing are clear | Messy or open-ended cases that do not fit a fixed answer set |
Prices checked 3 October 2026. Jev pricing comes from TypeSafe documentation. GPT-6 Luna pricing comes from OpenAI's model comparison. Decisions API pricing is not yet public.
What is a decision model?
A normal large language model takes context and generates tokens. Even if the application only needs one of three labels, the model is still a general text generator. Structured output can constrain the shape, but the underlying job is generation.
A decision model starts from a smaller contract. The application defines the permitted answers first. The model reads the state and returns a choice, a score, or a probability that the answer is yes. TypeSafe calls Jev a System One model. Its public API accepts text and structured text state, then evaluates typed questions against that state.
OpenAI announced the same broad interface at DevDay on 29 September. Its Decisions API focuses GPT-6 Luna on user-defined questions with finite predefined answers. The preview also accepts images, which matters for product, interface, moderation and computer-use decisions.
Jev is cheaper, but the size of the saving depends on the workflow
TypeSafe lists Jev 1.13 at $0.042 per million input tokens and does not charge for output tokens. OpenAI lists regular GPT-6 Luna at $0.10 per million input tokens and $0.50 per million output tokens on its model comparison page.
For a tiny classification request, that is a meaningful difference but not a hundredfold one. At 1,000 billed input tokens per decision, 1,000 Jev decisions cost about $0.042. The same input through regular Luna costs about $0.10 before output. If each Luna response uses 20 output tokens, the illustrative total is about $0.11.
The larger cost claims come from workloads where the interface changes how many calls are needed. A September 2026 University of Pennsylvania study compared Jev with GPT-5.6 Luna, Gemini 3.8 Flash and DeepSeek V4.1 Flash across 5,003 rubric-criterion pairs. Jev could evaluate multiple criteria against the same state in one request. The LLM judges were called once per criterion. Across the nine evaluation panels, the LLM judges cost 29 to 325 times as much as Jev and took 30 to 220 times as long.
That is real evidence, but it is evidence about that evaluation design. It should not be turned into a blanket claim that every Jev call costs 29 to 325 times less than an LLM call.
The quality evidence is more interesting than a simple leaderboard
In the same academic evaluation, only 8 of 27 paired comparisons between Jev and an LLM judge showed a statistically clear difference in accuracy. Jev's clear leads were mostly on binary criteria. Its clear deficits appeared on graded criteria, mainly against Gemini.
That pattern makes intuitive sense. "Does this candidate refer to the same hotel?" is closer to Jev's natural job than "rate the quality of this essay from one to five using a nuanced human convention."
There is also now a small amount of direct evidence for OpenAI's Decisions API. Every tested the preview against Jev. On a text-only replay of computer tasks, Decisions selected the correct control on 76 of 78 scored steps versus 73 for Jev, and was faster in that test. On a separate conversation-thread classification test, the two were effectively tied on accuracy and Jev had lower median latency.
Those are early product tests, not a general benchmark. The useful conclusion is that neither system has established a universal quality advantage. Task definition matters more than the logo on the API.
Where Jev fits today
Jev makes the most sense when the application already knows the answer space. Routing is the obvious example: sales, support or billing. Other good candidates include spam checks, lead qualification, policy checks, entity matching, relevance filtering, product classification and deciding which tool an agent should call next.
It is particularly attractive when the same state needs several decisions. Jev ingests the state once and can evaluate multiple typed questions in parallel. That is one reason its economics can improve quickly in rubric-style or multi-signal workflows.
There are limits. Jev is text-only today. TypeSafe says English is its strongest language. It is not the right model when the output itself must be prose, code, a calculation, a summary, or an explanation that depends on open-ended reasoning.
Where OpenAI's Decisions API could be stronger
The most obvious difference is multimodality. OpenAI says the Decisions API accepts image context. That opens cases Jev cannot handle directly: choosing the right control from a screenshot, comparing visual products, classifying an image, or making a bounded decision from both text and visual state.
It is also built on Luna rather than a separate text-only classifier. That may help on decisions that need more general reasoning before the final bounded choice. The tradeoff is that the product is still in limited preview and OpenAI has not published standalone pricing or a complete public API reference as of 3 October.
Until those details are public, a production comparison has to include regular GPT-6 Luna as the practical OpenAI option. Luna already supports structured outputs and image input, so developers can implement the same application contract themselves, even if the underlying inference is not optimized specifically for decisions.
A good architecture uses rules before either model
Decision models should not replace deterministic checks that already work.
Take product matching. If two retail pages expose the same GTIN, EAN, manufacturer SKU or model number, no AI call is needed. For hotels, an exact phone number, address, property identifier, or tight geographic match can resolve many cases deterministically.
The model belongs in the uncertain middle. Two hotel names are similar but not identical. A product title has changed between merchants. One listing says "Studio Bowling Bag" and another says "Leather Studio Bowling Bag Black." The deterministic score is not high enough to accept and not low enough to reject.
That is a natural decision-model call: same entity, variant, similar item, or unrelated. Jev is a strong first choice for text-only candidate pairs because it is cheap and already public. If the distinction depends on the product image, use a visual similarity model first or test OpenAI's multimodal Decisions API when access is available.
The practical model stack
For many production systems, the cheapest model is no model. Use exact identifiers and rules first. Then use embeddings or fuzzy matching to reduce the candidate set. Send only ambiguous cases to a decision model. Keep a general LLM for the small residue that actually requires open-ended reasoning or explanation.
This changes the role of the LLM. Instead of paying a general model to make every tiny branch in the application, you reserve it for the decisions that justify its flexibility.
Jev is the clearest proof of that architecture today. OpenAI's Decisions API is a strong signal that the pattern is becoming a category rather than a one-company experiment.
