endueendue

Jev in practice: cheaper LLM apps and faster 3D characters

TypeSafe's Jev picks one answer from options you define, in under half a second. How to use it as a router that cuts LLM costs, as the reflexes of a 3D character and as a guardrail, with diagrams and the limits to know first.

On a dark background, streams of dots flow in from the left and meet at a lavender circle in the middle. Of the three short branches leaving it, one ends in a filled green square. A dashed line arcs up to a large pink circle at the top right.

TypeSafe AI released Jev on September 15. It is the first model from the company founded by former OpenAI researcher Diogo Almeida, announced together with a $40 million funding round. Two numbers caught developers’ attention first: a response in 70 to 500 milliseconds, and $0.042 per million input tokens. Output tokens are not billed at all.

The low price follows from a narrow job. Jev picks one of the options you defined in advance and returns that choice with probabilities attached. It never writes sentences. TypeSafe calls this a System One model, after Daniel Kahneman’s name for fast, intuitive thinking.

Below is how Jev works, then three places it fits, each with a diagram: a router in front of your LLMs, fast reactions for 3D characters and NPCs, and checks on what other AI produces. The last sections cover the limits to know before you build on it.

Jev answers multiple-choice questions

Asking an LLM for a decision is like setting an essay question. The answer comes back as prose, and your code has to interpret it. With Jev you set a multiple-choice question instead. You send a state that describes the situation along with a few questions, and each question comes back as a typed answer with probabilities.

There are three kinds of question.

Question Asks Returns Example
Choice Which of these options? The chosen option, a probability per option, confidence Is this ticket for billing, technical or sales?
Score Where on this scale? A score, a probability per level, confidence How frustrated is the customer, from 0 to 2?
Noul Is this statement true? The probability of yes, from 0 to 1 Is this request urgent?

A Choice can offer up to 255 options. Questions in one request are evaluated in parallel, so adding questions barely changes the response time.

A flow diagram: a customer message as the state and three questions (which team, how frustrated, how urgent) go to Jev in one request. Back come technical with probability 0.85 and confidence 0.78, a score of 1.0, and an urgency probability of 1.0, which the code uses to route the ticket to the technical queue.
Redrawn from the quick start example in TypeSafe's documentation. The numbers are the documented response.

Because the shape of every answer is fixed in advance, Jev cannot return a malformed value. That is what TypeSafe means when it says Jev can’t hallucinate. It can still pick the wrong option, which is why the probabilities and confidence matter. Send low-confidence answers to a person or a larger model.

Jev in numbers

Item Value
Price $0.042 per million input tokens, output free
Response time 70 to 500 ms (TypeSafe’s figure, measured on the US West Coast)
Request size 64k tokens in total, 32k for the state plus the longest question
Rate limits 250,000 tokens per second, 1,200 requests per minute (may change with demand)
Input Text only. Turn images, audio and video into text first
Language Most accurate in English. Test other languages on your own data

TypeSafe published its own evaluation across four workflows, including security alert triage, invoice processing and customer service. Every model answered the same narrow questions.

Model Accuracy Cost per case Time per case
Jev 67.8% $0.0004 0.4 s
GPT-5.6 Terra 67.9% $0.0304 10.1 s
GPT-5.6 Sol 74.1% $0.0836 23.3 s
Claude Opus 5 73.1% $0.1761 37.8 s

On accuracy alone, Jev trails the largest models by five to six points. Against GPT-5.6 Terra, which scored about the same, it cost a seventy-sixth as much and took a twenty-fifth of the time.

These are TypeSafe’s own measurements. The reference answers are an average of GPT-6 Astra and Claude Fable 5.1, and no large independent reproduction has appeared yet. TypeSafe also says its evaluations were run from laptops on the US West Coast, and that only time can show whether today’s prices are sustainable.

Distance adds network time

TypeSafe’s service currently runs only on the US West Coast. On September 30 we timed unauthenticated requests to the API from Korea: 0.14 to 0.17 seconds round trip, of which the server spent 1 to 3 ms. If your users are far from California, add that to the published response times when you plan.

Use 1: a router in front of your LLMs

“Let Jev answer the easy questions and send only the hard ones to an LLM.” That is the short version of how Jev lowers LLM bills. More precisely, Jev decides where each request should be handled. Writing the answer is a later step.

  • Requests that code can handle directly, such as an order lookup, never reach an LLM.
  • Common questions get one of your prepared answers, picked with another Choice question.
  • Requests that need only a short explanation go to a small LLM.
  • Only requests that weigh several conditions or need a long new answer go to a large LLM.
  • Low confidence goes to a person.
A flow diagram: an incoming request goes through Jev. With confidence of at least 0.5 it branches four ways (handle in code, canned answer, small LLM, large LLM); below 0.5 it goes to a person or a large LLM. The decision costs about $0.00003 per request, a small LLM about $0.0002 and a large LLM about $0.0046.
The 0.5 confidence threshold follows the intent routing example in TypeSafe's documentation.

The request looks like this:

{
  "model": "jev-1.13.0",
  "state": { "message": "My bag still hasn't arrived. The order number is A-1042." },
  "questions": {
    "route": {
      "type": "choice",
      "instructions": "Where this request should be handled",
      "criteria": {
        "order_lookup": "Asks about an order or its delivery",
        "faq": "One of our prepared FAQ answers fully covers it",
        "small_llm": "Needs a short explanation, but the judgment is simple",
        "large_llm": "Needs several conditions weighed or a long new answer written",
        "human": "Refund dispute, legal issue, or strong complaint"
      }
    }
  }
}

The code that acts on the answer is short:

const { route } = res.answers;

if (route.confidence < 0.5) return sendToHuman(req); // unclear: a person decides

switch (route.choice) {
  case 'order_lookup': return lookupOrder(req);   // no LLM involved
  case 'faq':          return answerFromFaq(req); // pick a prepared answer
  case 'small_llm':    return askLLM(req, SMALL_MODEL);
  case 'large_llm':    return askLLM(req, LARGE_MODEL);
  default:             return sendToHuman(req);
}

How much it can save

The decision itself costs next to nothing. At 800 input tokens per request it is about $0.00003, or $33.60 for a million requests a month. Handling the same million requests entirely with Claude Sonnet 5.5, assuming 300 output tokens each, costs about $4,600, so the router adds roughly 0.7%.

The savings depend on how many requests finish without the large model. Suppose half end in code or a prepared answer, 35% go to GPT-6 Luna and only 15% reach Sonnet 5.5:

Setup Monthly cost (1M requests)
Everything on Sonnet 5.5 about $4,600
Routed by Jev (split above) about $804 (Jev $34, Luna $81, Sonnet 5.5 $690)

Your split will differ, so take the method rather than the numbers.

One caution. A request wrongly sent to a small model does not raise an error; the answer is simply worse. Run the router alongside your current setup for the first few weeks and compare the results.

Check the draft, escalate when needed

The order can also be reversed. A cheap model writes the answer first, Jev checks whether the retrieved sources support it, and a large model rewrites only the answers that fail. In an example OpenRouter published, this cascade gave fewer wrong answers than sending all 50 questions to GPT-6 Astra (0 against 2) at about 7% of the cost. Fifty questions is a small sample, so treat it as a sketch.

Use 2: fast reactions for 3D characters and NPCs

Speak to a character in a game or a virtual space and watch it stand frozen for a few seconds, and the moment is gone. If an LLM writes the character’s lines, those gaps are easy to create. This is where Jev’s speed pays off.

The trick is to separate the reaction from the line. The gestures and expressions a character can perform are prepared ahead of time through rigging and animation: a wave, a nod, a surprised face, a step back. Jev picks the reaction that fits the moment from that list. The animation set your rig supports becomes the options of a Choice question. The character plays the chosen reaction right away, and only when it has something to say does an LLM write the line in the background. Jev can also pick that line from ones your writers prepared.

A two-lane timeline: when the player speaks, within about 0.3 seconds the character plays the reaction Jev picked (wave, surprised face, look at player). If Jev decides a spoken line is needed, an LLM finishes it after 1 to 3 seconds and it plays with speech synthesis and lip sync. Three rules the engine keeps are listed underneath.
Timings are illustrative. Called from Korea, Jev's side also carries a 0.14 to 0.17 second network round trip.

Ask everything in one request. The body gesture, the expression and whether a line is needed all come back together.

{
  "model": "jev-1.13.0",
  "state": {
    "npc": "Hans, the village blacksmith. Gruff but kind.",
    "event": "The player walked up and asked: 'Could you fix my sword?'",
    "npc_is_busy": true,
    "relationship": "friendly"
  },
  "questions": {
    "reaction": {
      "type": "choice",
      "instructions": "Hans's immediate body reaction",
      "criteria": {
        "nod": "Acknowledges and agrees",
        "wave": "Greets the player",
        "keep_working": "Keeps hammering and glances up",
        "step_back": "Is startled or wary"
      }
    },
    "face": {
      "type": "choice",
      "instructions": "Hans's facial expression",
      "criteria": { "smile": "Pleased", "neutral": "Calm", "frown": "Annoyed" }
    },
    "needs_line": {
      "type": "noul",
      "instructions": "The player asked Hans something that needs a spoken answer"
    }
  }
}

A few rules on the engine side keep this reliable. They come from the community projects below and from a game development guide.

  1. Offer only actions that are possible right now. No key, no “open the door” option. Jev cannot break a game rule it never sees as an option.
  2. Don’t call it every frame. Call it when something changes: the player speaks, a threat appears, an action ends.
  3. Drop late answers. Tag each request with a world revision number and ignore any answer that arrives after the number has moved on.
  4. Never wait frozen. Send the request asynchronously and keep doing the current action. If no answer arrives in time, or confidence is low, fall back to a default behavior.
  5. Keep the API key out of the game build. The game calls Jev through your own server.

The cost is manageable. With a 1,000-token state a decision costs about 0.004 cents, so one NPC deciding once a second comes to about $0.15 an hour. TypeSafe’s Doom demo made 10 decisions a second for about $7 an hour. The limit of 1,200 requests a minute means a game with many NPCs should batch requests through its server.

People have already built on this pattern:

  • doom-jev turns the direction and distance of enemies, health and ammo into text and has Jev choose one of six actions. It made about 8 decisions a second with a median latency of 0.12 seconds, while code handled pathfinding and reflex moves.
  • jev-unreal-statetree adds a Jev decision task to StateTree in Unreal Engine 5.8. It sends one request when a state is entered and discards answers that arrive after the world revision has changed.
  • WorldKit is an NPC decision runtime. The engine enforces the game’s rules and Jev only chooses among the actions that are currently valid.
  • In jev-npc-interaction-prototype, every line the warehouse guard speaks is written in advance. Jev decides the guard’s action and whether the player’s story is credible, and letting the player in takes at least 0.50 action confidence and 0.55 credibility.

What about rigging itself?

Jev cannot build a skeleton or compute skin weights, and it cannot read images. What it can take on are decisions inside the pipeline that have a fixed set of answers.

Mapping bone names from characters made in different tools (mixamorig:LeftForeArm, L_lowerarm and so on) to the standard humanoid slots of your engine can be one Choice question per bone, with code keeping the parts that follow rules, such as left-right checks and hierarchy. Tagging animation clips with their purpose, like greeting, refusal or surprise, works the same way. We have not found public examples of either yet, so measure accuracy on a small sample first.

Use 3: other places worth trying

Use What you ask Jev Reference
Screening LLM input and output Is this a jailbreak attempt? How much harm would complying do? TypeSafe guardrails example
Checking agent tool calls Is this a risky call that should be blocked before it runs? LangChain harness post
Checking citations Does this quote support the claim? TypeSafe citation check example
Reranking search results How relevant is this passage to the query? Legal search example: top-1 accuracy from 5% to 18%
Moderating game chat Is this abuse or harassment? How severe? TypeSafe use case map
Voice agents Should this turn get a canned reply, a small model or a large one? Evalgent post: run it while speech recognition detects the end of the turn, so the delay stays hidden

Limits to know before you start

  • It writes nothing and explains nothing. You get probabilities only, so work that must record the reasoning behind a decision needs its own way to keep it.
  • It is weak with numbers, dates and counting. TypeSafe’s own documentation says to keep arithmetic and date comparison in code. In a game, pass “near” or “low” instead of raw distances and health.
  • It reads literally. Double negatives and questions that need several hops lose accuracy. Keep questions short and direct.
  • Long states full of unrelated detail make it less accurate. Send only the fields a question needs.
  • One plausible sentence can flip a decision. The JevOut paper, published September 24, added short, natural-looking context and turned 312 of 508 initially correct decisions (61.4%) into different answers. In 229 cases Jev gave the wrong answer a probability of 0.7 or more. Where user input goes straight into the state, double-check decisions that matter.
  • Probabilities hold across many answers. A single answer marked 0.9 is not guaranteed to be right.
  • The same input does not always give the same answer. There is no deterministic mode yet, which rules it out for multiplayer outcomes every player must see identically.
  • It is a hosted service on the US West Coast. If you need offline play or local data residency, look at open models. Kev, released under Apache-2.0, speaks the same API as TypeSafe, and its 27B model reportedly comes within about a point of Jev on unseen data (0.848 against 0.857).
  • It is still in early access. There is a waitlist and rate limits can change without notice. An alias like jev-latest moves when a new version ships, so once you have tuned thresholds, pin a version such as jev-1.13.0.

How to start

  1. Pick one recurring decision with a small set of possible answers: ticket routing, an NPC’s reaction, the risk of a tool call.
  2. Write the options precisely. Give each option one sentence on when to choose it, and add a “none of these” option.
  3. Score it on past records first. Run 100 to 200 cases where you know the right answer, see how accurate each confidence range is, and set your thresholds.
  4. Run it alongside your current setup. Compare results for a few weeks and watch for requests wrongly sent to a small model.
  5. Build the exit for low confidence. A person or a larger model in a service, a default behavior in a game.

And in endue

If you run agents in endue, the same principle carries over to model settings. Set a light, fast model as an agent’s default and switch to a stronger one only for the messages that need it. Pin a model on each routine so scheduled work keeps a steady cost. The settings are covered in Choosing a model, and an order for choosing models is in our previous post.

Sources