All articles

Jev AI: how the decision model works and where it fits

9 min read
AI Jev Automation
Three furry puppet characters sort cards against a plain light background: an orange puppet brings messages, a blue puppet considers a choice, and a green puppet points to boxes marked with a check, a question mark and a scale

Automation often needs a small decision about a piece of text. Which team should handle a request? Does a retrieved passage help answer a question? Is a customer asking for a refund? Does a product description match its category?

Jev, a decision model from TypeSafe AI, takes data and questions and returns a selected option, a score, or the probability of a positive answer. An agent system can use those values to choose its next step. That is the role described in the TypeSafe documentation.

I wanted to see how these checks might fit into existing workflows, so I tried Jev on support requests and work reports. This article explains its decisions, where they may be useful, and their limits. My own experiments come at the end, including what they failed to establish.

What distinguishes Jev

TypeSafe calls Jev a System One model, referring to quick judgments. Its practical specialization is a bounded question with a predefined answer format: choose a category, assess a property, or assign a score. Jev does not write an explanation of its choice or a reply to the user. The System One concept.

A general language model can also classify a request and return a structured result. Producing “billing team” alone therefore says little about why Jev might be useful. Vercel compares these approaches.

According to TypeSafe, Jev is fine-tuned using RLCD: reinforcement learning for calibrated decisions. The aim is to produce probabilities that match observed outcomes. If a well-calibrated model repeatedly assigns an event an 80% probability, that event should occur in roughly eight out of ten such cases. Calibration describes a group of predictions; it cannot guarantee an individual answer. TypeSafe’s account of the training method.

An agent system could use this information to handle clear and ambiguous cases differently. Calibration still needs to be checked on the data the system actually processes. My experiments did not test it.

How a Jev decision works

Consider a support message:

“My payment went through, but I cannot log in. Please help me recover access.”

Routing it requires the message, a question, and rules for the available teams. Suppose billing handles charges and refunds, technical support handles access and application problems, and anything outside those rules goes for separate review.

Under those rules, technical support is the expected destination: the customer wants access restored. Mentioning a payment does not determine the category by itself. This is an illustrative expected answer, not a measured Jev result.

The process is:

Message and context → question with criteria → Jev assessment → action by the agent system.

The input may be a single text or several related pieces: a conversation, order details, and a handling rule. TypeSafe calls this shared input state. The surrounding system supplies the evidence on which Jev makes a judgment. What state contains.

That distinction matters for a message such as “I was charged twice.” The text tells us what the customer reports. Establishing whether a duplicate charge actually occurred requires payment records. A confidently worded question cannot supply missing evidence.

After the assessment, the agent system decides what to do: assign a queue, request more information, or hand the case to a person. Model criteria and the rules for subsequent actions can be changed separately.

Three answer types

Jev offers three question types, depending on the decision required.

TypeWhat it returnsExample question
Choice A selected option and the probabilities of the options Which team should handle the request?
Noul The probability of “yes,” from 0 to 1 Is the customer explicitly asking for a refund?
Score A value on a defined scale, which can be fractional How much does the reported problem disrupt work?

Choice returns a selected option and the probabilities of the options. The system designer defines the categories and their meaning. When inputs may fall outside the list, an “other” or “insufficient information” option gives them an explicit destination. Without it, the model has to choose from categories that may all be unsuitable. Choice documentation.

Choice selects one answer. If an email asks both for a refund and for access to be restored, the system needs a rule for that combination. A mixed-request category or two independent questions about the presence of each request are possible approaches.

Noul returns the probability of a positive answer, from 0 to 1. An illustrative value of 0.9 leans toward “yes,” 0.1 toward “no,” and around 0.5 gives neither outcome a clear advantage. A low value can represent a fairly definite negative answer. Some integrations, including AI SDK, call this type boolean. Noul documentation.

Score requires descriptions of the levels on a scale. For example: a cosmetic defect; a broken feature with a workaround; or a blockage that prevents the main task. Those levels are numbered 0, 1 and 2. The result can be fractional because it is the average of the level numbers weighted by their probabilities. A score of 1.4 describes a position on that scale. It does not measure financial loss or the share of customers affected. Score documentation.

Several questions can use the same input. TypeSafe says they are evaluated independently against the shared context; an answer does not automatically become evidence for another question. The agent system combines the results, such as “refund requested” and “order details missing,” to choose the next step. Multiple questions and state.

What probabilities add

A selected option hides ambiguity. Suppose two requests both go to billing, with different probability distributions:

Illustrative examples, not measured Jev results.
OptionCase ACase B
Billing team 95% 46%
Technical support 3% 44%
Separate review 2% 10%

In the first case, one answer clearly dominates. In the second, the two teams are almost tied, although there is still a winner. Reading only the selected team would hide that difference.

For Choice and Score, TypeSafe also returns confidence, a measure of certainty calculated from the probability distribution. It is distinct from the probability of the selected option. Noul has no separate confidence value. How confidence works.

These values can support several paths through a workflow. Clear enough answers can be used automatically; ambiguous ones can go for review or receive additional context. The threshold depends on the mistakes the process can tolerate and needs testing against examples with known answers.

A high value can still accompany a wrong decision. A missing category or an ambiguous rule may produce a confident but unsuitable answer. Evaluation should therefore consider both the proportion of decisions handled automatically and the errors within that group.

Tasks worth trying

I would consider Jev when the same bounded question needs to be asked repeatedly about different texts. These are possible starting points:

Request routing

Input
A message and team handling rules
Jev decision
Which team to route it to, or whether it needs separate review

Classification

Input
A document or product description and category definitions
Jev decision
Which category to assign

Context selection

Input
A user question and a retrieved passage
Jev decision
Whether the passage helps answer the question

Claim checking

Input
A claim and its source text
Jev decision
Whether the source supports the claim

Editorial rules

Input
A text and a specific requirement
Jev decision
Whether there is a violation to flag for an editor

An agent’s next step

Input
The current situation and available actions
Jev decision
Which branch of the workflow to suggest

These directions appear in TypeSafe’s use-case map. They are candidates to test on your own data. A documented use case does not establish quality in a particular workflow.

For example, a search system may retrieve twenty passages. Sending all of them to the model that writes the answer may be unnecessary. An additional assessment can help select useful passages. Its success still depends on retrieval: evaluating the returned passages cannot recover a document the search missed.

Checking a claim has a similar requirement. If a draft says “refunds are available for a month” while the source states different terms, there is a specific comparison to make. Asking whether an entire article is true without supplying sources requires a much broader process.

Choosing an agent’s next step also needs a limited set of actions and clear conditions. Suggesting a tool does not grant permission to perform every operation it exposes. Access control and execution belong to the surrounding system.

The question needs a clear meaning

Short questions can hide the most ambiguity: “Is the task ready?”, “Is the customer unhappy?”, “Can this be closed?”

Closing a support request might require an employee to have replied, a customer to have confirmed the result, or evidence that the problem is resolved. Each interpretation needs different data. A single number cannot choose the intended policy for its author.

I would first break the decision into observable properties: whether the original question has an answer, whether the customer requested more help, and whether an action has evidence of completion. Then I would define which combination permits closure.

A rating scale needs equally clear levels. “Poor, acceptable, good” leaves plenty of room for interpretation. “Only a topic is named; the intended outcome is stated; both the outcome and its acceptance check are stated” gives people concrete differences to discuss. The model’s scores still need comparison with how people understand the scale.

A useful first candidate is a question that two people given the same data can usually agree on. If they need a call with the author, access to another system, or an investigation, the workflow should gather that information before evaluation or make it a possible next step.

Where Jev is a poor fit

Exact calculations and formal checks are easier to assign to ordinary code: calculate a total, compare dates, or check a required field. TypeSafe explicitly describes Jev 1.13’s difficulties with numbers, dates and multi-step reasoning.

Drafting an email, explaining an error, or developing a plan requires a generative step. Jev returns its predefined answer types without a detailed explanation of its judgment. That limits its role where a person needs to examine the reasoning. System One capabilities.

The current documentation describes text input. Screenshots, audio and video are not supported directly. TypeSafe also says English is the primary training language and performance in other languages is currently lower. A workflow using Russian therefore needs evaluation on its own wording. Supported input.

More context does not automatically improve a decision. The documented limitations of Jev 1.13 include sensitivity to irrelevant details and to instructions embedded in the input text. Both the relevance and the origin of supplied data matter. Jev 1.13 limitations.

What my experiments established

My main experiment concerned whether a work ticket was ready to start. I asked whether its description was sufficient, then used later clarification requests as a comparison. That proved to be a weak evaluation method: someone can start from an adequate brief, discover a new problem, and need clarification. Some labels were also produced by another LLM without manual review.

Another example involved messages in which an agent reported completion. A message containing a prepared snippet of passing test output received a high assessment for the presence of evidence. But the model received only that text, without logs from an actual run. This shows a response to the supplied text; it does not establish an ability to verify that the work happened.

These attempts helped me define the question and its required evidence more precisely. They do not establish Jev’s quality as a general verifier.

The service measurements were still useful. Three datasets contained 828 successful calls, with median durations between 383 and 477 ms. In one dataset, the 90th percentile reached 8.7 seconds: roughly one in ten successful calls took that long or longer. I also encountered rate-limit responses (HTTP 429) and server errors. The timings were measured by my client through Vercel AI Gateway, a service for accessing models. They exclude pauses my script inserted between attempts.

The successful calls used about 1.45 million input tokens. At the rate recorded in the project, $0.042 per million, the estimate is approximately six US cents. This calculation does not establish the amount actually charged. Current provider terms are available from TypeSafe and AI Gateway.

I see Jev as a candidate for frequent, small decisions in an already understood process. Classifying a message, selecting a passage, and assessing a specific property all allow the inputs and expected outcome to be described. That is where I would start: write the rule, collect real examples, and see which cases the model handles usefully and which still need review.