What Is Jev AI: The Model That Writes Nothing (and Costs 100x Less Because of It)
Everyone spent three years trying to get LLMs to return valid JSON. TypeSafe looked at the problem and asked a different question: what if the model simply couldn't write? The answer is called Jev, and understanding what Jev AI is means understanding a class of model that didn't exist six days ago.
On September 15, 2026, TypeSafe AI announced Jev as the first System One model. It doesn't complete text. It takes your program's state plus a set of typed questions and returns, in a single parallel pass, the chosen option, the score on your scale, and the probability of a yes. No prose, no parser, no try/catch wrapped around json_decode.
In this post: what Jev actually is, why it can't hallucinate (and why that doesn't mean it can't be wrong), the real math on pricing, and Python code that actually runs, calling all three primitives.
TL;DR
- What it is: the first System One model. State goes in, a typed decision comes out. Zero text generation.
- Stack:
jev-1.13.0viaPOST /v1/systemone, official SDKs for Python and JavaScript. - Cost/Access: $0.042 per million input tokens. Output is free. Requires a
TYPESAFE_API_KEY. - Latency: 70ms to 500ms end to end, versus 3s to 329s for frontier models in TypeSafe's own comparisons.
- Useful link: docs.typesafe.ai and
pip install typesafe-sdk.
What Jev AI is (and why a System One model can't hallucinate)
The name comes from Kahneman. System 1 is fast, intuitive judgment, the kind you make in one second just by looking at something. System 2 is slow, chained, deliberate reasoning. The LLMs we've been using since 2023 are System 2 machines improvising System 1: you ask for a three-way classification and the model builds a chain of reasoning, writes some JSON, and you pray the key comes back with the right name.
Jev flips the contract. The official documentation calls it a "frontier-intelligence function call": unstructured state goes in, a typed probabilistic decision comes out. You send a state (a string, a JSON object, an array of texts) and a dictionary of questions. Each question is evaluated in parallel and in isolation against the same state, and comes back as a value your code consumes directly.
Now the part that matters: the guarantee is a type guarantee, not a prompt guarantee.
In an LLM, "always answer with one of the three options" is an instruction. An instruction that gets followed very well in 2026, but still an instruction: the model has its entire vocabulary available at every token and you're asking for goodwill. If you want to understand why that's structural and not vendor laziness, start by understanding what an LLM does at every token — and then look at the layers we stack to tame structured output.
In Jev, the output space is the set you enumerated. The model returns a probability distribution over your options and nothing outside them. It has no way to invent a fourth category, no way to return "tecnico " with a trailing space, no way to break the schema in the middle of a traffic spike. The docs are blunt about it: "every answer is constrained to the options you supplied".
And here's the caveat the marketing material leaves out: not hallucinating is not the same thing as being right. Jev can't hand you an invalid value. It can, with all the calm in the world, hand you the wrong value with 0.94 confidence. The difference is that its error is a decision error, which you measure with evals and handle with thresholds, instead of a format error, which you handle with retries and defensive parsing. That's an entire class of bug leaving your codebase. But the line between "can't make things up" and "can't be wrong" is thin, and getting it wrong doesn't blow up on deploy Friday: it blows up three months later, as tech debt in the ticket nobody can explain.
And if you read that paragraph thinking "great, one more model category for me to keep up with," the problem isn't your pace: it's trying to keep up alone in a field where an entire class of model is born in six days. Environment solves what discipline can't, and that's what Clã Beer and Code is: every week, live, people filtering the noise together and putting architecture decisions on the table before they turn into incidents. It's paid, it's a subscription, and it's the same kind of conversation this post has in writing.
What changes when you remove autoregression: 70ms, $0.042/MTok, and free output
Autoregression is what makes an LLM expensive and slow: every output token depends on the previous one, so the response is serial by construction. Take generation out of the equation and three things change at once.
Latency. TypeSafe publishes 70ms to 500ms end to end, and presents that as 40x to 200x faster than frontier models on the same task. Third-party measurements put the p50 for a single question at around 236ms to 276ms, which is consistent with the advertised range. In practical terms: it fits inside an HTTP request without you needing a queue, a job, and a webhook.
Price. jev-1.13.0 costs $0.042 per million input tokens. Output costs zero, and TypeSafe writes "too cheap to meter" in the table, which is honest: there's no output to meter.
The math. Let's run the numbers on a concrete case, because "100x cheaper" is a headline number and headlines don't pay invoices. Assume 1 million classifications per month, 300 tokens of state in each one:
| Scenario | Input | Output | Total/month |
|---|---|---|---|
| Jev ($0.042/MTok, free output) | $12.60 | $0.00 | $12.60 |
| Frontier ($10/$50 per MTok, 60 tokens of JSON back) | $3,000 | $3,000 | $6,000 |
| Cheap model ($1/$5 per MTok) | $300 | $300 | $600 |
Against the frontier model, 476x. Against the cheap model of the current generation, 47x. The "100x" in the title is the average between the two extremes, and the math that matters is yours, not mine. The detail almost nobody notices in the table: half the savings come from output being zero. In a classification, the JSON coming back is short, but an output token costs five times an input token on most price sheets. That little slice disappears entirely.
Choice, Score, and Noul: the three primitives with running Python code
Jev's entire API fits into three question types. The docs call them primitives, and the analogy is a good one: they're small blocks you compose in code, not prompts you try to tune.
| Type | Answers | Returns |
|---|---|---|
Choice |
Which of these options? | choice, probabilities, confidence |
Score |
What level on this scale? | score, legend, probabilities, confidence |
Noul |
Is this true? | noul (0 to 1) |
Installation and first call
pip install typesafe-sdk
export TYPESAFE_API_KEY="sk-..."
The client reads the key from the environment and points at jev-latest by default. This code runs as is:
from typesafe_sdk import Choice, Noul, Score, TypeSafeClient
client = TypeSafeClient()
ticket = (
"Faz três dias que tento ligar minha conta do Stripe e a integração "
"quebra em toda tentativa. Estou perdendo venda. Preciso de ajuda urgente."
)
resposta = client.system_one(
state=ticket,
questions={
"time": Choice(
instructions="Qual time deve atender este ticket?",
criteria={
"cobranca": "Problemas de pagamento, fatura ou assinatura",
"tecnico": "Bugs, erros ou falha de integração",
"vendas": "Dúvidas de preço ou de conta comercial",
},
),
"irritacao": Score(
instructions="O quanto o cliente parece irritado",
criteria=[
"Calmo, só relatando o fato",
"Irritado, mas civilizado",
"Muito bravo, linguagem forte",
],
),
"e_urgente": Noul(
instructions="A mensagem expressa urgência ou sensibilidade a tempo",
),
},
)
print(resposta.answers["time"].choice) # "tecnico"
print(resposta.answers["time"].confidence) # 0.93
print(resposta.answers["time"].probabilities) # {"cobranca": 0.04, "tecnico": 0.93, ...}
print(resposta.answers["irritacao"].score) # 1.7
print(resposta.answers["e_urgente"].noul) # 0.96
print(resposta.usage.input_tokens, resposta.usage.output_tokens) # 118 0
Look at the last print. output_tokens is zero. That's not rounding: there's no output to bill.
What comes back, field by field
choice is the winning option, already a string from your set. probabilities is the full distribution, which lets you look at the runner-up before deciding. confidence summarizes how concentrated that distribution is, and the formula is out in the open in the docs:
def confidence(probabilidades: list[float]) -> float:
n = len(probabilidades)
pico = max(probabilidades)
return max(0.0, min(1.0, (n * pico - 1) / (n - 1)))
print(confidence([0.93, 0.04, 0.03])) # 0.895
print(confidence([1/3, 1/3, 1/3])) # 0.0
A uniform distribution gives a confidence of 0. A distribution pinned to one option gives 1. It's a measure of "how decided," not "how correct" — and that's what gives you a second axis to route on.
Two naming gotchas. Noul returns a probability from 0 to 1, not a boolean, and it's the only primitive without a confidence field (0.5 means "I don't know," not "somewhere in the middle"). And Score returns a float that can land between two levels of your scale: 1.7 on a three-position scale means "between irritated and very angry, leaning toward the latter."
Routing in code
The point of Jev isn't the answer. It's that the decision comes back in a format an if understands:
time = resposta.answers["time"]
irritacao = resposta.answers["irritacao"]
urgente = resposta.answers["e_urgente"]
if time.confidence < 0.70:
fila_humana(ticket_id, motivo="classificação incerta")
elif time.choice == "tecnico":
abrir_incidente(ticket_id, prioridade="alta" if urgente.noul > 0.8 else "normal")
elif time.choice == "cobranca":
rotear_cobranca(ticket_id)
if irritacao.score > 1.5:
marcar_para_resposta_prioritaria(ticket_id)
The thresholds are yours, they live versioned in the repo, and they change with a commit. When the business priority shifts, you change a number — you don't rewrite a prompt and pray nothing else regresses along with it.
Do not just follow the news — master it. Hands-on AI Engineering, live, every week, in the largest community in Brazil.
Join the ClãSpeculative fan-out: why the tenth question is almost free
This is the pattern that changes how you design the system, and it's where Jev's architecture pays for itself.
The model ingests the state once and evaluates every question in parallel against it. Adding questions barely moves the response time and costs only the tokens of the new questions, which are cheap. The state — which is usually 90% of the request — you pay for only once.
The practical consequence is what the docs call speculative fan-out: send every question your flow might need, including the ones that only make sense on one branch, and let the code throw away what doesn't matter.
questions = {
"categoria": Choice(
instructions="Qual a natureza deste ticket?",
criteria={
"bug": "Algo quebrado, erro ou falha de integração",
"cobranca": "Cobrança, fatura, reembolso, assinatura",
"feature": "Pedido de funcionalidade nova",
"conta": "Login, permissão, perfil, segurança",
},
),
# especulativas: só importam se categoria == "bug"
"severidade": Score(
instructions="Qual a severidade do problema relatado",
criteria=["Cosmético", "Degradado, com contorno", "Bloqueante, sem contorno"],
),
"tem_passos": Noul(instructions="O usuário descreve passos para reproduzir o problema"),
# especulativa: só importa se categoria == "cobranca"
"pede_reembolso": Noul(instructions="O usuário pede reembolso ou estorno explicitamente"),
}
In an LLM, that list would cost four calls or one mega-prompt that degrades with size. In Jev, it goes out in one call and the whole flow resolves in a single round trip.
And there's a third-party number for this. The parallel questions cookbook runs 13 questions (8 Noul, 2 Choice, 3 Score) over the GDPR Wikipedia article — about 54 thousand characters — comparing one call with 13 questions against 13 single-question calls: 12.2x cheaper and 10x faster, with identical answers. The standard deviation across repetitions was exactly 0.0 on most questions, in both formats. Batching doesn't add noise because each question is evaluated in isolation: one is not context for another.
That also kills a habit. In an LLM, stacking questions in the same prompt pollutes the window and drags down the quality of all of them. Not here. The questions don't see each other.
Jev's 5 traps: literalness, arithmetic, dates, context rot, and zero generation
TypeSafe publishes a page dedicated entirely to known defects, the jev-1.13 jaggedness, revised on September 17, 2026. It's the most honest document on the site and the one you should read before putting this in production.
1. Extreme literalness. "jev-1.13 answers the question you wrote, not the one you meant." Scope, negation, and implicit conditions are read to the letter. The docs' rule of thumb is excellent: when you look at a wrong answer and catch yourself explaining what you really meant, that explanation is the half that was missing from the instruction. Write the exact condition, and put the edge cases in the criteria.
2. Zero arithmetic. It's not a calculator and it doesn't count. Characters in a word, occurrences in a text, items in a list: the error grows with size. The way out is to iterate in code and ask one question per item:
itens = ["typesafe", "maçã", "california", "banana", "laranja"]
LIMIAR = 0.5
resultado = client.system_one(
{"itens": itens},
{
f"item_{i}": Noul(instructions=f"`itens[{i}]` é o nome de uma fruta?")
for i in range(len(itens))
},
)
total = sum(resultado.nouls[f"item_{i}"].noul > LIMIAR for i in range(len(itens)))
It's fan-out again: five questions, one call, and the sum happens in Python, where sums are reliable.
3. Dates are text to it. Jev reads a date as a string, not as an ordered quantity. Asking which comes first, how many days apart they are, or whether one falls inside a window is asking for errors — and it gets worse with mixed formats and relative references. The pattern is to split the work: extraction is judgment (month is a closed set of 12 options, day a set of 31, so it becomes a Choice with an explicit "not provided" option), arithmetic is code.
4. Context rot. Accuracy drops as the state grows with content irrelevant to the decision. Stray detail becomes a distractor. Filter first, send only the fields the question needs. The hard limit exists too: 64k tokens per request in total, and 32k for the state plus the largest question.
5. It generates nothing. Not even a little. It sounds obvious after this whole post, but it's the trap that catches the most people: you can force generation by chaining Choice, and "this won't work well and will be very slow". If you need text, you need another model. Jev decides; something else does the writing.
A bonus that saves you an afternoon: don't expect structural invariants. The same question as a Noul and as a yes/no Choice comes back with numbers that don't line up — the docs show a noul of 0.22 against probabilities["yes"] of 0.01 on the same ticket. And P(x) plus P(não x) doesn't sum to 1. Don't carry a threshold calibrated on one primitive over to another.
Quick FAQ
Can I swap my chatbot for Jev? No. It doesn't write any response at all. Its job is to decide what your code does next: which route, which severity, which passage is worth sending to the model that writes. The thing talking to the user is still an LLM.
Does it work well in Portuguese?
The docs are direct: English is the primary training language and where accuracy is best today. Other languages are supported, but not equally well. Test on your own content and use confidence as a safety net: below your threshold, send it to human review instead of acting.
Can I send images, audio, or PDFs? No. Input is text: a string, a JSON object, or an array of texts. Anything else has to become text or a structured field first. Current rate limit: 250 thousand tokens per second and 1,200 requests per minute, and TypeSafe warns that these numbers are moving.
Is there an open source alternative? There is: Laya, Apache 2.0, ~0.4B parameters and open weights on Hugging Face, plays the same game of typed questions with calibrated probabilities — the honest comparison between the two is in Laya vs Jev.
What to do with this
The Jev model isn't "a faster LLM." It's the admission that we'd been using the wrong tool for half the job: asking for deliberate reasoning and prose from a system that only needed to answer "which of these three" in 80 milliseconds. Three years of structured output engineering were, in large part, an architecture workaround.
What this changes in your system is a boundary question: where the decision ends and the writing begins. Every time your pipeline calls a frontier model just to get back one string from a set of five, you're paying for generation to generate nothing.
Exactly where that swap is worth making is a whole post of its own, and it exists: Replacing an LLM with Jev: 10 places where the migration pays for itself in a week.
{AI Engineer} — apaixonado por Laravel, arquitetura de software e construir produtos com impacto. Compartilho aqui tutoriais, descobertas e reflexões sobre o dia a dia de engenharia.
There is no shortage of content. What is missing is someone to untangle it: what matters now is how to implement it the right way. In the Clã you get that live, every week, with people who have already filtered out the noise.
Join the Clã