Laya vs Jev: the US$40 million decision AI against the free one that runs on your laptop
Six days old, US$40 million in seed funding, and Jev already has an Apache-2.0 competitor that runs offline and is 7x faster at P50.
The question is not which one is better.
It is why nobody had noticed that this job never needed an LLM.
This post compares Laya vs Jev using the numbers each side published, and records the prior-art fight that broke out on Hacker News two days after launch. Separating what you can check from what is still an accusation.
Update, 2026-10-06: the category is no longer a two-way fight. Cloudflare released Clef and Clef-flash with open weights on October 1, and OpenAI announced a Decisions API on top of GPT-6 Luna at DevDay. Both have their own section below, and the scoreboard has been revised.
TL;DR
- What it is: a head-to-head comparison between Jev (TypeSafe AI, closed API, US$40M seed) and Laya, the first open-source alternative to Jev with published weights (Convai Innovations, Apache 2.0, 421M parameters, runs locally).
- Stack/Models: Jev via REST API; Laya via
transformers(Python) or ONNX Runtime (@receptron/laya, Node/TypeScript). - Cost/Access: Jev is in early access by waitlist, US$0.042 per million input tokens, output free. Laya has open weights, zero inference cost and the GPU cost is yours.
- New in October: Cloudflare's Clef (27B) and Clef-flash (9B), Apache 2.0, API compatible with Jev's, hosted on Workers AI. And OpenAI's Decisions API on GPT-6 Luna, in limited preview with no public documentation.
- Useful links: Jev's official announcement | Laya's model card on Hugging Face | Clef announcement
- Status: living post. The prior-art dispute is still open; the update log is at the end.
The minimum context: what a System One model is
A System One model does not write text. It takes program state and returns a typed decision with a probability: which queue, which label, escalate or not. No token by token, no parsing crooked JSON, no retry because the model decided to philosophize.
If you have not seen the idea yet, the fundamentals are broken down in Jev: the AI model that writes nothing (in Portuguese). Here the subject is different: the comparison and the fight.
And it is a fight that matters to people building products, not to people collecting news. Choosing between a paid API and open weights for a decision that runs millions of times a day is an architecture decision, and a badly made architecture decision becomes technical debt in six months.
You have seen this choice made by tool preference instead of by numbers. That is exactly what Clã Beer and Code, our Portuguese-speaking community, exists to take out of a team: the bar there is what you build and can defend with measurement, and there are people deciding this in Python, Go, PHP and even embedded. With AI in the middle, the stack stopped being the argument. It is paid, it is a subscription, and anyone who only wants to collect launches will be frustrated.
The Hacker News thread: what Laya's author published in March 2025
On September 15, 2026 TypeSafe AI came out of stealth with Jev and US$40 million in seed funding led by DCVC. On the 17th, Nandakishor M, founder of Convai Innovations, opened a thread on Hacker News saying he had published the same idea in March 2025, with paper, weights and dataset open, and that it was now being announced as a new discovery.
What you can check yourself:
- VERIFIED. The paper exists and is dated. arXiv:2503.23303, SalesRLAgent: A Reinforcement Learning Approach for Real-Time Sales Conversion Prediction and Optimization, Nandakishor M, submitted March 30, 2025. The technique is PPO over sequence embeddings returning a turn-by-turn conversion trajectory, from 0.0 to 1.0, instead of generating text.
- VERIFIED. There is a second paper, arXiv:2510.01237, Confidence-Aware Routing for Large Language Model Reliability Enhancement, dated September 23, 2025, about confidence-based routing before generation.
- DISPUTED. That this work is the same architecture as Jev.
The counterpoint showed up in the HN discussion itself. In another thread, a commenter summarized the accusation and got a direct reply: "the paper does not describe a model architecture, it describes a system built on embeddings, rag, and orchestrators, they don't seem very similar to me."
And there is a detail Nandakishor himself admits: his model was trained for one task (conversion prediction in sales conversations), while Jev is credited with doing this zero-shot for any set of choices. Different scope.
There is also the acronym coincidence. TypeSafe calls Jev's training method RLCD, Reinforcement Learning for Calibrated Decisions. Laya's model card also describes its training as RLCD, and the project site says the September 2025 paper formalized that framework. I checked: the September paper does not use the term RLCD anywhere. It is recorded both ways, because it is exactly the kind of detail that will decide the argument.
As of this writing I have not found a public response from TypeSafe to the accusation. Nobody has shown code reuse. What is on the table is uncredited prior art, not copying.
Inside Jev: RLCD, 70 to 500 ms, US$0.042/MTok and up to 255 options
Everything below is VERIFIED in the sense of "published by TypeSafe". Not all of it is measured by a third party, and I flag where it is not.
| Item | What TypeSafe publishes |
|---|---|
| Training | RLCD, optimizing calibrated probabilities instead of human preference |
| Latency | 70 ms to 500 ms per call |
| Claimed gain | 40x to 200x faster than frontier LLMs (cited baseline: 3 to 329 seconds) |
| Price | US$0.042 per million input tokens, output "too cheap to meter" |
| Cardinality | up to 255 native options; beyond that, two-stage scoring |
| Hallucination | 0% by schema guarantee |
| Access | early access by waitlist, closed API |
Two honest caveats about this table.
First: "0% hallucination" is a type guarantee, not an accuracy measurement. The model cannot return something outside the schema. That does not mean it picks the right option. Those are different things, and it is worth reading with that filter.
Second: TypeSafe did not publish a public benchmark table in the announcement. The Jev accuracy numbers circulating today came, to a large extent, from measurements made by its competitor. Which is the next block.
Inside Laya: 421M on ModernBERT-large, 32.8 ms at P50 and 3x better ECE
Laya is the first open-source System One model with weights, dataset and paper published. Underneath, a bidirectional encoder with a decision head. No autoregressive generation, a single forward pass.
- Backbone: ModernBERT-large, 395M parameters, full fine-tune.
- Decision head: 2 transformer layers, one scorer per option marker and an
act/escalatehead. - Total: 421M in the English checkpoint, 512 tokens of context, a 192-token budget for options.
- Multilingual variant: mmBERT-base, 322M, 1024 tokens of context, 100+ languages.
- License: Apache 2.0, three checkpoints, the English one with about 808 MB of weights.
The latency numbers, measured on a Tesla T4 (a 2018 GPU, not an H100):
| Scenario | Laya |
|---|---|
| 1 question | 32.8 ms to 39.5 ms |
| 5 questions | 40.1 ms to 84.5 ms |
| 10 questions | 72.3 ms to 158.6 ms |
| Batch throughput | 103 to 332 questions per second |
In batch, marginal cost drops to about 7.2 ms per question. It is a different order of magnitude of problem.
Running it is one line of Python:
from laya import Router
router = Router(preload=True)
res = router.predict(ticket, questions)
print(res["answers"]["queue"]["choice"])
And there is a Node/TypeScript wrapper on top of ONNX Runtime, with no PyTorch and no Python at runtime, published under MIT (@receptron/laya):
import { Laya } from "@receptron/laya";
const laya = await Laya.load();
const result = await laya.systemOne(
{ subject: "Refund not received", body: "..." },
{
department: {
type: "choice",
instructions: "Which team should handle this?",
criteria: { billing: "refunds", support: "bugs", sales: "purchases" },
},
},
);
result.answers.department.choice; // "billing"
If your backend is PHP or Go, the path is the usual one: run Laya as an internal HTTP service and call it from inside the VPC. The expensive part is not the language, it is the GPU.
Laya vs Jev: the table of published numbers
This is where the Laya vs TypeSafe duel gets uncomfortable, because Jev was measured by Convai, the author of Laya. There is, so far, no neutral measurement of the two. Read with that in mind.
| Metric | Laya | Jev | Source of the number |
|---|---|---|---|
| typed-decisions (2,000 decisions) | 0.766 | 0.727 | Convai |
| AG News (4 labels) | 0.950 | 0.910 | Convai |
| DAIR Emotion (6 labels) | 0.595 | 0.480 | Convai |
| Banking77 (77 labels) | 0.425 | 0.870 | Convai |
| ECE (calibration, lower is better) | 0.081 | 0.246 | Convai |
| P50 latency | 32.8 ms | 236 to 276 ms | Convai |
| Usable languages | 45 of 51 | no published benchmark | Convai |
| Price per million tokens | 0 (self-hosted) | US$0.042 | each side |
ECE is the number I find most interesting and the one fewest people will look at. Expected Calibration Error measures whether the probability the model returns matches the actual frequency of being right. If it says 0.9 and is right 90% of the time, ECE is low. 0.081 against 0.246 is three times better, and that is what decides whether you can write if confidence > 0.85: auto_resolve() without it becoming an incident.
On the 7.8x latency figure, a correction nobody makes: 32.8 ms is local inference on a T4. Jev's 236 to 276 ms are API calls, with network, queueing and TLS in the middle. Comparing the two is comparing different things. The gap is still huge, but part of it is geography, not architecture.
Where Laya loses: more than 20 options, fine-tuning and per-domain calibration
This section exists because the model card itself is honest, and most posts about Laya ignored it.
High cardinality destroys the model. On Banking77, with 77 labels, Laya scores 0.425 against Jev's 0.870. Mechanical reason: the 192-token budget for options becomes 3 to 4 tokens per label. No room for a description, no room for criteria, just the squeezed name. In practice, above about 20 options you are in territory where Jev wins cleanly.
Zero-shot it does not work. The base checkpoint scores 0.362 on typed-decisions. Guessing the most common class scores 0.461. In other words: worse than the dumb heuristic. The model card does not hide it: "Laya is a fast base to specialise, not a zero-shot decision engine."
Calibration is not free. That 0.081 ECE is after fitting temperature per domain. The raw checkpoint comes out at 0.213 (0.285 for multilingual). If you do not refit temperature for your kind of question, the probability it returns is no good as an automation trigger.
Ordinal scales are the weak spot. SST-5, which is five-level sentiment, gives 0.372. Ordering intensity is not its strength.
Translated into a project decision: Laya is a fast base to specialize. It costs fine-tuning, a labeled dataset from your domain and a calibration step. Jev costs a credit card. That is the real trade, and it does not show up in any headline.
Do not just follow the news — master it. Hands-on AI Engineering, live, every week, in the largest community in Brazil.
Join the ClãClef joins the fight: Cloudflare, 27B and 9B, Apache 2.0
On October 1, 2026 Cloudflare released Clef and Clef-flash, the first models trained by the Workers AI team. They are decision models in the same mold as Jev, with one difference that changes the board: open weights and an infrastructure company behind them.
| Item | Clef | Clef-flash |
|---|---|---|
| Parameters | 27B | 9B |
| Backbone | Qwen3.8-27B | Qwen3.5-9B |
| License | Apache 2.0 | Apache 2.0 |
| Context | 64k tokens | 64k tokens |
| Input | text, JSON and images | text, JSON and images |
| Question types | choice, score, noul |
choice, score, noul |
| Median latency (Cloudflare) | 209.3 ms | 38.8 ms |
| Where it runs | Workers AI or your own weights | Workers AI or your own weights |
Cloudflare describes the API as "fully Jev-API compatible". The format is the one you have already seen: state plus a map of typed questions, and the response carries a probability for every option of every question. Up to 64 questions per request, according to the model documentation.
const response = await env.AI.run("@cf/cloudflare/clef-flash", {
state: "Checkout has been failing for every customer for the last hour.",
questions: {
urgent: { type: "noul", instructions: "Is this support request urgent?" },
team: {
type: "choice",
instructions: "Which team should handle this request?",
criteria: {
billing: "Payments, invoices, and refunds",
technical: "Outages, errors, and configuration",
sales: "Plans and upgrades",
},
},
},
});
Now the numbers. Cloudflare published ten decision benchmarks and says a Clef model comes first in seven of them. A sample:
| Benchmark | Clef | Clef-flash | Jev | Laya |
|---|---|---|---|---|
| BFCL (case exact) | 98.47 | 98.76 | 95.75 | 38.13 |
| API-Bank (accuracy) | 91.93 | 93.11 | 88.19 | 11.41 |
| BANKING77 (macro-F1) | 94.20 | 90.93 | 79.74 | 14.29 |
| CLINC150+OOS (macro-F1) | 97.43 | 66.77 | 89.27 | 3.19 |
| When2Call (accuracy) | 72.37 | 65.58 | 80.97 | 11.94 |
| BRIGHT (nDCG@10) | 45.91 | 39.26 | 47.52 | 19.90 |
Three readings the "Clef beats Jev" headline hides.
Jev still wins where it matters for agents. When2Call measures knowing when not to call a tool. Jev scores 80.97 against Clef's 72.37. Same on BRIGHT. And on TypeSafe's own workflow evals, the score is tight: Clef ahead on invoices and security incidents, Jev ahead on agent trace observability.
Laya collapses in this table, and that confirms the model card. 14.29 on BANKING77, 3.19 on CLINC150. Cloudflare does not say which checkpoint it used or whether it fine-tuned. The numbers are consistent with the zero-shot use Laya itself advises against, but that is my reading, not Cloudflare's statement. Either way, it is the first measurement of Laya by someone other than Convai, and it is not pretty.
It is the third different latency for the same Jev. TypeSafe publishes 70 to 500 ms. Convai measured 236 to 276 ms. Cloudflare measured a 524.1 ms median. Laya, which Convai measured at 32.8 ms, appears with a 5.8 ms median and a 222.5 ms p95 in Cloudflare's table. Hardware, network and batching change everything. Do not compare numbers across tables.
And two things Cloudflare did not publish. No ECE, which is the metric that decides whether the probability works as a trigger. And no neutral measurement: once again, whoever measured is whoever is selling.
On running it locally: Clef is not a T4 model. Back-of-the-envelope, 9B parameters at half precision already exceed a T4's 16 GB, and 27B calls for datacenter GPUs. The weights are open, but the hardware bar is a different category from Laya's 808 MB. Treat that as my estimate until someone publishes official requirements.
OpenAI Decisions API: GPT-6 Luna decides too
The news is true, with one adjustment: it is not a new model. At DevDay on September 29, 2026 OpenAI announced the Decisions API, an endpoint that constrains GPT-6 Luna, the cheapest model in the GPT-6 family, to questions with a finite set of answers. In OpenAI's own description, quoted by coverage of the announcement:
"Decisions API enables real-time decision-making by focusing Luna's intelligence on a specific set of user-defined questions with finite pre-defined answers."
What is known: it accepts images, it promises decisions in "less than a few hundreds of milliseconds end to end", it is in limited preview for selected customers and broad release is promised for "the coming days".
What is not known, as of today: price, request and response schema, option limit, measured latency and, above all, whether the API returns probabilities and whether they are calibrated. Regular Luna costs US$0.10 input and US$0.50 output per million tokens; whether the Decisions API follows that table, nobody has published.
That last point is what separates things. Jev, Laya and Clef are models trained to return a decision with a probability. The Decisions API, as described, is a frontier LLM with its answer space tied down. It may work very well. But without documented calibrated probability, it cannot go in the same table. I could not open OpenAI's official recap directly; the excerpts above come from coverage quoting the text.
It goes in here as context, not as a measured competitor. When documentation ships, it becomes a table row.
Paid API vs self-hosted: the real math by call volume
Let's do the math with an explicit premise, which is the only honest way to do it.
Premise: a typical decision with program state plus options uses about 600 input tokens. At US$0.042 per million, each decision costs US$0.0000252. That is, US$25.20 per million decisions.
| Monthly volume | Jev cost (estimated) | What Laya requires |
|---|---|---|
| 1 million | ~US$25 | one GPU idle 99% of the time |
| 10 million | ~US$252 | still GPU to spare |
| 50 million | ~US$1,260 | 1 T4 at ~20% usage |
| 267 million | ~US$6,700 | 1 T4 saturated at 103 req/s |
| 860 million | ~US$21,600 | 1 T4 saturated at 332 req/s |
The last two numbers come from the published throughput: 103 to 332 questions per second in batch on a T4, extrapolated to 30 days of saturation.
The reading is simple and probably goes against what you expected.
Below about 10 million decisions a month, Jev is cheaper. US$252 does not even pay for someone to configure, monitor and be on call for a GPU. Paying for the API here is the correct engineering decision.
Above a few hundred million, the math flips with room to spare. A T4 costs a fraction of US$6,700 a month on any cloud, and the same GPU absorbs the whole volume. In that regime, self-hosted is not ideology, it is margin.
And hosted Clef? Clef-flash on Workers AI costs US$0.09 per million input tokens. On the same 600-token premise, that is about US$54 per million decisions, a little more than double Jev. You pay the difference for having no waitlist, for having the weights if you want to leave, and for a stated policy of not reading, storing or training on your requests. I did not find the price of the 27B Clef.
And there is the axis no price resolves: Laya runs offline. If your data cannot leave the VPC because of a contract, privacy law or the justified paranoia of an enterprise customer, the cost comparison ends before it starts. One side is eligible, the other is not. Clef is also eligible on that criterion, as long as you have the GPU for it.
Verified vs disputed: the honest scoreboard
| Claim | Status |
|---|---|
| Jev launched on 2026-09-15 with US$40M seed (DCVC) | VERIFIED |
| Jev price: US$0.042/MTok input, output free | VERIFIED (published by TypeSafe) |
| Jev supports up to 255 native options | VERIFIED (published by TypeSafe) |
| Laya is Apache 2.0, 421M, ModernBERT-large | VERIFIED (public model card) |
| Laya does 32.8 ms P50 on a T4 | VERIFIED (benchmark published by Convai) |
| Laya loses badly above ~20 options | VERIFIED (the model card itself admits it) |
| Laya is 7.8x faster than Jev | VERIFIED with a caveat: competitor's measurement, local vs network |
| Clef and Clef-flash are Apache 2.0, 27B and 9B, Qwen backbone | VERIFIED (published by Cloudflare) |
| A Clef model leads 7 of 10 decision benchmarks | VERIFIED with a caveat: Cloudflare's own measurement |
| Clef-flash is 13x faster than Jev at the median | VERIFIED with a caveat: Cloudflare's measurement, third different latency for Jev |
| Clef returns calibrated probabilities | NO PUBLISHED NUMBER: no ECE in the announcement |
| OpenAI announced a Decisions API on GPT-6 Luna | VERIFIED (DevDay, 2026-09-29, limited preview) |
| The Decisions API returns calibrated probability | NOT PUBLISHED |
| The March 2025 paper exists and is dated | VERIFIED |
| The March 2025 paper describes Jev's architecture | DISPUTED |
| TypeSafe used uncredited work | DISPUTED |
| Code was copied | NO EVIDENCE PRESENTED by either side |
This post does not arbitrate the last three rows. It cannot, with what is public today.
Quick FAQ
Can I swap Jev for Laya without rewriting code?
The call shape is compatible on purpose: you pass state plus a map of typed questions and get answers[key].choice with a probability. What is not portable is quality: without fine-tuning on your domain, raw Laya will make your metrics worse.
Do I need a GPU to run Laya? The published numbers are all on GPU (Tesla T4). There is no CPU benchmark in the model card. With 421M parameters and a single pass, CPU is plausible for low volume, but treat that as a hypothesis to test, not a fact.
Does Laya replace my LLM? No. Neither of them writes text. They replace the LLM call you use today only to decide something: route a ticket, classify intent, pick a tool, decide whether to escalate to a human. Generation stays with the LLM.
Clef or Laya? Different animals. Laya has 421M parameters, runs on an old GPU and requires fine-tuning. Clef has 9B or 27B, handles many options and images with no additional training according to Cloudflare's numbers, and needs serious hardware if you want to run it at home. If the question is "which one do I test tomorrow without a dataset", it is hosted Clef-flash.
Which one do I use today, October 6, 2026? Few options, very high volume, data that cannot leave the house and budget for fine-tuning: Laya. Many options, image input or a team without a dataset: Clef-flash on Workers AI, which has no waitlist. Deciding when to call a tool inside an agent: Jev still measures better, if you have access. OpenAI's Decisions API: wait for the documentation. None of these answers is permanent, and the September 21 answer already changed in two weeks.
What stands
Six days separated a US$40 million model from an 808 MB competitor that runs on a 2018 GPU. This is not about Jev being bad. It is about the whole category having been expensive for a reason that no longer holds: we were using a generative model for a task that was never generative.
Two weeks later, Cloudflare and OpenAI came in. When an infrastructure company opens the weights and the biggest lab in the world creates an endpoint for it, it has stopped being a startup bet. It has become a layer.
The prior-art dispute probably will not be resolved, and an official open-source Jev is on nobody's roadmap. What it exposes is more useful: calibrated typed decisions were not an open problem waiting for a lab with a nine-figure seed. It was a bidirectional encoder with a head on top, and someone had written that in a sales paper in March 2025 with nobody looking.
If you want the fundamentals before the comparison, start with Jev: the AI model that writes nothing. If you want the bigger pattern, open weights catching up with a closed API in weeks, the Qwen 3.8 27B case is the same film with a different cast (both in Portuguese).
Update log
This is a living post and will be updated as the dispute evolves.
- 2026-10-06: added Cloudflare's Clef and Clef-flash (released October 1, Apache 2.0) and OpenAI's Decisions API on GPT-6 Luna (announced September 29, limited preview, no documentation). Scoreboard and FAQ revised. First third-party measurement of Laya, made by Cloudflare. Still no neutral measurement and no public response from TypeSafe to the prior-art dispute.
- 2026-09-21: published. Jev in early access since September 15; Laya published on September 19; prior-art thread opened on HN on September 17, with follow-up discussion the same week. No public response from TypeSafe so far. No neutral measurement of the two models.
If you run an independent benchmark of the two, send it over. It goes in here with credit.
{AI Engineer} — apaixonado por Laravel, arquitetura de software e construir produtos com impacto. Compartilho aqui tutoriais, descobertas e reflexões sobre o dia a dia de engenharia.
There is no shortage of content. What is missing is someone to untangle it: what matters now is how to implement it the right way. In the Clã you get that live, every week, with people who have already filtered out the noise.
Join the Clã