~/beer-and-code
▪ next event Workshop: Jev na Prática para Devs · 14 Oct · 19h — duração de 2h a 3h, ao vivo via Google Meet save your seat ›
~ / news / laya-vs-jev-open-source-decision-ai $
News

Laya vs Jev: The $40 Million Decision AI vs. the Free One That Runs on Your Laptop

LS Lucas Souza · · 14 min read
Laya vs Jev: The $40 Million Decision AI vs. the Free One That Runs on Your Laptop

Six days old, $40 million in seed funding, and Jev already has an Apache-2.0 competitor that runs offline and is 7x faster at P50.

The question isn't which one is better.

It's why nobody had noticed this job didn't need an LLM.

This post compares Laya vs Jev using the numbers each side published, and documents the prior-art fight that blew up on Hacker News two days after the launch. Separating what you can check from what is still an accusation.

TL;DR

  • What it is: a head-to-head comparison between Jev (TypeSafe AI, closed API, $40M seed) and Laya, the first open source alternative to Jev with published weights (Convai Innovations, Apache 2.0, 421M parameters, runs locally).
  • Stack/Models: Jev via REST API; Laya via transformers (Python) or ONNX Runtime (@receptron/laya, Node/TypeScript).
  • Cost/Access: Jev is in early access behind a waitlist, $0.042 per million input tokens, output free. Laya has open weights, zero inference cost, and the GPU bill is yours.
  • Useful links: official Jev announcement | Laya model card on Hugging Face
  • Status: living post. The prior-art dispute is still open; the update log is at the bottom.

The minimum context: what a System One model is

A System One model doesn't write text. It takes program state and returns a typed decision with a probability: which queue, which label, escalate or don't. No token-by-token output, no mangled JSON to parse, no retry because the model decided to philosophize.

If you haven't seen the idea yet, the fundamentals are broken down in Jev: the AI model that doesn't write anything. This post is about something else: the comparison and the fight.

And it's a fight that matters to people who build product, not to people who collect news. Choosing between a paid API and open weights for a decision that runs millions of times a day is an architecture decision, and a badly made architecture decision turns into tech debt in six months.

You've seen that choice made on tool preference instead of numbers. That's exactly what the Clã Beer and Code exists to get out of your team: in there, the yardstick is what you build and can defend with measurements, and there are people making this call in Python, Go, PHP, and even embedded — with AI in the mix, the stack stopped being the argument. It's paid, it's a subscription, and anyone who just wants to collect launches is going to be frustrated.

The Hacker News thread: what Laya's author published in March 2025

On September 15, 2026, TypeSafe AI came out of stealth with Jev and a $40 million seed round led by DCVC. On the 17th, Nandakishor M, founder of Convai Innovations, opened a thread on Hacker News saying he had published the same idea in March 2025, with an open paper, weights, and dataset, and that it was now being announced as a new discovery.

What you can check yourself:

  • VERIFIED. The paper exists and it's dated. arXiv:2503.23303, SalesRLAgent: A Reinforcement Learning Approach for Real-Time Sales Conversion Prediction and Optimization, Nandakishor M, submitted March 30, 2025. The technique is PPO over sequence embeddings returning a turn-by-turn conversion trajectory, from 0.0 to 1.0, instead of generating text.
  • VERIFIED. There's a second paper, arXiv:2510.01237, Confidence-Aware Routing for Large Language Model Reliability Enhancement, from September 23, 2025, on confidence-based routing before generation.
  • DISPUTED. That this work is the same architecture as Jev.

The pushback showed up in the HN thread itself. In another thread, a commenter summarized the accusation and got a blunt reply: "the paper does not describe a model architecture, it describes a system built on embeddings, rag, and orchestrators, they don't seem very similar to me."

And there's a detail Nandakishor himself admits: his model was trained for one task (conversion prediction in sales conversations), while Jev is being credited with doing this zero-shot for any set of choices. Different scope.

Then there's the acronym coincidence. TypeSafe calls Jev's training method RLCD, Reinforcement Learning for Calibrated Decisions. Laya's model card also describes its training as RLCD, and the project site says the September 2025 paper formalized that framework. I went and checked: the September paper doesn't use the term RLCD anywhere. I'm putting both on the record, because this is exactly the kind of detail that will settle the argument.

As of publication, I haven't found a public response from TypeSafe to the accusation. Nobody has shown code reuse. What's on the table is uncredited prior art, not copying.

Inside Jev: RLCD, 70 to 500 ms, $0.042/MTok, and up to 255 options

Everything below is VERIFIED in the sense of "published by TypeSafe." Not all of it is measured by a third party, and I flag where it isn't.

Item What TypeSafe publishes
Training RLCD, optimizing calibrated probabilities instead of human preference
Latency 70 ms to 500 ms per call
Claimed gain 40x to 200x faster than frontier LLMs (cited baseline: 3 to 329 seconds)
Price $0.042 per million input tokens, output "too cheap to meter"
Cardinality up to 255 native options; above that, two-stage scoring
Hallucination 0% by schema guarantee
Access early access behind a waitlist, closed API

Two honest caveats about that table.

First: the "0% hallucination" is a type guarantee, not an accuracy measurement. The model can't return something outside the schema. That doesn't mean it picks the right option. Those are different things, and it's worth reading with that filter on.

Second: TypeSafe didn't publish a public benchmark table in the announcement. The Jev accuracy numbers floating around today came, for the most part, from measurements run by its competitor. Which is the next section.

Inside Laya: 421M on ModernBERT-large, 32.8 ms at P50, and 3x better ECE

Laya is the first open source System One model with published weights, dataset, and paper. Under the hood, it's a bidirectional encoder with a decision head. No autoregressive generation, just a single forward pass.

  • Backbone: ModernBERT-large, 395M parameters, full fine-tune.
  • Decision head: 2 transformer layers, one scorer per option marker, and an act/escalate head.
  • Total: 421M in the English checkpoint, 512 tokens of context, a 192-token budget for the options.
  • Multilingual variant: mmBERT-base, 322M, 1024 tokens of context, 100+ languages.
  • License: Apache 2.0, three checkpoints, the English one at roughly 808 MB of weights.

The latency numbers, measured on a Tesla T4 (a GPU from 2018, not an H100):

Scenario Laya
1 question 32.8 ms to 39.5 ms
5 questions 40.1 ms to 84.5 ms
10 questions 72.3 ms to 158.6 ms
Batch throughput 103 to 332 questions per second

In batch, the marginal cost drops to about 7.2 ms per question. That's a different order of magnitude of problem.

Running it is one line of Python:

from laya import Router

router = Router(preload=True)
res = router.predict(ticket, questions)

print(res["answers"]["queue"]["choice"])

And there's a Node/TypeScript wrapper on top of ONNX Runtime, with no PyTorch and no Python at runtime, published under MIT (@receptron/laya):

import { Laya } from "@receptron/laya";

const laya = await Laya.load();

const result = await laya.systemOne(
  { subject: "Refund not received", body: "..." },
  {
    department: {
      type: "choice",
      instructions: "Which team should handle this?",
      criteria: { billing: "refunds", support: "bugs", sales: "purchases" },
    },
  },
);

result.answers.department.choice; // "billing"

If your backend is PHP or Go, the path is the same as always: stand Laya up as an internal HTTP service and call it from inside the VPC. The expensive part isn't the language, it's the GPU.

Laya vs Jev: the table of published numbers

This is where the Laya vs. TypeSafe matchup gets uncomfortable, because the one who measured Jev was Convai, Laya's author. So far, there is no neutral measurement of the two. Read with that in mind.

Metric Laya Jev Source of the number
typed-decisions (2,000 decisions) 0.766 0.727 Convai
AG News (4 labels) 0.950 0.910 Convai
DAIR Emotion (6 labels) 0.595 0.480 Convai
Banking77 (77 labels) 0.425 0.870 Convai
ECE (calibration, lower is better) 0.081 0.246 Convai
P50 latency 32.8 ms 236 to 276 ms Convai
Usable languages 45 of 51 no published benchmark Convai
Price per million tokens 0 (self-hosted) $0.042 each side

ECE is the number I find most interesting and the one the fewest people will look at. Expected Calibration Error measures whether the probability the model returns matches how often it's actually right. If it says 0.9 and it's right 90% of the time, ECE is low. 0.081 against 0.246 is three times better, and that's what decides whether you can write if confidence > 0.85: auto_resolve() without it turning into an incident.

On the 7.8x latency figure, a correction nobody makes: 32.8 ms is local inference on a T4. Jev's 236 to 276 ms is an API call, with network, queueing, and TLS in the middle. Comparing the two is comparing different things. The gap is still huge, but part of it is geography, not architecture.

▪ Clã Beer and Code

Do not just follow the news — master it. Hands-on AI Engineering, live, every week, in the largest community in Brazil.

Join the Clã

Where Laya gets beat: more than 20 options, fine-tuning, and per-domain calibration

This section exists because the model card itself is honest, and most posts about Laya ignored it.

High cardinality wrecks the model. On Banking77, with 77 labels, Laya scores 0.425 against Jev's 0.870. The reason is mechanical: the 192-token budget for options works out to 3 to 4 tokens per label. There's no room for a description, no room for criteria, all that's left is the name squeezed in. In practice, above about 20 options you're in territory where Jev wins clean.

It's useless zero-shot. The base checkpoint scores 0.362 on typed-decisions. Guessing the most common class scores 0.461. In other words: worse than the dumb heuristic. The model card doesn't hide it: "Laya is a fast base to specialise, not a zero-shot decision engine."

Calibration doesn't come for free. That 0.081 ECE is after tuning temperature per domain. The raw checkpoint comes out at 0.213 (0.285 for the multilingual one). If you don't refit temperature for your kind of question, the probability it returns is worthless as an automation trigger.

Ordinal scales are the weak spot. SST-5, which is five-level sentiment, comes in at 0.372. Ranking intensity is not its strong suit.

Translated into a project decision: Laya is a fast base to specialize. It costs you fine-tuning, a labeled dataset from your domain, and a calibration step. Jev costs you a credit card. That's the real trade-off, and it doesn't show up in any headline.

Let's do the math with an explicit assumption, which is the only honest way to do it.

Assumption: a typical decision with program state plus options uses about 600 input tokens. At $0.042 per million, each decision comes out to $0.0000252. That is, $25.20 per million decisions.

Monthly volume Jev cost (estimated) What Laya requires
1 million ~$25 a GPU sitting idle 99% of the time
10 million ~$252 still GPU to spare
50 million ~$1,260 1 T4 at ~20% utilization
267 million ~$6,700 1 T4 saturated at 103 req/s
860 million ~$21,600 1 T4 saturated at 332 req/s

The last two numbers come from the published throughput: 103 to 332 questions per second in batch on a T4, extrapolated to 30 days of saturation.

The takeaway is simple and probably the opposite of what you expected.

Below about 10 million decisions a month, Jev is cheaper. $252 doesn't even cover the cost of someone setting up, monitoring, and being on call for a GPU. Paying for the API here is the correct engineering decision.

Above a few hundred million, the math flips with room to spare. A T4 costs a fraction of $6,700 a month on any cloud, and that same GPU absorbs the entire volume. In that regime, self-hosted isn't ideology, it's margin.

And then there's the axis no price fixes: Laya runs offline. If your data can't leave the VPC because of a contract, LGPD, or an enterprise customer's justified paranoia, the cost comparison is over before it starts. One side is eligible, the other isn't.

Verified vs disputed: the honest scoreboard

Claim Status
Jev launched on September 15, 2026 with a $40M seed (DCVC) VERIFIED
Jev pricing: $0.042/MTok input, output free VERIFIED (published by TypeSafe)
Jev supports up to 255 native options VERIFIED (published by TypeSafe)
Laya is Apache 2.0, 421M, ModernBERT-large VERIFIED (public model card)
Laya hits 32.8 ms P50 on a T4 VERIFIED (benchmark published by Convai)
Laya loses badly above ~20 options VERIFIED (the model card itself admits it)
Laya is 7.8x faster than Jev VERIFIED with a caveat: the competitor's measurement, local vs network
The March 2025 paper exists and is dated VERIFIED
The March 2025 paper describes Jev's architecture DISPUTED
TypeSafe used uncredited work DISPUTED
Code was copied NO EVIDENCE PRESENTED by either side

This post doesn't rule on the last three rows. It can't, with what's public today.

Quick FAQ

Can I swap Jev for Laya without rewriting my code? The call shape is compatible on purpose: you pass state plus a map of typed questions and get back answers[chave].choice with a probability. What isn't portable is the quality: without fine-tuning on your domain, raw Laya will make your metrics worse.

Do I need a GPU to run Laya? The published numbers are all on GPU (Tesla T4). There's no CPU benchmark in the model card. With 421M parameters and a single pass, CPU is plausible for low volume, but treat it as a hypothesis to test, not a fact.

Does Laya replace my LLM? No. Neither of them writes text. They replace the LLM call you use today just to decide something: route a ticket, classify intent, pick a tool, decide whether to escalate to a human. Generation stays with the LLM.

Which one do I use today, September 21, 2026? Few options, high volume, data that can't leave the building: Laya, with a budget for fine-tuning. Lots of options, low volume, small team: Jev. Prototyping: Jev, because it doesn't require a dataset. None of these answers is permanent, and the second half of the table could change in a month.

What's left standing

Six days separated a $40 million model from an 808 MB competitor that runs on a GPU from 2018. This isn't about Jev being bad. It's about the whole category having been expensive for a reason that no longer holds: we were using a generative model for a task that was never generative.

The prior-art dispute probably won't get resolved, and an official open source Jev isn't on anyone's roadmap. What it exposes is more useful: calibrated typed decisions weren't an open problem waiting for a lab with a nine-figure seed. It was a bidirectional encoder with a head on top, and someone had written that down in a sales paper in March 2025 with nobody looking.

If you want the fundamentals before the comparison, start with Jev: the AI model that doesn't write anything. If you want the bigger pattern, open weights catching up to a closed API in weeks, the Qwen 3.8 27B case is the same movie with a different cast.

Update log

This is a living post and will be updated as the dispute evolves.

  • September 21, 2026 — published. Jev in early access since September 15; Laya published on September 19; prior-art thread opened on HN on September 17, with spin-off discussion that same week. No public response from TypeSafe so far. No neutral measurement of the two models.

If you run an independent benchmark of the two, send it over. It goes in here with credit.

Lucas Souza
Written by
Lucas Souza

{AI Engineer} — apaixonado por Laravel, arquitetura de software e construir produtos com impacto. Compartilho aqui tutoriais, descobertas e reflexões sobre o dia a dia de engenharia.

▪ Clã Beer and Code

There is no shortage of content. What is missing is someone to untangle it: what matters now is how to implement it the right way. In the Clã you get that live, every week, with people who have already filtered out the noise.

Join the Clã
Meet the Clã Beer and Code
playing