Jev vs Laya: Hosted Decisions or Your Own GPU
Laya is the most direct answer to “can I do this without an API bill”. It is an open-weights decision model published by ConvAI Innovations under Apache 2.0, with the weights on Hugging Face and a design aimed squarely at running on hardware you already own.
So the comparison is not really about accuracy tables. It is the old question in a new costume: do you want a per-call price and no infrastructure, or no per-call price and infrastructure.
Background
What Laya Is
Per its own site, Laya is a non-autoregressive decision model built by ConvAI Innovations, released under Apache 2.0 with weights on Hugging Face, and intended to be self-hosted rather than consumed as an API. Its multilingual checkpoint is built on mmBERT-base, and the project publishes single-question and batched latency figures measured on a GPU, along with accuracy and calibration numbers against its own test sets.
Those are the project’s own measurements, which is how we treat them: as a reason to evaluate it, not as a result we are repeating. If the decision matters, run both on a few hundred of your own labelled examples. Everyone’s benchmark looks good on the data they chose.
Architecturally the two are close cousins. Both skip text generation, both return a distribution over options you declare, both are built for the same class of problem. The differences that will actually affect your project are operational.
Source for the factual claims above: laya.convaiinnovations.com, checked 2026-09-25. We link it so you can check it rather than take our word for it.
Side by side
What Actually Differs
The rows that change your architecture. Laya’s column reflects its own published description, checked on 2026-09-25.
| Jev (via jev-ai.org) | Laya | |
|---|---|---|
| Weights | Closed, hosted by the vendor | Open, Apache 2.0, on Hugging Face |
| How you run it | One HTTPS call; nothing to operate | Your own hardware, your own serving stack |
| Marginal cost per decision | Input tokens at $0.042/1M; output free | No licence fee; you pay for the GPU whether it is busy or not |
| Data leaves your network | Yes — to us, and on to the model provider | No, if you host it yourself |
| Fine-tuning on your data | Not available | Yes — that is a main reason to pick open weights |
| Scaling to a burst | Provider capacity, subject to rate limits | Your capacity planning problem |
| Who fixes a regression | The vendor, on their schedule; pin a version to avoid surprises | You, on yours — which is both the cost and the benefit |
How to choose
Neither One Wins Every Time
Choose the hosted model when
- You do not have a GPU budget, an inference platform or someone whose job includes keeping one healthy.
- Your volume is bursty. A model you pay for per call costs nothing at 3am; a reserved GPU costs the same at 3am as at noon.
- You want to ship this week. Getting to a working decision endpoint is an API key and twenty lines, not a deployment.
- You want zero-shot quality on a label set that changes often, without owning a retraining loop.
- You would rather spend your evaluation effort on your own data than on serving infrastructure.
Choose Laya when
- Your data cannot leave your network. This is the argument that ends the discussion, and no amount of API convenience answers it.
- You already run GPUs and have spare capacity, which makes the marginal cost of a decision effectively zero.
- Your volume is high and steady — the point where a reserved GPU beats per-call pricing arrives sooner than people expect.
- You want to fine-tune on your own labelled data, which open weights allow and a hosted API does not.
- You need to pin a model that will still behave identically in three years, independent of anyone’s roadmap.
We sell one of these two, so read the right-hand column with that in mind — and then run both on two hundred of your own labelled examples, which settles it better than any page on the internet can.
Questions
Jev vs Laya: Common Questions
Is Laya a drop-in replacement for Jev?
No. They are separate models from separate teams with separate APIs, and the request shapes differ. The concepts port cleanly — declared options, a distribution back — but the integration does not copy across.
Which one is more accurate?
Both projects publish favourable numbers on their own evaluations, which is what every project does. The only comparison worth acting on is one you run on a few hundred of your own labelled examples, and it is cheap to run: an afternoon of work and a few cents of API calls.
Can I use both?
Yes, and it is a reasonable architecture. Run the open model for the bulk of your volume and send the low-confidence tail to a hosted model, or use the hosted one to label a training set for the one you fine-tune.
What does self-hosting really cost?
The GPU, the serving stack, the monitoring, the on-call, and the engineer-days that go into all of it. That can be excellent value at high steady volume and terrible value at low bursty volume. Work out your decisions-per-month before deciding, not after.
More
Other Comparisons
Jev vs SemIf / OpenJev
The self-hosted reimplementation that reads option logits straight from an open model.
Jev vs djev
The multimodal alternative: same idea, but it accepts images and camera frames.
Jev vs LLM structured output
JSON mode guarantees the output parses. It does not tell you how sure the model was.
The Cheapest Way to Settle It
Load a scenario, paste in your own text, and see what the distribution says. Editing costs nothing and needs no account.
