Comparison

Jev vs Laya: Hosted Decisions or Your Own GPU

Laya is the most direct answer to “can I do this without an API bill”. It is an open-weights decision model published by ConvAI Innovations under Apache 2.0, with the weights on Hugging Face and a design aimed squarely at running on hardware you already own.

So the comparison is not really about accuracy tables. It is the old question in a new costume: do you want a per-call price and no infrastructure, or no per-call price and infrastructure.

Background

What Laya Is

Per its own site, Laya is a non-autoregressive decision model built by ConvAI Innovations, released under Apache 2.0 with weights on Hugging Face, and intended to be self-hosted rather than consumed as an API. Its multilingual checkpoint is built on mmBERT-base, and the project publishes single-question and batched latency figures measured on a GPU, along with accuracy and calibration numbers against its own test sets.

Those are the project’s own measurements, which is how we treat them: as a reason to evaluate it, not as a result we are repeating. If the decision matters, run both on a few hundred of your own labelled examples. Everyone’s benchmark looks good on the data they chose.

Architecturally the two are close cousins. Both skip text generation, both return a distribution over options you declare, both are built for the same class of problem. The differences that will actually affect your project are operational.

Source for the factual claims above: laya.convaiinnovations.com, checked 2026-09-25. We link it so you can check it rather than take our word for it.

Side by side

What Actually Differs

The rows that change your architecture. Laya’s column reflects its own published description, checked on 2026-09-25.

 Jev (via jev-ai.org)Laya
WeightsClosed, hosted by the vendorOpen, Apache 2.0, on Hugging Face
How you run itOne HTTPS call; nothing to operateYour own hardware, your own serving stack
Marginal cost per decisionInput tokens at $0.042/1M; output freeNo licence fee; you pay for the GPU whether it is busy or not
Data leaves your networkYes — to us, and on to the model providerNo, if you host it yourself
Fine-tuning on your dataNot availableYes — that is a main reason to pick open weights
Scaling to a burstProvider capacity, subject to rate limitsYour capacity planning problem
Who fixes a regressionThe vendor, on their schedule; pin a version to avoid surprisesYou, on yours — which is both the cost and the benefit

How to choose

Neither One Wins Every Time

Choose the hosted model when

  • You do not have a GPU budget, an inference platform or someone whose job includes keeping one healthy.
  • Your volume is bursty. A model you pay for per call costs nothing at 3am; a reserved GPU costs the same at 3am as at noon.
  • You want to ship this week. Getting to a working decision endpoint is an API key and twenty lines, not a deployment.
  • You want zero-shot quality on a label set that changes often, without owning a retraining loop.
  • You would rather spend your evaluation effort on your own data than on serving infrastructure.

Choose Laya when

  • Your data cannot leave your network. This is the argument that ends the discussion, and no amount of API convenience answers it.
  • You already run GPUs and have spare capacity, which makes the marginal cost of a decision effectively zero.
  • Your volume is high and steady — the point where a reserved GPU beats per-call pricing arrives sooner than people expect.
  • You want to fine-tune on your own labelled data, which open weights allow and a hosted API does not.
  • You need to pin a model that will still behave identically in three years, independent of anyone’s roadmap.

We sell one of these two, so read the right-hand column with that in mind — and then run both on two hundred of your own labelled examples, which settles it better than any page on the internet can.

Questions

Jev vs Laya: Common Questions

Is Laya a drop-in replacement for Jev?

No. They are separate models from separate teams with separate APIs, and the request shapes differ. The concepts port cleanly — declared options, a distribution back — but the integration does not copy across.

Which one is more accurate?

Both projects publish favourable numbers on their own evaluations, which is what every project does. The only comparison worth acting on is one you run on a few hundred of your own labelled examples, and it is cheap to run: an afternoon of work and a few cents of API calls.

Can I use both?

Yes, and it is a reasonable architecture. Run the open model for the bulk of your volume and send the low-confidence tail to a hosted model, or use the hosted one to label a training set for the one you fine-tune.

What does self-hosting really cost?

The GPU, the serving stack, the monitoring, the on-call, and the engineer-days that go into all of it. That can be excellent value at high steady volume and terrible value at low bursty volume. Work out your decisions-per-month before deciding, not after.

The Cheapest Way to Settle It

Load a scenario, paste in your own text, and see what the distribution says. Editing costs nothing and needs no account.