Comparison

Jev vs Zero-Shot Classification with an NLI Model

The Hugging Face zero-shot pipeline is how a lot of teams first solve this: take an NLI model, turn each candidate label into a hypothesis — "this text is about billing" — and score entailment. No labelled data, no API bill, runs on your own machine.

It is a genuinely clever trick and it gets you surprisingly far. The places it stops are specific and worth knowing before you build on it.

Background

How Zero-Shot NLI Classification Works, and Where It Strains

The pipeline runs one forward pass per candidate label, scoring how much the text entails the hypothesis built from that label. With ten labels that is ten forward passes, so latency and compute scale with the size of your label set rather than staying flat.

The scores are per-label entailment probabilities, normalised across labels afterwards. That normalisation is a convenience, not a calibration: the resulting numbers do not reliably mean "70% likely to be this label", and they shift when you add or remove a label from the candidate set. Thresholding them tends to work on the examples you tried it on and not on the ones you did not.

It is also sensitive to phrasing in a way that is hard to debug. The hypothesis template, the exact wording of a label, and its length all move the scores, and there is no place to write down what a label actually means beyond its name. That is the piece a decision model gives you: a description per label, which is where the boundary cases belong.

Side by side

What Actually Differs

Both need no labelled data. After that they diverge quickly.

 Jev (via jev-ai.org)Zero-shot NLI
Labelled data neededNoneNone
Cost per callInput tokens at $0.042/1MYour own compute
Cost of a tenth labelOne more line in the criteria objectA tenth forward pass, every call
Where label meaning livesA description per label, in the requestThe label string and the hypothesis template
CalibrationTrained for; verify on your dataEntailment scores normalised after the fact
Stability when the label set changesOther labels keep their meaningNormalised scores shift across the whole set
Scores and scalesOrdered score questions with a tier distributionNot a natural fit; needs to be faked with labels
Several questions about one textOne call, charged onceA separate run per question

How to choose

Neither One Wins Every Time

Choose a decision model when

  • Your labels need explanations. "Escalate" means something specific in your business and the word alone does not carry it.
  • You have more than a handful of labels, where per-label forward passes stop being cheap.
  • You need an ordered scale, which entailment scoring does not naturally express.
  • You want to threshold on the probability and have that threshold keep its meaning as the label set evolves.
  • You want several judgements about one text in one pass.

Stay with zero-shot NLI when

  • You are exploring. It costs nothing to try and tells you quickly whether the task is separable at all.
  • You have two or three labels whose names are genuinely self-explanatory.
  • Nothing may leave your network and you would rather not host something larger.
  • It is a batch job with no latency budget, where running ten forward passes per item is simply fine.
  • It already works on your data. Measure before replacing anything — "it works" beats "it is more principled".

We sell one of these two, so read the right-hand column with that in mind — and then run both on two hundred of your own labelled examples, which settles it better than any page on the internet can.

Questions

Jev vs Zero-Shot NLI Classification: Common Questions

Can I just add descriptions to my label names?

You can lengthen the hypothesis, and it sometimes helps. It also changes the score distribution in ways that are hard to predict, because the model is judging entailment of a longer sentence rather than reading a definition. A criteria description is read as a definition, which is a different thing.

Is a decision model just a better-packaged zero-shot classifier?

It is the same problem attacked differently. Zero-shot NLI repurposes an entailment model one label at a time; a decision model scores the declared options in one pass and is trained so those numbers are usable as probabilities. The second is the part that is hard to replicate by packaging.

What about embeddings plus cosine similarity?

Good for retrieval, weak for classification with real criteria. Similarity is symmetric and topical, and it cannot express "this belongs in escalate because a legal threat was made". It also cannot do yes/no questions about a text at all.

How do I compare them fairly?

Label two or three hundred examples of your own, run both, and look at accuracy and at what the confidence scores do. If zero-shot wins on your data, use it — the aim is a system that works, not one that uses the newer idea.

The Cheapest Way to Settle It

Load a scenario, paste in your own text, and see what the distribution says. Editing costs nothing and needs no account.