Jev vs Zero-Shot Classification with an NLI Model
The Hugging Face zero-shot pipeline is how a lot of teams first solve this: take an NLI model, turn each candidate label into a hypothesis — "this text is about billing" — and score entailment. No labelled data, no API bill, runs on your own machine.
It is a genuinely clever trick and it gets you surprisingly far. The places it stops are specific and worth knowing before you build on it.
Background
How Zero-Shot NLI Classification Works, and Where It Strains
The pipeline runs one forward pass per candidate label, scoring how much the text entails the hypothesis built from that label. With ten labels that is ten forward passes, so latency and compute scale with the size of your label set rather than staying flat.
The scores are per-label entailment probabilities, normalised across labels afterwards. That normalisation is a convenience, not a calibration: the resulting numbers do not reliably mean "70% likely to be this label", and they shift when you add or remove a label from the candidate set. Thresholding them tends to work on the examples you tried it on and not on the ones you did not.
It is also sensitive to phrasing in a way that is hard to debug. The hypothesis template, the exact wording of a label, and its length all move the scores, and there is no place to write down what a label actually means beyond its name. That is the piece a decision model gives you: a description per label, which is where the boundary cases belong.
Side by side
What Actually Differs
Both need no labelled data. After that they diverge quickly.
| Jev (via jev-ai.org) | Zero-shot NLI | |
|---|---|---|
| Labelled data needed | None | None |
| Cost per call | Input tokens at $0.042/1M | Your own compute |
| Cost of a tenth label | One more line in the criteria object | A tenth forward pass, every call |
| Where label meaning lives | A description per label, in the request | The label string and the hypothesis template |
| Calibration | Trained for; verify on your data | Entailment scores normalised after the fact |
| Stability when the label set changes | Other labels keep their meaning | Normalised scores shift across the whole set |
| Scores and scales | Ordered score questions with a tier distribution | Not a natural fit; needs to be faked with labels |
| Several questions about one text | One call, charged once | A separate run per question |
How to choose
Neither One Wins Every Time
Choose a decision model when
- Your labels need explanations. "Escalate" means something specific in your business and the word alone does not carry it.
- You have more than a handful of labels, where per-label forward passes stop being cheap.
- You need an ordered scale, which entailment scoring does not naturally express.
- You want to threshold on the probability and have that threshold keep its meaning as the label set evolves.
- You want several judgements about one text in one pass.
Stay with zero-shot NLI when
- You are exploring. It costs nothing to try and tells you quickly whether the task is separable at all.
- You have two or three labels whose names are genuinely self-explanatory.
- Nothing may leave your network and you would rather not host something larger.
- It is a batch job with no latency budget, where running ten forward passes per item is simply fine.
- It already works on your data. Measure before replacing anything — "it works" beats "it is more principled".
We sell one of these two, so read the right-hand column with that in mind — and then run both on two hundred of your own labelled examples, which settles it better than any page on the internet can.
Questions
Jev vs Zero-Shot NLI Classification: Common Questions
Can I just add descriptions to my label names?
You can lengthen the hypothesis, and it sometimes helps. It also changes the score distribution in ways that are hard to predict, because the model is judging entailment of a longer sentence rather than reading a definition. A criteria description is read as a definition, which is a different thing.
Is a decision model just a better-packaged zero-shot classifier?
It is the same problem attacked differently. Zero-shot NLI repurposes an entailment model one label at a time; a decision model scores the declared options in one pass and is trained so those numbers are usable as probabilities. The second is the part that is hard to replicate by packaging.
What about embeddings plus cosine similarity?
Good for retrieval, weak for classification with real criteria. Similarity is symmetric and topical, and it cannot express "this belongs in escalate because a legal threat was made". It also cannot do yes/no questions about a text at all.
How do I compare them fairly?
Label two or three hundred examples of your own, run both, and look at accuracy and at what the confidence scores do. If zero-shot wins on your data, use it — the aim is a system that works, not one that uses the newer idea.
More
Other Comparisons
The Cheapest Way to Settle It
Load a scenario, paste in your own text, and see what the distribution says. Editing costs nothing and needs no account.
