Comparison

Jev vs a Fine-Tuned Classifier

This is the comparison that most deserves to be taken seriously, and the one most often skipped. A fine-tuned DistilBERT or similar, trained on ten thousand of your own labelled examples, is fast, costs almost nothing per call, runs anywhere, and on its own task is very hard to beat.

The reason people still reach for a general model is not that classifiers are bad. It is that the labels, the training loop and the retraining loop are a permanent piece of infrastructure, and most label sets do not sit still long enough to amortise it.

Background

What a Fine-Tuned Classifier Gets You

Take a small encoder model, add a classification head, train it on your labelled data, serve it on a CPU. Inference is single-digit milliseconds, the cost per call rounds to zero, nothing leaves your network, and the model is yours forever. On a stable task with abundant labels this is not a compromise, it is the right answer.

The costs are all up front and all recurring. You need labelled data, which means annotation guidelines, annotators, and the discovery that your annotators disagree with each other about 15% of the time. You need a training pipeline, an evaluation set, and someone who notices when production drifts away from it. Adding a label means relabelling and retraining.

The honest summary is that a classifier is cheap per call and expensive per change, and a decision model is the other way round. Which one wins depends entirely on how often your label set changes and how much labelled data you already have.

Side by side

What Actually Differs

The trade is between cost-per-call and cost-per-change. Almost everything follows from that.

 Jev (via jev-ai.org)Fine-tuned classifier
Labelled data needed to startNoneHundreds to tens of thousands per label
Time to first working versionMinutesWeeks, most of it annotation
Adding or splitting a labelEdit a string and deployRelabel, retrain, re-evaluate, redeploy
Cost per callInput tokens at $0.042/1MEffectively zero on hardware you already run
LatencyNetwork round trip; upstream P50 ~0.22sSingle-digit milliseconds, in-process
Ceiling on a stable task with lots of dataGood zero-shot, no task-specific trainingUsually higher — it learned your exact distribution
Long tail and rare labelsHandles them from the descriptionPoor until you have enough examples of them
Data residencyLeaves your networkStays put

How to choose

Neither One Wins Every Time

Choose a decision model when

  • You have no labelled data yet, which is where every project starts.
  • The label set changes: product adds a queue, policy adds a category, a new failure mode appears.
  • Volume is moderate, or spiky enough that a per-call price beats owning capacity.
  • You need several different judgements about the same text, where a classifier per judgement means a model zoo.
  • The tail matters. Rare categories are exactly where a trained classifier has least to go on.

Train a classifier when

  • The task is stable, high-volume and well defined — spam, language ID, a fixed taxonomy that has not moved in two years.
  • You already have labelled data, or you get it free as a by-product of an existing workflow.
  • Per-call cost is the binding constraint. At hundreds of millions of decisions, a model you own is not close.
  • Latency has to be in-process: no network hop, no external dependency in the request path.
  • Data cannot leave your network, and you would rather train than self-host something larger.

We sell one of these two, so read the right-hand column with that in mind — and then run both on two hundred of your own labelled examples, which settles it better than any page on the internet can.

Questions

Jev vs a Fine-Tuned Classifier: Common Questions

Can I use one to build the other?

Yes, and it is the best answer for many teams. Ship the decision model on day one, keep its inputs and outputs, spot-check a sample for quality, and you have a labelled dataset accumulating as a by-product. Train the classifier when the volume justifies it and the labels have stopped moving.

How much labelled data do I actually need?

For a small encoder and a handful of labels, a few hundred clean examples per label gets you something usable and a few thousand gets you something good. The number that surprises people is the annotation-agreement number: budget for the discovery that your own guidelines are ambiguous.

Is a classifier more accurate?

On its trained task with enough data, usually yes. Off that distribution, on a new label, or on a rare category, usually not. Which of those describes your production traffic is the real question.

What about a hybrid?

Run the classifier for the volume and send its low-confidence cases to the decision model. You get the classifier’s cost on the easy majority and better behaviour on the tail, which is where the classifier was always weakest.

The Cheapest Way to Settle It

Load a scenario, paste in your own text, and see what the distribution says. Editing costs nothing and needs no account.