Missed our launch?

Watch it here
Oumi AI

Jev vs. Fine-Tuning on BANKING77: Do You Need Jev or Just a Small Fine-Tuned Model?

When to use Jev for classification without fine-tuning, and when a small fine-tuned model is simpler. A real-world comparison on the BANKING77 benchmark.

By Ryan Arman and Mary Reagan

September 28, 2026

Share
AI summary
Markdown

We recently fine-tuned Qwen3 4B Instruct on BANKING77 in Oumi in about 30 minutes and achieved approximately 93% accuracy. That result caught our attention because BANKING77 is also being used to demonstrate a very different approach to classification: Jev.

Jev is an interesting new approach to AI decision-making. Instead of generating open-ended text, it returns a decision from a defined set of choices. One of its most appealing promises is that you can use it for classification without task-specific fine-tuning. But BANKING77 shows an important tradeoff between avoiding training and keeping your production system simple.

What it takes for Jev to reach 92.4%

BANKING77 is a public benchmark with 77 banking customer-support intents, including card payment issues, cash withdrawal problems, transfers, and account questions.

In an experiment with Jev, the model achieved roughly 79% accuracy on an initial screening set without labeled examples in context.

To get to 92.4% accuracy on the full BANKING77 test set, the system added another component.

For every test query, it:

  1. Searched the labeled BANKING77 training set,
  2. Found the 24 most relevant examples,
  3. Included those examples and their correct labels in Jev's input, and
  4. Asked Jev to make its classification.

That's a strong result. But it's useful to be precise about what it represents. The 92.4% system isn't simply Jev classifying a query on its own. It's retrieval-augmented few-shot classification + Jev. And that changes the production tradeoff.

You're Avoiding Training, But Adding Retrieval

There is nothing inherently wrong with this architecture.

In fact, it can be a great choice if your categories change constantly or you want to stand up a new classification system without training and deploying a task-specific model. But if you already have thousands of labeled examples, it's worth asking a different question:

Is it simpler to retrieve labeled examples for every request, or teach a small model the task once?

With the retrieval approach, your labeled dataset effectively becomes part of your production inference system. You need to maintain the examples, search them for every request, send the selected examples to the model, and think about what happens when examples or categories change. The alternative is to encode that information into the model through fine-tuning.

We tried the fine-tuning approach in Oumi

Using Oumi, we fine-tuned Qwen3 4B Instruct on BANKING77.

The result:

  • ~30 minutes of training.
  • ~93% accuracy.
  • No retrieval step required at inference time.

That puts a small fine-tuned model at roughly the same accuracy as the retrieval-augmented Jev system while requiring a much simpler inference path. The original BANKING77 work also reported 93.66% accuracy with a fine-tuned BERT classifier, so the idea that this task is well suited to fine-tuning isn't new.

What has changed is how easy it has become to train and deploy these specialized models using Oumi. Fine-tuning a 4B model no longer needs to be a multi-week ML project. In our experiment, it took about half an hour.

And you don't have to give up generative capabilities

Another interesting benefit of using a small language model instead of a traditional classifier.

It doesn't have to return only:

card_payment_wrong_exchange_rate

You can ask the model to return structured output with the class and an explanation:

Intent: card_payment_wrong_exchange_rate
Reason: The customer is asking why the exchange rate used for a card purchase differs from the expected rate.

You still get the behavior of a classifier, but you retain the capabilities of a language model when they are useful.

So when does Jev make sense?

There are absolutely cases where avoiding fine-tuning is attractive. If your categories change constantly, you are prototyping a new taxonomy, or you already have retrieval infrastructure in production, updating examples may be easier than retraining a model.

That's Jev's interesting promise. But BANKING77 illustrates the other side of the tradeoff.

If you already have a relatively stable taxonomy and thousands of labeled examples, you may not need to carry those examples through your inference pipeline forever. You can train a small model once and let the model learn the task.

At Oumi, we're increasingly seeing that small specialized models are cheap and fast enough to train that fine-tuning should be considered alongside prompt- and retrieval-based approaches, not treated as a last resort.

Jev asks an interesting question: What if you didn't need to fine-tune a classifier?

BANKING77 on Oumi suggests an equally interesting question: What if fine-tuning the classifier takes only 30 minutes?

FAQ

Is Jev better than fine-tuning for BANKING77?
Not clearly. Jev reached 92.4% with retrieval-augmented few-shot classification, while the Oumi fine-tuned Qwen3 4B Instruct result reached approximately 93% without retrieval at inference time. The better choice depends on taxonomy stability and operational constraints.
When should you fine-tune a small language model?
Fine-tune when you have a stable task, labeled examples, and a repeatable deployment path. It can give you a simpler production system while preserving structured and generative outputs.
When does Jev make more sense?
Jev is attractive when categories change frequently, you are testing a new taxonomy, or retrieval infrastructure is already in place.
Jev vs. fine-tuning: which should you choose?
Choose Jev when your taxonomy changes often, you need to prototype quickly, or your team already operates retrieval infrastructure. Choose fine-tuning when the taxonomy is stable, you have labeled examples, and you want a simpler inference path with no per-request retrieval step.

For this BANKING77 comparison, the key tradeoff is not training versus no training. It is whether you want to maintain retrieval as part of inference or encode the task into a small specialized model. The Jev experiment reported a mean elapsed time of 0.533 seconds and 0.713 seconds at p95 for its full implementation, including retrieval, scheduling, and network communication. Those figures describe that implementation, not universal Jev latency.

The State of
Owned AI

Which AI should you own

Research report · Coming soon

The State of Owned AI: Which AI Should You Own

Where enterprise AI creates proprietary data, measurable value, and a learning advantage.

600+

AI decisions mapped

300+

external proof points

39

real-world implementation stacks analyzed

13,700+

senior leaders represented

15+

benchmarks audited

It's not published yet. Get on the list and you'll read it before everyone else.

By clicking “Get early access”, you agree to Oumi processing your personal data in accordance with its Privacy Notice.