HN Simulatornew | past | comments | lists | submitlogin

For _your_ classification it's unacceptable. The OP seems to have anticipated this and mentions you can fine tune it for your use case. Did you try that?

I don't think the point is to displace Jev, but to show it's possible to build an MVP on open weights without years of work and millions of dollars.

Why (presumably) an engineer would dismiss exploring a lightweight, custom alternative to locking into a fashionable PaaS, I'll never know.

help



> The OP seems to have anticipated this and mentions you can fine tune it for your use case. Did you try that?

You can already so that with classification models such as ModernBERT, at 0.4B.

Jev's value is its zero shot performance without having to fine-tune.


I am sure I am missing something obvious here, but why is that valuable? Like, what kinds of projects are there where you need to classify stuff but are unable to make a bespoke model targeting the specific problem?

For my org, it meant we could trial classifiers across various internal systems with little to no engineering effort. In one case we ended up building our own classifier instead of Jev, but in others we kept Jev because it was zero-effort for a great impact.

But like how often are you gonna do this in general? Why does dev time or effort really matter here when either way you are building something to just, you know, actually use going forward? Its not like one needs to build a new classifier everyday.

Because a lot of people employed as engineers can’t build classifiers. They’re able to glue libraries and tools together, build UIs, and write APIs, but don’t have the curiosity or creative problem solving to learn and master a new domain. Even when “master” is scoped to something like this.

On top, most EMs wouldn’t take a risk on an exploration of something “unknown” (to them) and couldn’t get buy-in from a PM.

I say this as an EM. Interview hundreds of people and, while, yes, some people don’t interview well, you might be shocked at the level of creative thinking. Even when “creative” is narrowly scoped to “this is a solved problem in a related domain.


I guess this makes sense. I am no business guy, but if my product/company was focused on some sort of classification problem, my naive intuition would be to focus on hiring guys that can do it, rather than try to make the problem easier for them. But perhaps at the end of the day this is still cheaper? It's the same reason why we use Postgres instead of hire database experts to build something special?

> if my product/company was focused on some sort of classification problem

It's more likely the product is focused on something else, but a classification model could come in handy...


Tons. For example, most web and mobile app developers won’t know where to start with making a bespoke model (and would likely have no interest in making one), but they will have lots of usecases for a classifier.

its a 1 hour project on claude code on ur local machine. lol...

Only for folks who already know what they’re doing.

What I said will only make sense if you take yourself out of your current context and think entirely from the perspective of someone who knows little-to-nothing about ML.

It’s the same mistake folks on HN made when Dropbox launched, drawing comparisons to rsync and other Unix tools as if they were somehow equivalent.


You need to have enough high quality data to train with, knowledge how to do it, and developer time.

In practice, that's enough of a barrier to not even try the approach on a number of cases where it might potentially be useful.

I wouldn't be surprised if Jev turned out to be a "gateway drug" that validates approach on a use case, the team gathers experience and labeled data, and switches to an in house locally tuned model to minimize costs.


Exactly. And that labeled data could be collected by just recording what they feed jev and what the decision is.

> unable to make a bespoke model targeting the specific problem?

There is a fixed cost (and some maintenance) to e.g. fine tuning ModernBERT.

Maybe once you include all of that it might be a half-day to a day of engineering time to set everything up in a maintainable fashion.

For Jev, it takes all of 30 seconds of prompting. And it's not that much more expensive to deploy vs. a BERT model.


The biggest disadvantage of Jev is that it's a proprietary product and you need to send them your data. A bespoke solution makes much more sense in many scenarios.

Anything where you don’t have a decent amount of training data.

Not sure the task at hand here. But if it doesn’t require any reasoning/thinking and it’s just a classification task, it’s worth a shot to look into training your own classifier

I’ve run some benchmarks. Using embeddings + logistic classifier, the architecture matches or beats Jev and Laya in all basic classification tasks (datasets tested: AG News, Emotion, MASSIVE Intent, Banking77) The type of task in which it does really well, especially against Laya, is classification with >50 classes

The classifiers also run in <1ms, so they can be very fast and precise at the same time

But this architecture has no “reasoning”, so it performs rather poorly on tasks that require it, like the ones from the XLNI dataset (Jev/Laya do a lot better on this one)

For the latter cases, you could use add a local lightweight LLM, something like a Gemma model. Or even some basic MLP, depending on the tasks/data


What do you use to determine that a particular task in a heterogeneous pile of tasks requires reasoning? The logistic classifier itself is too dumb to recognize the details of the problem that make it reasoning-sensitive (IIRC recognizing the “fiddliness” of a given problem requires a recognizer at least as complex as the problem itself.) And if you’re using the lightweight LLM for that, then you may as well skip the classifier and just use the LLM all the time, since that eval step is already going to be dominating your response time anyway.

My understanding of Jev is that it’s a replacement for the LLM you’d necessarily need to use to identify reasoning-sensitive workloads in a heterogeneous mix, where Jev will be cheaper than an actual LLM and so act as an actual optimization / de-bottlenecking change.


I’m in the process of piecing together the different task/dataset-specific classifiers

Depending on how much overfit, you can go from routing deterministically based on features/shape of the input data, all the way to training a routing model (which could be a classifier too). I’ll need to experiment to find the best approach

For completely unseen/unexpected, I’ve also experimented routing to a local LLM: request comes in, if there’s a marching classifier, send it there, otherwise send to LLM+training. As the system learns more tasks, the % of requests that go to the LLM go down over time


"Using embeddings + logistic classifier, the architecture matches or beats Jev and Laya in all basic classification"

Have I understood correctly that you trained only the logistic classifier, but didn't need to train the embedding model?

If so, I'm curious whether you compared that approach (A) with:

B) Jev only, with a single output.

C) Jev with multiple outputs fed into a logistic classifier.

Obviously C has cons (can't be self-hosted, needs some up-front work on deciding the shape of the output) but it might be somewhat more interpretable. (And I suppose it might have better performance?)


You are correct, I didn’t train the embeddings model

Here's a gist with code you can use to test the Banking77 dataset: https://gist.github.com/nicobrenner/056a5aaff5d0119c0032ecda...

The gist uses BAAI/bge-large-en-v1.5, which is 1.2GB approx. You can replace it for all-MiniLM-L6-v2 (91 MB @ fp32 or 45 MB quantized fp16) small enough for mobile/edge. With all-MiniLM-L6-v2 it still gets 93.0% on Banking77, only 1.3 points behind bge-large at 15x smaller

I haven’t compared different ways of sending requests to Jev

The data to train the classifiers comes from the datasets used to test them (not from Jev)


I have not personally reviewed the benchmarks, but recall seeing a post where a Linear SVM with bag-of-words features outperformed Jev on many simple NLP classification tasks, like the ones you cite. You don't even need embeddings!

I've used LLM's to classify things for numerous projects and it's always really hard to beat linear classifier/decision tree over simple embeddings or even the sklearn HashingVectorizer. However, you DO need to trust your validated data, ground truths for this to work - but you should have these anyway to validate a Jev or similar solution.

Needing fine-tuning for the use-case completely changes the product category

You don't seem to understand what the point of Jev is when you say "you can fine tune it for your use case". Building your own classifiers for your specific business problems is the type of work we all used to do back in 2016 or so. It costs very much. Jev is a cheap and fast general purpose classifier.



Guidelines | FAQ | Lists | API | Security | DMCA | Apply to YC | Contact

Search: