Skip to main content
October 3, 202611:06

Jev, But Free and On Your Mac

By Samuel Gregory

About this video

Most people are using massive AI models for tasks that a tiny model could do 10x faster. In this video, I explore the world of Decision Models like Jev, Tev, and Nimble to show you how to build lightning fast, deterministic AI workflows that run entirely on your local hardware for free. Key Takeaways: - Decision models are System 1 AI designed for speed and specific classification tasks. - Cloud models like Jev can be 8.8x faster than Claude Haiku. - You can run models like Tev 0.8b locally with response times under 100ms. - A hybrid workflow allows you to use local AI first and only fall back to expensive cloud models when confidence is low. - We demonstrate a real-world email automation setup using n8n and Ollama.

Most AI implementations are burning money for no reason.

We have reached a point where using a frontier model like Claude or GPT 4 for basic classification is like hiring a rocket scientist to sort your post. It is overkill, it is slow, and it is draining your budget. The rise of decision models, specifically System 1 AI, is changing the landscape by providing deterministic, lightning fast responses for a fraction of the cost or even for free on local hardware.

What Are Decision Models?

Decision models are specialised AI architectures designed to provide specific answers to specific questions. Unlike traditional LLMs that predict the next token in a conversational sequence, these models act as classifiers. They take a state (like an email or a support ticket) and match it against a set of criteria.

The key advantage is determinism. You get a categorised output with a confidence score, rather than a wordy explanation you have to parse.

The Speed Gap: Claude vs Jev

In recent testing, Claude Haiku, known for being one of the fastest frontier models, took 4.3s to answer a simple yes or no query. In contrast, Jev provided the same answer in 0.33s. That is 8.8x faster. When you scale this across 1m requests, the time savings are astronomical.

Going Local with Tev and Nimble

The real magic happens when you move these workflows to your own hardware. Models like Tev (available in 0.8b and 4b sizes) and Nimble (a 9b model based on Qwen) allow you to run these decisions on a local machine.

During a real-life test of email filtering:

  • Tev 0.8b responded in 97ms.
  • Jev responded in 329ms.
  • Claude trailed behind at 2.4s.

Real World Implementation

I implemented this on an M1 Mac using n8n to poll emails. The workflow is simple:

  1. Fetch the email.
  2. Build a request for the local Ollama instance.
  3. Ask specific questions: Is this a sponsorship? Is it urgent?
  4. Route the decision based on the confidence level.

If a local model like Tev 4b shows low confidence (below 0.5), the system can automatically hand the task off to a larger cloud model. This hybrid approach ensures you only pay for high level intelligence when you actually need it.

Conclusion

The era of the "everything model" is ending for enterprise workflows. By utilising small, local decision models, you can achieve near instantaneous results for $0 in compute costs. If you are still waiting 5 seconds for a chatbot to tell you an email is spam, you are doing it wrong.

Transcript▾

Jev has taken the AI world by storm and has spawned a plethora of decision-making models that save on time, cost, resources, and the key selling point is deterministic. And in this video, I am going to put the two types of models to the test where we find out that Claude answers a yes or no question in 4.3s, whereas the decision-making models answered it in a fraction of that. But if you are a fan of the channel, you know that we like to get down and dirty with local AI models. So I was wondering, can we run this locally on our own hardware? And the answer is yes, you can. So in this video, we are going to go over what are decision models. We are going to run some comparisons between local models and cloud-based models, and we are going to run it in a real-life workflow. So stay tuned for that. So let us talk about what decision models are and what you can do with them. In this case, we have got a user that says, "Hi, I was charged twice for my pro subscription this month. Can you sort this out?" We pass the model a question with a set of potential options. And if we just look at the actual code here, you can see that we ask for the model using Jev in this instance, and the state is I was charged twice, can you sort this out? And then we are asking it, "What are they asking for? What is it that this user is asking for? Is it a refund, a replacement, information, or cancellation?" So we run this here, you can see that Jev has decided in pretty much 0.5s, whereas Claude Haiku, their smallest, fastest model, took 4s to respond. Now that is 8.8x faster. And it is pretty much decided that they want a refund. Claude even has a 0.95 confidence level that this is a refund. Jev decided instantly that it was indeed a refund. So this is what the sort of speed that we are talking about, and the price as well. Like I said, they do not even charge for output tokens. So if that comparison alone isn't worth a like and subscribe, I do not know what is. But since then, there has been a plethora of local models. We have got Tev 1 and Nimble. Nimble was a decision model fine-tuned from Qwen 3.5 9b, whereas Tev is 3.5 4b. They both take 64 questions at the same time, which is really, really cool. Nimble is Apache 2.0 licence, whereas Tev is MIT licence. Now, licencing is not my wheelhouse. However, if you are building this for production, then it is worth checking out and understanding what the two differing licencing allows you and does not allow you to do. Now, Nimble have a 9b parameter model, and you can sort of see the structure of the quest here. This is obviously running all on Ollama, and it is the system one endpoint. This goes into what they are talking about the system one. They Jev will be releasing a system two version, which is a thinking model, but this system one seems to be proliferating through the the industry as a new endpoint. ask for the Nimble model on Ollama. We pass a state, the the thing that we are acting upon, and we are giving it a set of our questions that we want to know against that state. And in this instance, we want to know, "Does the state contain a greeting?" And the criteria is either true or false. Tev, they have two models. They have a 4b and a 0.8b. And I would say in this instance, you want to try both models. You want to see if given your use case and what is producing the best results. This is up to you to decide, but obviously the 0.8b is going to be a hell of a lot faster than the 4b. However, 4b is already half the size of Nimble. The actual data structure and the the the request is all the same. So, let us run all these models against each other, including Jev, and how they compare to standard LLMs. So, here we are. This is a use case of filtering out emails for me that which emails are sponsorship emails. Now, I have got all of the models loaded here. We have got Tev, a 0.8, 4b, Nimble 9, Jev, and Claude. And then we are going to run against these, and we can see the results here as they filter through. Jev is steaming ahead there. It is pretty much already done. The local models just falling a little bit behind it, but already done. And Claude is following very slowly behind. Nimble 9b seems to be having an issue here. I am wondering if we have a problem. There we go. There it is. I would not worry too much about that. We can look at the actual results here. The average is 1.3s cuz that time to first load is something that can be quite slow. Once it is going, it is going. So, 1.3s, 256ms for Tev, which is, you know, again, half rapid. 97ms for the Tel Aviv one, 329ms for Jev, but 2.4s for Claude. Now, looking at some of the results here, pitch This is what the truth is. So, this is a pitch. This is a pitch. This is a pitch. We should be seeing near enough 100% on all of these All of these emails whereas the rest we are sort of seeing a low percentage. Now, Tev Tev 1 is showing that it thinks this is a sponsorship pitch at 74% confidence level on a sponsorship pitch, but we know that it is not. This is obviously one of the downsides of a smaller model. It is not as smart, you know? So, that is why we are seeing those reds here, but generally across the board, all of them are getting it right until we get to Matt, the co-host one. Now, this is an email here that is a asking about sponsorships, but is not actually a sponsorship email. And these are where the models are falling short. It definitely is not a sponsorship email, we know that, whereas the models tend to think it does. And then the Claude model is absolutely certain because it can do a little bit more inference. And so with that, we can build potentially a system which, if the if the confidence level is around 50 or there's a threshold below 90%, then we hand it off to a Claude model or a or a Jev even. You know, if the local model cannot figure it out, then we hand it off to Jev. If Jev cannot figure it out, we hand it off to the the frontier model or something like that, which is able to get pretty quickly. So these are not these magical models that just know everything, but it does a damn good job at making these decisions for us. So let us look at real-life workflow here. This is a local instance on an n8n running on my M1 Mac here. And essentially, I am polling my emails. We build the request and it is again, what I am simply asking, what kind of email is this? And I want to know if it is a sponsorship email, it is a consultancy lead, it is a podcast viewer, admin, user, whatever. I can set up my criteria however I like on this. Basically, you could pass it a few questions and get a few results from them. So, does this email need a reply, a personal reply from me? And then it is either true or false. How urgent is it? Is it Can it wait a week, within a few days, today? And then ask for rates. Does this sender ask for rates, prices, or a media kit? So there's a few questions I can pose against the original statement that they made. And now I am deciding, I am sending the Ollama my request that I have just built up in the previous step. And with that, we have got to we can parse the decision. Again, there is just some code here that actually runs through and just formats everything nicely, and then we route those decisions based on the the parsing of that data. And then, we send it off. And this builds the actual response to the email. I draft the email, and I label that draft as well. If it needs a review, then it sends a message to me. And then, if it is just it simply just needs a label on it, then I will just file it away and put a label on it. So, with that, let let us just test this out. So, I can say something like, "Hey, yo, we think you'd be a great person to work with. Are you taking sponsorships? We think your audience will love it." Send that off now. And on here, if we go to executions, if that's if we bring this up here, you can see this running. I mean, that was literally instantaneous. So, it is a it needs my reply. It is a sponsorship opportunity. There's the subject and who it was sent from. And then, I should respond within a few days. And just quickly here, here's what the response looks like. So, we've got my answers to my category question. It comes back with a type, which is choice. The choice is indeed sponsorship. And these are the probabilities that it has. So, the highest number wins, which is a sponsorship. And it's got a 0.51 confidence level in that it's a sponsorship. Things it do need to reply. Urgency, it thinks that I need to respond in a few days. And it does not think I need to ask for rates there. And that is an a prime example of where we would use a decision model and the sort of speed we are getting from that. Did not cost me anything as well because it's run locally, but if this was Jev, then you can sure as hell imagine that this wouldn't cost an awful lot whatsoever. Now, here's the Ollama benchmarks. So, Tev 4b is scoring an accuracy of 73%. Nimble is going at 75, so a little bit better. However, it is a lot bigger, and Jev is scoring the best with 76%. But, it's worth knowing that confidence isn't a sign of accuracy. It's just how confident is it against the answer. So, the ultimate fix here is to run these models against the sorts of data sets that you're going to be using and and seeing how accurate it is for what you want to use it for. That's what these results show. So, ultimately, Jev is going to be the better solution here. However, when you get a chance to run it for free on your local machine, then the answer really is just down to what you want. So, there we go. That I hope that gives you a better idea of what Jev is and these decision models, and of course, how they differentiate, how you can get them running locally on your machine, sort of performance you can expect from them, as well as a real-world example on how I would use a decision model. Next, I'll build a decision model that chooses between local and cloud-based models, so you only pay when you need to. Like, subscribe if you haven't already. Until next time, keep on vibing.