Skip to main content
September 15, 20268:45

Hybrid Compute SOLVES Local AI

By Samuel Gregory

About this video

The era of sending every single byte of your sensitive data to a distant server is officially dead. In this video, I dive deep into Perplexity's new Hybrid Compute update for Apple Silicon, showing you how to save credits and protect your privacy by running AI models locally on your Mac. Key Takeaways: - Automatic delegation between local and cloud models for maximum efficiency. - Save up to 60% on cloud credits by utilising local tokens for file processing. - The Privacy Gate: A 0.66bn parameter model that identifies and blocks PII from leaving your machine. - Hardware requirements: Why 24gb to 32gb of RAM is the new baseline for AI power users. - Real-world testing: Cloud vs. Hybrid performance benchmarks on legal documents.

Perplexity Hybrid Compute: The Privacy Revolution Your Mac Has Been Waiting For

The era of sending every single byte of your sensitive data to a distant server is officially dead. With the launch of Perplexity Hybrid Compute, we are witnessing a fundamental shift in how artificial intelligence operates on personal hardware. This update brings automatic delegation: local models handle the sensitive data and privacy, while the cloud handles the heavy lifting and web research.

The Power of Local Sub-Agents

For the first time, we have a system that understands the nuance of 'where' a task should be performed. In my testing, running a complex legal document review through the cloud utilised 2.26k credits. By simply toggling on Hybrid mode, that same task only cost 773.8 credits because 12k tokens were processed locally on my machine.

This is not just about saving money: it is about data sovereignty. When the local sub-agent takes over, it reads your files and formats the data without that information ever touching an external server.

The Privacy Gate: Your Digital Bouncer

At the heart of this system is the Privacy Gate. This is a specialised 0.66bn parameter model that runs entirely on your Mac. Its sole job is to identify Personally Identifiable Information (PII) and ensure it stays behind your local firewall. Perplexity has utilised an open-source model from Hugging Face that benches incredibly well against much larger competitors, proving that you do not need a massive cloud cluster to maintain security.

The Hardware Tax

There is no such thing as a free lunch. To run these models effectively, you need Apple Silicon with a minimum of 24gb of unified memory. Whilst I tested this on an M5 Max with 124gb of RAM, Perplexity suggests that 32gb is the sweet spot for their flagship local model. If you are running a lighter machine with 16gb, you can still participate by using the Gemma 4b model, though your results may vary.

The Honest Truth About Speed

We must address the elephant in the room: local inference is slower. In my demonstration, the cloud-only task was nearly instantaneous in its initialisation, whereas the hybrid task took 7.73 minutes to complete. My recommendation is to reserve Hybrid Compute for tasks where speed is not the essence. If you have recurring morning reports or sensitive legal audits that can run in the background, this is a game-changer.

Perplexity has nailed the balance between local privacy and frontier model power. It is time to start treating our local hardware as the secure vault it was meant to be.

Transcript

This has become a hell of a lot more important now with Perplexity's latest update, hybrid computer. And what this finally gives us is automatic delegation of local models where it matters, such as privacy and data retention, and cloud models where they shine in computing and planning. But what I am going to show you is that whilst hybrid does save you credits, it does cost you something. And I am going to show you exactly that. Everything we will be going over will be in the time stamps below. We will do a hybrid versus cloud demo right off the bat so you can see this in action. Go over the system requirements, the pricing, actually installing and getting it up and running. We will discuss where the local tasks are actually initialised versus where the cloud tasks are initialised. We will talk about privacy gate, what are the different responsibilities for these different models. We will go over some gotchas. So definitely stay tuned for that one and then I will give you my final verdict. So thank you Perplexity for sponsoring this video. Let's get on with it. Now to set a baseline, what I am going to do is hit computer here. I am actually going to turn off hybrids. So we are just using the remote models. Going to attach some files and say check the DPA appendix and whether the promised clauses are covered in our insurance documents. Let's send that off. And I am getting updates on my phone so I can just go for a walk and double check things and review things if it needs to ask me a question. That used 2.26k credits altogether and worked for about 8 minutes. So, not too bad. Did we need to use Astra? I mean, this is legal work, so arguably yes. But this would probably be one of the most priciest models you can use. So, just take that with a grain of salt. So, with that baseline set, let's see what cost savings or any other savings we can count on by turning on that hybrid mode. So, when we are in a new task here, you just want to make sure your hybrid mode is actually selected on. And then you can choose what your local model is if you have downloaded a few of those ones. And then what your cloud model is. We are going to put Astra because of course that is the model we used for the completely cloud-based solution. And we are basically going to put in the same prompt. So, check the documents, make sure selected, touch the files, and letter it. And here we go. Here is the local sub agent working away. We can see the CPU, the memory, the GPU being 100% utilised there. And then actually how many tokens we are saving by using a local model where and when it is necessary. And here we go. I obviously have to blur this out because this is a genuine policy document. This is something I am actively working on right now. But what is really interesting is that you can see that it worked for 7.73 minutes. We used 773.8 credits. And now that is taken from my credit on my account, but it used 12k local tokens. Now, is the result significantly worse or better because it used a local model? You might be able to argue those points. However, the cost saving and the sensitivity saving is undeniable. If that isn't worth a like and subscribe, then I don't know what is. Let's take a look at the stuff that matters. So, hybrid compute is specifically an Apple silicon feature that Perplexity have launched in their app and it is designed to run on 24gb minimum of unified memory. 32gb for the best results. Now, I have an M5 Max with 124gb. Don't be deterred by the fact that I have a powerful machine. Perplexity tell you you can run it on 24gb and 32gb machines. Now, it is secure by default. I will be going over this, but they use a PII, that is personal identifiable information, a special classifier that runs on your machine to determine what is secure and personal information and what is not and what it can actually hand off to the cloud models. Now, how much does it cost? Now, you get it on Perplexity Pro for $20 a month. That gives you 4k bonus credits. On the max plan for $200 a month. You get 35k bonus credits with reoccurring 10k monthly credits. So, I do think this is the best version because you get those monthly reoccurring credits. However, there is a lot more flexibility with the pro plan. You can top up as and when you need it with an automatic top up once it reach below a certain threshold. Let me show you how you get this thing downloaded and installed and choosing the right model. If you go to the orchestrator here and you will get these setup commands here. If we set that up, it is going to describe to you a little bit about the initial task going to the cloud to understand the actual shape of the task, the structure of what needs to be done. And that will be the deciding factor on what runs locally on your Mac versus what runs continues to run in the cloud. Now, you're going to get a choice of three models here: the Perplexity model, which is the recommended one, which needs 32gb of RAM, which I have. Or if you have a slightly lighter machine, you might go for the Gemma 4b one, which needs 16gb of RAM. Or you might want to choose Quen 35b that needs 32gb of RAM. We are going to go for the Perplexity model here. And then that is going to just download the model. It is worth you leaving down in the comments what machine you have and what model you chose so other people can determine which is the best model for them. And then when that is done, we can head on all through and when we start a new task, we can just go and click hybrid here. So that means anything that will do, they will dedicate the tasks that are relevant to the local model to the local model and the cloud model to the cloud model. Now what I want to dig into here is where exactly it offset the tasks locally. So it has found the policy documents checking names and folders. Loading skills that it has access to. So it looks like this one here: read only insurance document inspection on this map exact folder. Now it looks like local sub agent started. So it is when it detected some insurance something quite sensitive that it is going to hand that off to the local agent. I think reading files especially is another one of those things it can do locally. That is not a taxing task. And what I'd strongly recommend as well, if we go into settings here and under local inference, you want to set your privacy gate, enable that. It downloads a small model that does the checking where it needs to divvy up that sensitive data. It is in of itself an open-source model on hugging face. And here is the model card for that actually. And it is a 0.66bn parameter model for PII detection and it benches very very well against some of the best models here. Plus is open source and runs completely locally on your machine and you can even download it on hugging face. Now it is worth saying that anonymised research material leaves the machine. So the data itself stays on the machine but the shape of the work doesn't. So while the local model handles all the reading, digesting, formatting, the cloud model handles the research, the reasoning, the heavy duty tasks, and of course all of the web research. So what is really the gotcha here? It really comes down to the fact that local inference is going to be a lot slower than a completely remote solution. As wild as that sounds, we are still not at the stage where local inference is like as fast. So, my genuine honest recommendation here is to hand this off to slow running tasks where you don't need a response immediately. They could be prescheduled as well. If you need something to run every morning, as an example, or again, you're out and about and you just receive an email that you need to check against some local files on your computer. You're not going to be back at your computer right in anytime soon. That is where I think this is just amazing. If speed isn't of the essence, this is going to become really handy. And just that reassurance that the secure tasks are going to remain on your machine and you know that there's some intelligent mechanism behind this to hand the relevant work off to those remote models and keep all the stuff that can be local. So I think this really does answer the question of does this utilise the best of both worlds, the local where it matters, the frontier remote models where it matters. I honestly think Perplexity have nailed this one. So, I am so thankful they sponsored this video. Links for everything will be down below. Like, subscribe if you haven't already. Until next time, keep on vibing.