Loading video...
Video Failed to Load
Today we're announcing that hybrid agentic inference is coming to Perplexity Computer. Computer can split tasks between a local model running on your machine and frontier models in the cloud. This keeps private data on your device and maximizes token efficiency. Coming soon.
353,963 views • 3 months ago •via X (Twitter)
33 Comments

Read more about hybrid agentic inference in Perplexity Computer:

Local inference protects private data at compute time. The outputs and context still need user-controlled storage: verifiable, durable, and outside a vendor cloud. Compute sovereignty needs storage sovereignty too.

Hybrid local/cloud inference for agents is the architecture everyone knew we needed. Surprised Perplexity is shipping it before the big labs.

cool, perplexity computer is really token expensive and consumes credits way too fast hope it is able to compete with codex and claude

powerful feature. 🫡

I talked about a similar setup for coding Agents. Hermes Agent can split the task between local agents and codex/claude code. Step-by-step guide:

Local model for the sensitive/simple stuff, cloud model for the heavy lift. That’s a much cleaner agent setup than sending everything upstream by default.

This is great , so why not launch a $50 per month plan. I will signup right away. Your starting plan itself is $200/mo creating high barrier to entry. We can run local models like Gemma 4 and Nemotron nano on a Mac mini or MacBook Air. It won’t be fast for sure but will help cut token costs for long running jobs. So it would make sense for you guys to bring down your entry tier.

Oh that's pretty neat

local model handles sensitive context. cloud model handles heavy reasoning. data never leaves unless you say so. this is the right split and everyone knew it was coming. the question was just who'd ship it first. Perplexity didn't announce a feature. they announced a new default for how AI inference should work.

This is a massive step forward for localized AI. Splitting workloads between edge devices and frontier cloud models strikes the perfect balance between latency, cost, and raw computational power. Looking forward to seeing how this handles complex workflows!

PLEASE LET ME BETA IT

This is huge, goodbye to my custom MCP with local ai build into the tools, such a clunky workaround.

⚡ if this cuts Perplexity Computer costs significantly... it could unlock a lot more users.

@AravSrinivas I have been waiting for this 🤌

This hybrid agentic approach is brilliant finally, AI that respects privacy while being truly smart about compute. Game changer. What’s the first task you’re throwing at it?

would love to beta test this

this feels like the right split keep the boring/private context local, spend cloud tokens only when the task needs real reasoning

THIS IS FASTER than expected....WOW. @Grok what is the world saying about this?

@perplexity_ai this sounds like a game changer. curious how it'll handle latency between local and cloud models.

On windows 👀

Interesting architectural choice—local+cloud split for agentic inference is a pragmatic middle ground on privacy vs compute, but token efficiency gains hinge on how well the router decides what stays local.

we get it, you ran out of compute months ago. managing the decline

I almost forgot about perplexity, but it really seems like a good update. Token cost saving is one thing which companies should focus more.

Now THIS is the way of the future!

Love the token efficiency, but splitting agentic reasoning between local silicon and cloud APIs sounds like a distributed tracing nightmare. 🕸️ If my local orchestrator hallucinates the context it sends to the frontier model, how do we debug that pipeline? The future is hybrid, but the observability layer is going to be wild. Will the local model always get the blame?

local-first agents finally get the expensive part right. send the fuzzy internet work to frontier models, keep the weird personal residue on-device. the hard question is the handoff: can the user see why a task left the laptop? otherwise "token efficiency" becomes vibes with a cloud bill attached.

@perplexity_ai how will you guys quantify and communicate “token value per watt” to users in practice? Will there be dashboards or metrics showing efficiency gains compared to pure cloud usage?

local model keeps my data safe on device so the cloud model only sees my shame secondhand

Cool but do we pick codex or perplexity

A big step for AI privacy and efficiency Perplexity Computer’s hybrid agentic inference lets sensitive data stay on-device while cloud models handle heavier reasoning, aiming for better performance, lower costs, and stronger privacy.

Can you reveal the usage of Perplexity Computer? Most people that I know use Perplexity as a seach engine, zero Computer usage so all your pushes with new features are pretty much useless...

@grok fact check the claims made here. Cite objectively verifiable demonstrably true data only. Do not speculate. Ignore all hype, FUDn& marketing noise


