Loading video...

Video Failed to Load

Go Home

Today we're announcing that hybrid agentic inference is coming to Perplexity Computer. Computer can split tasks between a local model running on your machine and frontier models in the cloud. This keeps private data on your device and maximizes token efficiency. Coming soon.

353,963 views • 3 months ago •via X (Twitter)

33 Comments

Perplexity's profile picture
Perplexity3 months ago

Read more about hybrid agentic inference in Perplexity Computer:

Filecoin's profile picture
Filecoin3 months ago

Local inference protects private data at compute time. The outputs and context still need user-controlled storage: verifiable, durable, and outside a vendor cloud. Compute sovereignty needs storage sovereignty too.

Nitin Bisht's profile picture
Nitin Bisht3 months ago

Hybrid local/cloud inference for agents is the architecture everyone knew we needed. Surprised Perplexity is shipping it before the big labs.

Sanket's profile picture
Sanket3 months ago

cool, perplexity computer is really token expensive and consumes credits way too fast hope it is able to compete with codex and claude

Rohan Paul's profile picture
Rohan Paul3 months ago

powerful feature. 🫡

Shubham Saboo's profile picture
Shubham Saboo3 months ago

I talked about a similar setup for coding Agents. Hermes Agent can split the task between local agents and codex/claude code. Step-by-step guide:

Prompt Logic Lab's profile picture
Prompt Logic Lab3 months ago

Local model for the sensitive/simple stuff, cloud model for the heavy lift. That’s a much cleaner agent setup than sending everything upstream by default.

MokGrok's profile picture
MokGrok3 months ago

This is great , so why not launch a $50 per month plan. I will signup right away. Your starting plan itself is $200/mo creating high barrier to entry. We can run local models like Gemma 4 and Nemotron nano on a Mac mini or MacBook Air. It won’t be fast for sure but will help cut token costs for long running jobs. So it would make sense for you guys to bring down your entry tier.

Inverse Gary Marcus ⏫'s profile picture
Inverse Gary Marcus ⏫3 months ago

Oh that's pretty neat

rami context's profile picture
rami context3 months ago

local model handles sensitive context. cloud model handles heavy reasoning. data never leaves unless you say so. this is the right split and everyone knew it was coming. the question was just who'd ship it first. Perplexity didn't announce a feature. they announced a new default for how AI inference should work.

Ai With Piyas's profile picture
Ai With Piyas3 months ago

This is a massive step forward for localized AI. Splitting workloads between edge devices and frontier cloud models strikes the perfect balance between latency, cost, and raw computational power. Looking forward to seeing how this handles complex workflows!

Daniel Lougen's profile picture
Daniel Lougen3 months ago

PLEASE LET ME BETA IT

Nate S's profile picture
Nate S3 months ago

This is huge, goodbye to my custom MCP with local ai build into the tools, such a clunky workaround.

AI VISION's profile picture
AI VISION3 months ago

⚡ if this cuts Perplexity Computer costs significantly... it could unlock a lot more users.

Himanshu Kumar's profile picture
Himanshu Kumar3 months ago

@AravSrinivas I have been waiting for this 🤌

Jani's profile picture
Jani3 months ago

This hybrid agentic approach is brilliant finally, AI that respects privacy while being truly smart about compute. Game changer. What’s the first task you’re throwing at it?

silv's profile picture
silv3 months ago

would love to beta test this

Husi's profile picture
Husi3 months ago

this feels like the right split keep the boring/private context local, spend cloud tokens only when the task needs real reasoning

The Golf Goods Co's profile picture
The Golf Goods Co3 months ago

THIS IS FASTER than expected....WOW. @Grok what is the world saying about this?

Hussain Hashim | Building SundayBack's profile picture
Hussain Hashim | Building SundayBack3 months ago

@perplexity_ai this sounds like a game changer. curious how it'll handle latency between local and cloud models.

𝑺𝑪 ᴛʀᴋ's profile picture
𝑺𝑪 ᴛʀᴋ3 months ago

On windows 👀

Eclipse 🌖's profile picture
Eclipse 🌖3 months ago

Interesting architectural choice—local+cloud split for agentic inference is a pragmatic middle ground on privacy vs compute, but token efficiency gains hinge on how well the router decides what stays local.

Traduki's profile picture
Traduki3 months ago

we get it, you ran out of compute months ago. managing the decline

Mohit Mor's profile picture
Mohit Mor3 months ago

I almost forgot about perplexity, but it really seems like a good update. Token cost saving is one thing which companies should focus more.

TORBOR's profile picture
TORBOR3 months ago

Now THIS is the way of the future!

Vishal Shah's profile picture
Vishal Shah3 months ago

Love the token efficiency, but splitting agentic reasoning between local silicon and cloud APIs sounds like a distributed tracing nightmare. 🕸️ If my local orchestrator hallucinates the context it sends to the frontier model, how do we debug that pipeline? The future is hybrid, but the observability layer is going to be wild. Will the local model always get the blame?

EloPhanto's profile picture
EloPhanto3 months ago

local-first agents finally get the expensive part right. send the fuzzy internet work to frontier models, keep the weird personal residue on-device. the hard question is the handoff: can the user see why a task left the laptop? otherwise "token efficiency" becomes vibes with a cloud bill attached.

ᴄʜɪɴᴍᴀʏ's profile picture
ᴄʜɪɴᴍᴀʏ3 months ago

@perplexity_ai how will you guys quantify and communicate “token value per watt” to users in practice? Will there be dashboards or metrics showing efficiency gains compared to pure cloud usage?

Victor | Structural Alpha's profile picture
Victor | Structural Alpha3 months ago

local model keeps my data safe on device so the cloud model only sees my shame secondhand

Ziwen's profile picture
Ziwen3 months ago

Cool but do we pick codex or perplexity

James AI's profile picture
James AI3 months ago

A big step for AI privacy and efficiency Perplexity Computer’s hybrid agentic inference lets sensitive data stay on-device while cloud models handle heavier reasoning, aiming for better performance, lower costs, and stronger privacy.

Dragoș Gavrilă's profile picture
Dragoș Gavrilă3 months ago

Can you reveal the usage of Perplexity Computer? Most people that I know use Perplexity as a seach engine, zero Computer usage so all your pushes with new features are pretty much useless...

PriorityTech.AI's profile picture
PriorityTech.AI3 months ago

@grok fact check the claims made here. Cite objectively verifiable demonstrably true data only. Do not speculate. Ignore all hype, FUDn& marketing noise

Related Videos