On August 25, 2026, Perplexity released Portable Computer, a version of its Perplexity Computer agent platform that runs entirely on the user's own machine. Built together with NVIDIA, it starts on the NVIDIA DGX Spark desktop AI system. The model, the conversation, and the work in progress all stay on the device, and anything handled locally consumes no credits.

A design that keeps work off the cloud by default

Perplexity Computer is an agent platform that combines models, files, tools, and web access to complete multi-step tasks. Portable Computer moves that entire stack onto the device. The orchestrator, planner, tool router, scheduler, task queue, and local search index all run on the machine itself.

Some jobs still need the outside world: fetching current information, driving a browser, calling a connected app, or handing a difficult judgment to a stronger model. Those steps cross over to the cloud, but only after the user approves them. The system shows what would leave the device before sending anything.

Supported connectors include Google Drive, Gmail, Slack, and GitHub. Perplexity offers one example: triage GitHub issues on device against overnight bug reports that arrived in Gmail, then post the top three with suggested owners to Slack by morning. Local-first does not mean disconnected.

Dictation also runs locally. Using NVIDIA's Nemotron 3.5 ASR model, both transcription and file actions happen on the machine, so audio covering confidential matters never reaches the cloud. Code and tool execution run inside an isolated sandbox, which Perplexity says matches the security of the cloud version.

Rebuilding the scaffolding around a smaller model

The technically interesting part is that Perplexity redesigned the model and the harness together. A harness is the scaffolding around a model: the prompts, the tool set, and the orchestration logic. General-purpose harnesses assume a frontier model that can absorb long contexts, navigate a broad tool surface, and plan over long horizons. Push those assumptions onto a small on-device model and things fall apart quickly.

Qwen 3.8 27B, for example, advertises a 260,000-token context window, but Perplexity found in testing that its behavior degrades past roughly 100,000 tokens. So the team kept the system prompt short, limited the always-resident tool set, and split every other capability into skills that load only when needed. The harness also compacts stale context by summarizing it once a task runs long.

Connectors were handled differently as well. Services like Gmail and GitHub are normally exposed as MCP servers, but their tool definitions eat a large share of the context. Perplexity rewrote the most-used ones as compact command-line tools. It also added hooks that monitor the health of a task and trigger self-verification when something goes wrong.

Sandboxing is absolute. If the isolated environment is unavailable, the harness shuts itself down rather than running tools unprotected. That is a deliberate break from open-source harnesses that execute commands with the user's full permissions by default.

What the benchmarks show, and what gap remains

Perplexity published a research post the same day with its own evaluation results. On an internal benchmark of 53 everyday knowledge-work tasks, running the same Qwen 3.8 27B model, its harness scored 82.6 percent, compared with 77.6 percent for Pi and 74.0 percent for Hermes. It also used the fewest tokens, at 520,000. Swapping in PPLX 27B, a version of Qwen 3.8 27B post-trained for the harness, raises the score to 85.4 percent.

On BrowseComp, a web research benchmark of 1,266 tasks, the score reached 66.7 percent against 50.2 percent for Pi and 43.9 percent for Hermes. On ParseBench-100, which measures reading of PDFs and screenshots, it reached 65.1 percent versus 34.6 percent for Hermes and 13.9 percent for Pi.

The limits are visible too. On Terminal Bench 2.1, a set of hard coding tasks, the local model alone scored 59.6 percent. Enabling escalation, where the local model consults a stronger cloud model when it gets stuck, lifts that to 73.0 percent, still short of the 82.4 percent that Claude Opus 5 reaches on its own. API cost works out to 0.415 USD per run with escalation and 0.65 USD for Claude Opus 5 alone. In other words, escalation closes about three-fifths of the gap to the frontier at roughly two-thirds of the frontier's cost.

These figures come from Perplexity's own evaluations and have not been independently reproduced. The company does say it plans to release the 53-task benchmark publicly, with a technical report on the model training to follow.

Where it runs

Portable Computer is available to Pro and Max subscribers on the NVIDIA DGX Spark. Linux comes first, with Windows support following. Support for PCs with NVIDIA RTX GPUs is also planned, and setup is a one-click install from the Perplexity app.

The DGX Spark itself is built on the GB10 Grace Blackwell platform, with a 20-core Arm CPU and 128GB of unified memory in a desktop-sized box. The premise behind this launch is that hardware capable of running a 27B-class model continuously now fits on a desk.

Two models are available at launch, Qwen 3.8 27B and PPLX 27B, with NVIDIA Nemotron 3.5 Lightning, a 30B open model, announced as coming soon.

Summary

Portable Computer is an attempt to move local AI from "does it run" to "can it do the job." Perplexity rebuilt the scaffolding around the weaknesses of a small model, then routed only the hard parts to the cloud with user approval. That trade-off meaningfully widens the range of confidential work that can stay on the device. Coding still favors frontier models by a clear margin, but between improving hardware and fast-moving open models, that gap looks likely to narrow.