More demanding tasks—like deep research
Perplexity has launched a new feature that allows users to split AI tasks between cloud servers and local devices, aiming to reduce operational expenses while maintaining performance. The update, announced in early September 2026, enables intelligent routing of workloads based on complexity and sensitivity, letting lighter tasks run locally and heavier ones in the cloud. This hybrid approach mirrors strategies used in enterprise computing but brings them to consumer-facing AI assistants for the first time at scale. The system evaluates each request in real time, deciding whether to process it on the user’s device or in Perplexity’s cloud infrastructure. Simple queries, such as summarizing short texts or answering factual questions, can be handled locally using smaller, optimized models.
Breaking news
Dell Unveils 14S Laptop to Compete With Apple’s New Budget Model
Microsoft is enabling a Windows 11 security feature that can hurt gaming performance
Dell Unveils Colorful New 14S Laptop Aimed at Students
Apple Unveils Eight New Devices in September 2026More demanding tasks—like deep research, multi-step This division helps minimize latency for everyday use while cutting down on cloud computing costs associated with constant high-volume processing. How the Split-Processing Model Works Perplexity’s technology uses a lightweight classifier to assess the computational demand of each incoming task. If the model determines the task can be completed accurately with a local model—such as a distilled version of its main AI—it processes it on-device. Otherwise, it securely transmits the request to the cloud. The company says this method reduces unnecessary cloud usage by up to 40% in typical user scenarios, based on internal testing. Privacy is also enhanced since sensitive data never leaves the device when processed locally. Can Users Control Where Tasks Are Processed? Currently, the routing decision is automated, but Perplexity plans to introduce user controls in a future update.
These would allow individuals to prioritize either speed, privacy
These would allow individuals to prioritize either speed, privacy, or cost savings by adjusting how aggressively the system uses local versus cloud resources. For now, the feature operates transparently in the background, with no visible indication to the user about where a task was handled. The update is rolling out gradually across Perplexity’s mobile and desktop applications. Frequently Asked Questions Does this feature work offline? Yes, tasks routed to the local AI can be processed without an internet connection, though initial setup and model updates require online access. Will this affect the accuracy of responses? No, Perplexity states that accuracy remains unchanged because the system only uses local models for tasks they are proven to handle reliably. Is this available on all devices? The feature is being rolled out to devices with sufficient processing power, starting with newer smartphones and laptops, and will expand to older models over time.