This infrastructure shift addresses growing concerns regarding data privacy and the rising costs of AI inference, which refers to the process of a trained model generating an output from new data.
By processing confidential materials locally, professionals such as lawyers can cross-reference private client data with public information without exposing sensitive details to external servers.
Furthermore, because local processing does not incur cloud service fees, Perplexity suggests this method can reduce the overall cost of complex tasks for both individual subscribers and enterprise clients.
The feature is currently available to Pro, Max, and enterprise subscribers using Apple Silicon Macs with at least 32GB of memory.
Users can choose from several local Large Language Models (LLMs), including versions of Google’s Gemma and Alibaba’s Qwen, which the Perplexity app installs automatically without requiring technical command-line setup.
While the company acknowledges that local models may be less capable than the most advanced cloud systems, a built-in visualization tool allows users to monitor their computer's hardware usage and token consumption to balance performance against privacy and cost.