CORE: Reducing UI Exposure in Mobile Agents via Collaboration Between Cloud and Local LLMs
Gucongcong Fan, Chaoyue Niu, Chengfei Lyu, Fan Wu, Guihai Chen
TL;DR
CORE addresses the privacy risk of uploading full mobile UI context to cloud LLMs by introducing an asymmetric cloud-local collaboration framework. It uses layout-aware XML block partitioning to structure UI content, a co-planning phase where local LLMs propose sub-tasks and cloud LLMs select the best, and a multi-round co-decision-making phase where the cloud progressively accumulates blocks for confident decisions. Empirical results on DroidTask and AndroidLab show up to 55.6% UI exposure reduction with small task performance penalties, along with substantial reductions in sensitive UI elements uploaded (up to 70.49%), demonstrating a practical privacy-preserving trade-off. The approach is modality-agnostic and extensible to multimodal inputs, suggesting broad applicability for privacy-preserving mobile automation.
Abstract
Mobile agents rely on Large Language Models (LLMs) to plan and execute tasks on smartphone user interfaces (UIs). While cloud-based LLMs achieve high task accuracy, they require uploading the full UI state at every step, exposing unnecessary and often irrelevant information. In contrast, local LLMs avoid UI uploads but suffer from limited capacity, resulting in lower task success rates. We propose $\textbf{CORE}$, a $\textbf{CO}$llaborative framework that combines the strengths of cloud and local LLMs to $\textbf{R}$educe UI $\textbf{E}$xposure, while maintaining task accuracy for mobile agents. CORE comprises three key components: (1) $\textbf{Layout-aware block partitioning}$, which groups semantically related UI elements based on the XML screen hierarchy; (2) $\textbf{Co-planning}$, where local and cloud LLMs collaboratively identify the current sub-task; and (3) $\textbf{Co-decision-making}$, where the local LLM ranks relevant UI blocks, and the cloud LLM selects specific UI elements within the top-ranked block. CORE further introduces a multi-round accumulation mechanism to mitigate local misjudgment or limited context. Experiments across diverse mobile apps and tasks show that CORE reduces UI exposure by up to 55.6% while maintaining task success rates slightly below cloud-only agents, effectively mitigating unnecessary privacy exposure to the cloud. The code is available at https://github.com/Entropy-Fighter/CORE.
