PerceptionGuard: Privacy-Aware Split Inference for Multimodal Mobile Applications
DOI:
https://doi.org/10.30871/jaic.v10i4.13255Keywords:
Split Inference, On-Device AI, Privacy-Preserving Inference, Vision-Language Models, Mobile Computing, Multimodal AIAbstract
Multimodal mobile AI applications increasingly rely on cloud-based large language models (LLMs) for complex reasoning over visual inputs, but raw-image cloud upload creates substantial privacy exposure, high inference costs, and unacceptable latency for interactive use cases. This paper proposes PerceptionGuard, a four-layer split-inference architecture where on-device vision-language models (VLMs) handle privacy-sensitive perception and adaptive routing, sending only compact, privacy-preserving representations to cloud LLMs for higher-order reasoning. Three representation modes are defined and evaluated: dense embeddings (Mode A), structured scene graphs (Mode B), and redacted natural-language captions (Mode C). The architecture incorporates information-bottleneck filtering and calibrated differential-privacy noise to resist membership inference and embedding-inversion attacks. Experimental evaluation across three representative mobile workloads, accessibility visual question answering on VizWiz, augmented reality scene understanding, and visual document search on DocVQA, demonstrates that Mode B achieves task accuracy within approximately 5 to 7 percentage points of cloud-only baselines while reducing cloud token cost by over 60 percent and achieving meaningful reductions in membership-inference attack success. A learned adaptive router outperforms confidence-threshold cascade baselines on cost-accuracy Pareto frontiers. PerceptionGuard is implemented as open-source Android and iOS libraries and contributes design heuristics for practitioners building privacy-respecting multimodal mobile applications at scale.
Downloads
References
[1] D. Yu, "Cost-efficient multimodal LLM inference via cross-tier GPU heterogeneity," arXiv preprint arXiv:2603.12707, 2026. Available: https://arxiv.org/pdf/2603.12707
[2] W. Oliveira, "Less is more: engineering challenges of on-device small language model integration in a mobile application," arXiv preprint arXiv:2604.24636, 2026. Available: https://arxiv.org/pdf/2604.24636
[3] L. Cai, Y. Zhang, R. Zhang, Y. Liu, T. Jiang, D. Niyato, W. Ni & A. Jamalipour, "Federated agentic AI for wireless networks: fundamentals, approaches, and applications," arXiv preprint arXiv:2603.01755, 2026. Available: https://arxiv.org/pdf/2603.01755
[4] D. Amebley & S. Dibbo, "Are neuro-inspired multi-modal vision-language models resilient to membership inference privacy leakage?" arXiv preprint arXiv:2511.20710, 2025. Available: https://arxiv.org/pdf/2511.20710
[5] I. C. Ngong, Z. Reza & J. P. Near, "Differentially private multimodal in-context learning," arXiv preprint arXiv:2603.04894, 2026. Available: https://arxiv.org/pdf/2603.04894
[6] V. Chandra & R. Krishnamoorthi, "On-device LLMs: state of the union, 2026," Industry Report, 2026. Available: https://v-chandra.github.io/on-device-llms/
[7] T. M. Pham, P. T. Nguyen, S. Yoon, V. D. Lai, F. Dernoncourt & T. Bui, "SlimLM: an efficient small language model for on-device document assistance," in Proc. 63rd Annual Meeting of the Association for Computational Linguistics (Volume 3: System Demonstrations), pp. 436–447, 2025. Available: https://aclanthology.org/2025.acl-demo.42.pdf
[8] Y. Tussa, A. Heredia & N. Roy, "Lessons learned from developing a privacy-preserving multimodal wearable for local voice-and-vision inference," arXiv preprint arXiv:2511.11811, 2025. Available: https://arxiv.org/pdf/2511.11811
[9] Z. Liu, C. Zhao, F. Iandola, C. Lai, Y. Tian, I. Fedorov, Y. Xiong, E. Chang, Y. Shi, R. Krishnamoorthi, L. Lai & V. Chandra, "MobileLLM: optimizing sub-billion parameter language models for on-device use cases," in Proc. Forty-First International Conference on Machine Learning, 2024. Available: https://openreview.net/pdf?id=EIGbXbxcUQ
[10] D. Gurari, Q. Li, A. J. Stangl, A. Guo, C. Lin, K. Grauman, J. Luo & J. P. Bigham, "VizWiz grand challenge: answering visual questions from blind people," in Proc. IEEE CVPR, pp. 3608–3617, 2018. Available: https://openaccess.thecvf.com/content_cvpr_2018/papers/Gurari_VizWiz_Grand_Challenge_CVPR_2018_paper.pdf
[11] M. Mathew, D. Karatzas & C.V. Jawahar, "DocVQA: a dataset for VQA on document images," in Proc. IEEE/CVF WACV, pp. 2200–2209, 2021. Available: https://openaccess.thecvf.com/content/WACV2021/papers/Mathew_DocVQA_A_Dataset_for_VQA_on_Document_Images_WACV_2021_paper.pdf
[12] X. Zhang, R. Razavi-Far, H. Isah, A. David, G. Higgins & M. Zhang, "A survey on deep learning in edge-cloud collaboration: model partitioning, privacy preservation, and prospects," Knowledge-Based Systems, vol. 310, p. 112965, 2025. Available: https://www.sciencedirect.com/science/article/abs/pii/S0950705125000139
[13] P. K. A. Vasu, F. Faghri, C. L. Li, C. Koc, N. True, A. Antony, G. Santhanam, J. Gabriel, P. Grasch, O. Tuzel & H. Pouransari, "FastVLM: efficient vision encoding for vision language models," in Proc. CVPR, 2025. Available: https://machinelearning.apple.com/research/fastvlm-efficient-vision-encoding
[14] M. Huh, F. Xu, Y. H. Peng, C. Chen, D. Gurari, E. Choi & A. Pavel, "Long-form answers to visual questions from blind and low vision people," in Workshop on Demographic Diversity in Computer Vision, CVPR, 2025. Available: https://openreview.net/pdf?id=92BJhiZjWa
[15] A. Alhindi, S. Al-Ahmadi & M. M. B. Ismail , "Advancements and challenges in privacy-preserving split learning: experimental findings and future directions," International Journal of Information Security, vol. 24, no. 3, p. 125, 2025. Available: https://link.springer.com/article/10.1007/s10207-025-01045-9
[16] Y. Wang, G. Zhong, Y. Duan, Y. Cheng, M. Yin, & R. Yang, "Efficient and privacy-preserving deep inference towards cloud-edge collaborative," Applied Soft Computing, vol. 180, p. 113381, 2025. Available: https://www.sciencedirect.com/science/article/abs/pii/S1568494625006921
[17] Y. Shu, S. Li, T. Dong, Y. Meng & H. Zhu, "Model inversion in split learning for personalized LLMs: new insights from information bottleneck theory," arXiv preprint arXiv:2501.05965, 2025. Available: https://arxiv.org/pdf/2501.05965
[18] Y. Dong, W. Luo, X. Wang, L. Zhang, L. Xu, Z. Zhou & L. Wang, "Multi-task federated split learning across multi-modal data with privacy preservation," Sensors, vol. 25, no. 1, p. 233, 2025. Available: https://www.mdpi.com/1424-8220/25/1/233
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Sridhar Muthineni

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
Authors who publish with this journal agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License (Attribution-ShareAlike 4.0 International (CC BY-SA 4.0) ) that allows others to share the work with an acknowledgement of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgement of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work (See The Effect of Open Access).








