PerceptionGuard: Privacy-Aware Split Inference for Multimodal Mobile Applications

Authors

  • Sridhar Muthineni Optum Services Inc., USA

DOI:

https://doi.org/10.30871/jaic.v10i4.13255

Keywords:

Split Inference, On-Device AI, Privacy-Preserving Inference, Vision-Language Models, Mobile Computing, Multimodal AI

Abstract

Multimodal mobile AI applications increasingly rely on cloud-based large language models (LLMs) for complex reasoning over visual inputs, but raw-image cloud upload creates substantial privacy exposure, high inference costs, and unacceptable latency for interactive use cases. This paper proposes PerceptionGuard, a four-layer split-inference architecture where on-device vision-language models (VLMs) handle privacy-sensitive perception and adaptive routing, sending only compact, privacy-preserving representations to cloud LLMs for higher-order reasoning. Three representation modes are defined and evaluated: dense embeddings (Mode A), structured scene graphs (Mode B), and redacted natural-language captions (Mode C). The architecture incorporates information-bottleneck filtering and calibrated differential-privacy noise to resist membership inference and embedding-inversion attacks. Experimental evaluation across three representative mobile workloads, accessibility visual question answering on VizWiz, augmented reality scene understanding, and visual document search on DocVQA, demonstrates that Mode B achieves task accuracy within approximately 5 to 7 percentage points of cloud-only baselines while reducing cloud token cost by over 60 percent and achieving meaningful reductions in membership-inference attack success. A learned adaptive router outperforms confidence-threshold cascade baselines on cost-accuracy Pareto frontiers. PerceptionGuard is implemented as open-source Android and iOS libraries and contributes design heuristics for practitioners building privacy-respecting multimodal mobile applications at scale.

Downloads

Download data is not yet available.

References

[1] D. Yu, "Cost-efficient multimodal LLM inference via cross-tier GPU heterogeneity," arXiv preprint arXiv:2603.12707, 2026. Available: https://arxiv.org/pdf/2603.12707

[2] W. Oliveira, "Less is more: engineering challenges of on-device small language model integration in a mobile application," arXiv preprint arXiv:2604.24636, 2026. Available: https://arxiv.org/pdf/2604.24636

[3] L. Cai, Y. Zhang, R. Zhang, Y. Liu, T. Jiang, D. Niyato, W. Ni & A. Jamalipour, "Federated agentic AI for wireless networks: fundamentals, approaches, and applications," arXiv preprint arXiv:2603.01755, 2026. Available: https://arxiv.org/pdf/2603.01755

[4] D. Amebley & S. Dibbo, "Are neuro-inspired multi-modal vision-language models resilient to membership inference privacy leakage?" arXiv preprint arXiv:2511.20710, 2025. Available: https://arxiv.org/pdf/2511.20710

[5] I. C. Ngong, Z. Reza & J. P. Near, "Differentially private multimodal in-context learning," arXiv preprint arXiv:2603.04894, 2026. Available: https://arxiv.org/pdf/2603.04894

[6] V. Chandra & R. Krishnamoorthi, "On-device LLMs: state of the union, 2026," Industry Report, 2026. Available: https://v-chandra.github.io/on-device-llms/

[7] T. M. Pham, P. T. Nguyen, S. Yoon, V. D. Lai, F. Dernoncourt & T. Bui, "SlimLM: an efficient small language model for on-device document assistance," in Proc. 63rd Annual Meeting of the Association for Computational Linguistics (Volume 3: System Demonstrations), pp. 436–447, 2025. Available: https://aclanthology.org/2025.acl-demo.42.pdf

[8] Y. Tussa, A. Heredia & N. Roy, "Lessons learned from developing a privacy-preserving multimodal wearable for local voice-and-vision inference," arXiv preprint arXiv:2511.11811, 2025. Available: https://arxiv.org/pdf/2511.11811

[9] Z. Liu, C. Zhao, F. Iandola, C. Lai, Y. Tian, I. Fedorov, Y. Xiong, E. Chang, Y. Shi, R. Krishnamoorthi, L. Lai & V. Chandra, "MobileLLM: optimizing sub-billion parameter language models for on-device use cases," in Proc. Forty-First International Conference on Machine Learning, 2024. Available: https://openreview.net/pdf?id=EIGbXbxcUQ

[10] D. Gurari, Q. Li, A. J. Stangl, A. Guo, C. Lin, K. Grauman, J. Luo & J. P. Bigham, "VizWiz grand challenge: answering visual questions from blind people," in Proc. IEEE CVPR, pp. 3608–3617, 2018. Available: https://openaccess.thecvf.com/content_cvpr_2018/papers/Gurari_VizWiz_Grand_Challenge_CVPR_2018_paper.pdf

[11] M. Mathew, D. Karatzas & C.V. Jawahar, "DocVQA: a dataset for VQA on document images," in Proc. IEEE/CVF WACV, pp. 2200–2209, 2021. Available: https://openaccess.thecvf.com/content/WACV2021/papers/Mathew_DocVQA_A_Dataset_for_VQA_on_Document_Images_WACV_2021_paper.pdf

[12] X. Zhang, R. Razavi-Far, H. Isah, A. David, G. Higgins & M. Zhang, "A survey on deep learning in edge-cloud collaboration: model partitioning, privacy preservation, and prospects," Knowledge-Based Systems, vol. 310, p. 112965, 2025. Available: https://www.sciencedirect.com/science/article/abs/pii/S0950705125000139

[13] P. K. A. Vasu, F. Faghri, C. L. Li, C. Koc, N. True, A. Antony, G. Santhanam, J. Gabriel, P. Grasch, O. Tuzel & H. Pouransari, "FastVLM: efficient vision encoding for vision language models," in Proc. CVPR, 2025. Available: https://machinelearning.apple.com/research/fastvlm-efficient-vision-encoding

[14] M. Huh, F. Xu, Y. H. Peng, C. Chen, D. Gurari, E. Choi & A. Pavel, "Long-form answers to visual questions from blind and low vision people," in Workshop on Demographic Diversity in Computer Vision, CVPR, 2025. Available: https://openreview.net/pdf?id=92BJhiZjWa

[15] A. Alhindi, S. Al-Ahmadi & M. M. B. Ismail , "Advancements and challenges in privacy-preserving split learning: experimental findings and future directions," International Journal of Information Security, vol. 24, no. 3, p. 125, 2025. Available: https://link.springer.com/article/10.1007/s10207-025-01045-9

[16] Y. Wang, G. Zhong, Y. Duan, Y. Cheng, M. Yin, & R. Yang, "Efficient and privacy-preserving deep inference towards cloud-edge collaborative," Applied Soft Computing, vol. 180, p. 113381, 2025. Available: https://www.sciencedirect.com/science/article/abs/pii/S1568494625006921

[17] Y. Shu, S. Li, T. Dong, Y. Meng & H. Zhu, "Model inversion in split learning for personalized LLMs: new insights from information bottleneck theory," arXiv preprint arXiv:2501.05965, 2025. Available: https://arxiv.org/pdf/2501.05965

[18] Y. Dong, W. Luo, X. Wang, L. Zhang, L. Xu, Z. Zhou & L. Wang, "Multi-task federated split learning across multi-modal data with privacy preservation," Sensors, vol. 25, no. 1, p. 233, 2025. Available: https://www.mdpi.com/1424-8220/25/1/233

Downloads

Published

2026-08-07

How to Cite

[1]
S. Muthineni, “PerceptionGuard: Privacy-Aware Split Inference for Multimodal Mobile Applications”, JAIC, vol. 10, no. 4, pp. 3166–3176, Aug. 2026.

Similar Articles

<< < 46 47 48 

You may also start an advanced similarity search for this article.