Privacy-Preserving Heterogeneous Multi-LLM Federated Inference for Cognitive Diagnosis
cs.CR, cs.AI, cs.CL, cs.LG
Submitted: 2026-09-01
Updated: 2026-09-01
Comments: 16 pages, 3 figures, Accepted to Findings of EMNLP 2026
Code: https://github.com/manasa2107/priv
License: http://creativecommons.org/licenses/by/4.0/
The gist: Significant challenges remain in AI-driven educational systems in balancing privacy preservation with accurate cognitive diagnosis.
Terminology
Abstract
Significant challenges remain in AI-driven educational systems in balancing privacy preservation with accurate cognitive diagnosis. To overcome this, we propose a federated inference framework in which several commercial LLM APIs collaborate without requiring access to raw student data or proprietary model internals. Using multiple federated entities, such as LLaMA-3.3-70B, GPT-4o-mini, and Claude-3-Haiku, our framework builds upon a heterogeneous multi-LLM architecture. The predictions generated by these entities are combined with epsilon-local differential privacy by adding Laplace noise locally to each entity's prediction output before aggregation, while residual-based aggregation mitigates model heterogeneity. Our approach is predicated on an honest-but-curious trust paradigm in which API providers are presumed not to abuse submitted queries, and our differential privacy mechanism shields the published diagnostic results from external inference. We conduct rigorous privacy-utility analysis showing strong privacy guarantees with minimal accuracy loss, and extensive real-world evaluations across three educational benchmarks confirm the framework's practical usability and cross-domain generalizability.
Sources
- GPT-4 Technical Report
- Sparks of Artificial General Intelligence: Early experiments with GPT-4
- Training Verifiers to Solve Math Word Problems
- Differentially Private Federated Learning: A Client Level Perspective
- LLaMA: Open and Efficient Foundation Language Models
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs