Computational Efficiency through Momentum-Guided Split Distillation: Rethinking Federated Learning for Edge Autonomy
Conférence : Communications par affiche dans un congrès international ou national
IoT ecosystems increasingly require adaptive intelligence under heterogeneous data and resource conditions [1]. In such distributed, privacy-sensitive, and evolving environments, Federated Learning (FL) enables collaborative edge intelligence without transmitting raw data [2]. However, IoT intelligence requires temporal reasoning to capture evolving and context-dependent dynamics [3]. This makes temporal reasoning valuable yet costly for constrained devices. FL often assumes clients can train and synchronize full models [4]. Split and federated split learning reduce this burden by moving part of the model to a server [5]. Yet they maintain server-side dependency, creating a training–deployment mismatch with autonomous edge inference. This is amplified by non-I.I.D. data, heterogeneous sensing modalities, and divergent temporal learning directions [6], where a single global model with uniform updates cannot support personalized temporal behavior. We investigate computational efficiency through a momentum-guided federated split distillation framework for temporal edge intelligence. We adopt a U-shaped split design in which clients retain a lightweight deployable model comprising a reservoir representation module, compact temporal student, and personalized output module, while the server hosts a higher-capacity temporal teacher used only during training. The teacher adapts to client needs and provides refined activations for distillation, transferring server-side temporal knowledge into lightweight edge models (Fig. 1). The novelty is momentum-guided teacher fusion for personalized split distillation. By tracking client-induced teacher displacements, the server derives path-aware momentum vectors that capture persistent temporal learning trajectories. Clients with aligned trajectories are grouped to receive cluster-specialized teacher updates, making collaboration trajectorycompatible rather than global (Fig. 2). We improve edge efficiency by distilling server-side temporal knowledge into lightweight local students, thereby removing server-side inference dependency, while trajectory-aware teacher specialization yields more personalized and effective updates that reduce learning inefficiencies in heterogeneous temporal IoT ecosystems.