DeepJSCC for video semantic communication with general semantic preservation.

Li, Junting; Chen, Xuechen; Deng, Xiaoheng · Neural Netw · 2026

basic_science · Level V

Where this comes from

Abstract

Video-based intelligent applications such as autonomous driving, video understanding, and telemedicine rely heavily on the availability of robust and transferable semantic representations. In practical systems, however, video signals transmitted over wireless links are often exposed to severe and unpredictable distortions caused by bandwidth limitations, time-varying noise, and transmission impairments. Under such conditions, high reconstruction quality does not necessarily guarantee the preservation of semantic structures that support downstream reasoning, leading to degraded temporal coherence and poor generalization to unseen tasks. Existing learning-based video transmission and compression methods predominantly optimize pixel-level reconstruction fidelity, while the preservation of general, task-agnostic semantic representations remains insufficiently explored. To address this challenge, we propose a novel end-to-end deep joint source-channel coding (DeepJSCC) framework for video semantic transmission, which explicitly perserves general semantic information for diverse downstream video tasks over wireless channels. Our framework firstly employs a Video-level Semantic Alignment Module (VSAM) to reduce semantic distance between original and reconstructed videos by video-level contrastive learning objectives. Additionally, we propose the Temporal Consistency Learning Module (TCLM) to enhance temporal coherence by predicting future-frame semantic embeddings, enabling more stable temporal dynamics in reconstructed videos. The experiments demonstrate that the proposed framework achieves reconstruction quality comparable to existing DeepJSCC-based and traditional transmission schemes, while delivering consistent performance gains across multiple video downstream tasks without any task-specific fine-tuning.