Multimodal Action Recognition via Causality-Inspired Graph Representation Learning.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 42678833.
- Also identified by DOI 10.1109/TIP.2026.3727908.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Multimodal human activity recognition (HAR) benefits from complementary skeleton, inertial, and visual observations. However, many learning-based models still treat relationships among modalities, joints, and sensor variables as symmetric associations. This limits their ability to represent asymmetric information flow and can weaken the preservation of modality-specific cues during feature fusion. We propose Causality-Inspired Structure Representation Learning (CSRL), a multimodal HAR framework that uses directional dependency modeling as a structural prior for representation learning. CSRL first estimates transfer-entropy-based graphs from temporal entities, including skeleton joints and IMU sensor variables. These graphs provide asymmetric priors that guide recognition-oriented graph learning in the representation space. CSRL further combines hybrid contrastive learning with an encoder-decoder architecture to learn modality-invariant, modality-specific, and structure-aware representations in a unified framework. This design encourages cross-modal alignment while retaining local motion cues that are important for fine-grained action discrimination. Experiments on five public HAR benchmarks, including UTD-MHAD, MMAct, CZU-MHAD, NTU RGB+D, and NTU RGB+D 120, show that CSRL consistently improves accuracy, F1 score, and recall over competitive supervised and contrastive baselines. These results support TE-guided directional structure modeling as a practical and interpretable prior for multimodal action recognition.