MoFTSS: Motion Generation With Frequency and Text State Space Models.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 42065978.
- Also identified by DOI 10.1109/TNNLS.2026.3683909.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Text-driven diffusion models have achieved remarkable performance in human motion generation. However, these generative works struggle to generate high-quality motion consistent with textual descriptions. The primary reasons are: 1) insufficient fine-grained motion modeling due to the motion representations being difficult to distinguish in latent diffusion; and 2) inconsistencies between motions and textual descriptions due to misalignment in the multimodal space. To overcome these limitations, this work proposes the Motion generation with Frequency and Text State Space models (MoFTSS) including two main modules: frequency state space model (FreqSSM) and text state space model (TextSSM). Specifically, FreqSSM derives fine-grained representations by decomposing sequences into low-frequency and high-frequency components. This allows it to guide the generation of static poses (e.g., sitting, lying) and fine-grained motions (e.g, transitions, stumbling). For consistency between text and motion, TextSSM treats text features as a semantic modulation term within the SSM, enabling dynamic filtering of motion features consistent with textual semantics. Extensive experiments suggest that our MoFTSS achieves superior performance on the text-to-motion generation task. Notably, it attains the lowest FID of 0.181 on the HumanML3D dataset, significantly lower than the 0.421 achieved by MLD.