MoFTSS: Motion Generation With Frequency and Text State Space Models.

Li, Chengjian; Shu, Xiangbo; Cui, Qiongjie; Xia, Haifeng; Yao, Yazhou; Tang, Jinhui · IEEE Trans Neural Netw Learn Syst · 2026

basic_science · Level V

Where this comes from

Abstract

Text-driven diffusion models have achieved remarkable performance in human motion generation. However, these generative works struggle to generate high-quality motion consistent with textual descriptions. The primary reasons are: 1) insufficient fine-grained motion modeling due to the motion representations being difficult to distinguish in latent diffusion; and 2) inconsistencies between motions and textual descriptions due to misalignment in the multimodal space. To overcome these limitations, this work proposes the Motion generation with Frequency and Text State Space models (MoFTSS) including two main modules: frequency state space model (FreqSSM) and text state space model (TextSSM). Specifically, FreqSSM derives fine-grained representations by decomposing sequences into low-frequency and high-frequency components. This allows it to guide the generation of static poses (e.g., sitting, lying) and fine-grained motions (e.g, transitions, stumbling). For consistency between text and motion, TextSSM treats text features as a semantic modulation term within the SSM, enabling dynamic filtering of motion features consistent with textual semantics. Extensive experiments suggest that our MoFTSS achieves superior performance on the text-to-motion generation task. Notably, it attains the lowest FID of 0.181 on the HumanML3D dataset, significantly lower than the 0.421 achieved by MLD.