AnimateAnyMesh++: A Flexible Feed-Forward Framework for High-Fidelity Text-Driven Mesh Animation.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 42606975.
- Also identified by DOI 10.1109/TPAMI.2026.3724593.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Recent advances in 4D content generation have attracted increasing attention, yet creating high-quality animated 3D models remains challenging due to the complexity of modeling spatio-temporal distributions and the scarcity of 4D training data. We present AnimateAnyMesh++, a feed-forward framework for text-driven animation of arbitrary 3D meshes with substantial upgrades in data, architecture, and generative capability. First, we expand the DyMesh-XL dataset by mining dynamic content from Objaverse-XL, increasing the number of unique identities from 60K to 300K and substantially broadening category and motion diversity. Second, we redesign DyMeshVAE-Flex with power-law topology-aware attention and vertex-normal-enhanced features, which significantly improves trajectory reconstruction, local geometry preservation, and mit igates trajectory-sticking artifacts. Third, we introduce archi tectural changes to both DyMeshVAE-Flex and the rectified flow (RF) generator to support variable-length sequence training and generation, enabling longer animations while preserving reconstruction fidelity. Extensive experiments demonstrate that AnimateAnyMesh++ generates semantically accurate and tem porally coherent mesh animations within seconds, surpassing prior approaches in quality and efficiency. The enlarged DyMesh XL, the upgraded DyMeshVAE-Flex, and variable-length RF to gether deliver consistent gains across benchmarks and in-the-wild meshes. We will release code, models, and the expanded DyMesh XL at https://github.com/JarrentWu1031/AnimateAnyMesh-pp upon acceptance of this manuscript to facilitate research in 4D content creation.