Toward Accurate Procedure Planning in Instructional Videos: Visual State Generation Helps Task-Selective Diffusion.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 41364568.
- Also identified by DOI 10.1109/TPAMI.2025.3641798.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Procedure planning in instructional videos entails predicting an action sequence that transitions a given start state to a desired goal state. This task is particularly challenging due to two key sources of uncertainty: limited visual observations and an enormous decision space. The former results in multiple plausible plan variations due to missing intermediate visual states, while the latter complicates prediction by requiring selection from a large set of potential actions. Unlike prior work that addresses these issues implicitly, we propose an explicit solution. To mitigate the first challenge, we employ image generation models to synthesize diverse intermediate visual states using various text prompts, followed by a prompt selection module integrated within a diffusion model. To tackle the second challenge, we introduce a task-selective diffusion model that applies a task-specific mask to constrain the action space. As the effectiveness of this mask depends on accurate task classification, we further enhance visual representation by leveraging pre-trained vision-language models to generate action-aware, text-enriched multimodal embeddings. Extensive experiments on three benchmark datasets validate the superior performance of our proposed approach.