Scanpath Prediction in Panoramic Videos Via Expected Code Length Minimization.
Where this comes from
- Record sourced from PubMed, PMID 42184184.
- Also identified by DOI 10.1109/TPAMI.2026.3696331.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Scanpath prediction in panoramic videos is a challenging task due to the spherical geometry and multimodality of the input, and the inherent uncertainty and diversity of the output. To give a complete treatment of these characteristics, we first present a simple criterion for scanpath prediction based on principles from lossy data compression. This criterion suggests minimizing the expected code length of quantized scanpaths, corresponding to fitting a discrete conditional probability model via maximum likelihood. We condition the probability model on two modalities: a viewport sequence as the deformation-reduced visual input and a set of relative past scanpaths projected onto respective viewports as the aligned path input. Furthermore, we parameterize it by a product of discretized Gaussian mixture models to capture the uncertainty and diversity of scanpaths from different humans. In doing so, the training of the probability model does not rely on the specification of "ground-truth" scanpaths for imitation learning. We also introduce a proportional-integral-derivative (PID) controller-based sampler to generate realistic human-like scanpaths from the learned probability model. Experimental results demonstrate that our method consistently produces better quantitative scanpath results in terms of prediction accuracy (by comparing to the assumed "ground-truths") and perceptual realism (through machine discrimination) over a wide range of prediction horizons. We additionally verify the perceptual realism improvement via a formal psychophysical experiment and the generalization improvement on several unseen panoramic video datasets.