Learning Gaze Synthesizer via 3D-eye Controlled Diffusion and Cross-domain Feature Alignment.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 42391078.
- Also identified by DOI 10.1109/TIP.2026.3707758.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Appearance-based supervised methods with full-face image input have made tremendous advances in recent gaze estimation tasks. However, intensive human annotation requirement inhibits current methods from achieving industrial level accuracy and robustness. Although methods based on generative AI can synthesize eye images to expand self-annotated eye data, these methods usually have limited model capacity and cannot effectively inject gaze information, resulting in poor quality, monotonous texture, and inaccurate gaze direction of the generated eye images. To alleviate the above challenge, we propose a novel gaze data synthesizer framework, in which a 3D-eye model that can flexibly manipulate the gaze direction is used to finely control eye image synthesis based on a stable diffusion large generative model, so that high-quality eye images with arbitrary gaze angle can be synthesized. At the same time, when training the gaze feature extractor, we propose a cross-domain feature alignment module to minimize the feature distribution discrepancy between real samples and synthetic ones, to pursue domain-invariant (shape, texture, etc.) gaze representation. Both qualitative and quantitative experimental results demonstrate that the proposed scheme generates high-quality gaze images and also achieves superior gaze estimation performances over state-of-the-art.