TS-Binaural: Visually Guided Binaural Audio Generation by Temporal-Spatial Dynamic Analysis.

Liu, Shulin; Cheng, Haonan; Lian, Zhicheng; Ye, Long; Zhang, Qin · IEEE Trans Neural Netw Learn Syst · 2026

basic_science · Level V

Where this comes from

Abstract

Visually guided binaural audio generation (VGBAG) aims at estimating spatial information from visual images to reconstruct binaural audio, yet struggles to model dynamic audio-visual relationships in complex multisource scenarios. This work focuses on musical instrument performance videos, a critical case characterized by a variable number of sound sources, multiple performance forms, and diverse stage effects. These dynamic characteristics are ignored by existing methods, resulting in inaccuracy of the generated binaural audio. To overcome these limitations, in this article, we propose a novel VGBAG method (TS-Binaural) based on temporal-spatial dynamic analysis (TSDA). Different from the overall audio-visual content analysis in previous work, the TS-Binaural employs a novel "separation-later-mixing" strategy. It splits complex audio-visual scenes into separate units to improve accuracy. The TS-Binaural comprises two key components: a TSDA module and a temporal-spatial information fusion (TSIF) guided binaural audio generation module. Furthermore, in the TSDA module, we design an automatic mask generation-based performance analyzer (AMGPA) to prevent intersource interference. The AMGPA integrates novel position detection logic to identify and separate the activation region, which refers to a separate visual region consisting of a player and his/her musical instrument being played at different times in the video. Extensive quantitative and qualitative comparisons demonstrate that our TS-Binaural outperforms the state-of-the-art VGBAG methods.