Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction
AuthorsKaisi Guan†‡**, Xihua Wang†‡, Zhengfeng Lai, Xin Cheng†, Peng Zhang, Xiaojiang Liu, Ruihua Song†, Meng Cao
Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction
AuthorsKaisi Guan†‡**, Xihua Wang†‡, Zhengfeng Lai, Xin Cheng†, Peng Zhang, Xiaojiang Liu, Ruihua Song†, Meng Cao
This study focuses on Text-to-Sounding-Video (T2SV) generation, which aims to generate a video with synchronized audio from text, with both modalities aligned to the text conditions. Despite progress in joint audio-video training, two critical challenges remain: (1) text conditioning is a bottleneck—shared captions (TV=TA) trigger modal interference, while a gap persists between dense training captions and concise inference user prompts, and (2) the optimal fusion mechanism for cross-modal feature interaction remains unclear. To address the first challenge, we first propose the Cross-Referential Rewriter (CRR) caption framework, a dual-agent pipeline where a Semantic Checker extracts grounded Semantic Anchors and a Cross-Modal Rewriter generates disentangled caption pairs (TV and TA), eliminating modal interference and bridging the training-inference gap.
Putting Captions to the Test: Evaluating Video Caption Quality through Multiple-Choice Question Answering
September 11, 2026research area Computer Vision, research area Data Science and Annotationconference ACL
Evaluating video captioning remains a critical challenge for Visual Large Language Models (VLLMs). Existing metrics primarily rely on matching generated text against ground-truth references. This paradigm suffers from the “one-to-many” nature of video description, where high-quality captions are often penalized for lexical mismatches or valid shifts in visual focus. Furthermore, such assessments are typically one-dimensional, failing to provide a…
Promoting Cross-Modal Representations to Improve Multimodal Foundation Models for Physiological Signals
October 28, 2024research area Methods and Algorithmsconference NeurIPS
Many healthcare applications are inherently multimodal, involving several physiological signals. As sensors for these signals become more common, improving machine learning methods for multimodal healthcare data is crucial. Pretraining foundation models is a promising avenue for success. However, methods for developing foundation models in healthcare are still in early exploration and it is unclear which pretraining strategies are most effective…