Watching Synthetic Videos: Aligning Cross-modal Representations with Visual Synthesis for Zero-shot Video Captioning figure
AlphaXiv 中文概览(可滚动查看)