Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework figure
AlphaXiv 中文概览(可滚动查看)