Compositional Cross-Modality Translation via Whole-Volume Multitask Latent Flow Matching figure
AlphaXiv 中文概览(可滚动查看)