LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback figure
AlphaXiv 中文概览(可滚动查看)