Reverse to Advance: Teleoperation-Cost Effective Hard Policy Learning from Reversed Easy Tasks figure
AlphaXiv 中文概览(可滚动查看)