Robust Average-Reward Markov Decision Processes: Minimax-Optimal Learning via Plug-in Reductions figure
AlphaXiv 中文概览(可滚动查看)