Robust Average-Reward Markov Decision Processes: Minimax-Optimal Learning via Plug-in Reductions
作者:Yuepeng Yang, Yuxin Chen, Yuejie Chi · 单位:Department of Statistics and Data Science, Yale University.Department of Statistics and Data Science, the Wharton School, University of Pennsylvania. · 会议/期刊:arXiv preprint · 方向:cs.LG · 发布日期:2026-08-10