EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning
作者:Shuoqin Zhang, Tongtong Cheng, Xiru Gao, Jinzhuo Peng, Bin Zheng, Jiahao Tu, Ke Wang, Jia Pan, Zhe Hu, Kai Liu · 单位:are with the National Elite Institute of Engineering, Chongqing University, Chongqing 401135, China.Tongtong Cheng and Kai Liu are with the College of Computer Science, Chongqing University, Chongqing 400044, China.Jinzhuo Peng is with the College of Mechanical and Vehicle Engineering, Chongqing University, Chongqing 400044, China.Jia Pan is with the Department of Computer Science, The University of, Hong Kong, Hong Kong, China.This work was conducted in collaboration with industry partner Chengdu · 会议/期刊:arXiv preprint · 方向:cs.RO · 发布日期:2026-08-05