f-divergence constrained policy improvement
Original
Boris Belousov and Jan Peters · 2017
Later among the works it cites.
Boosting the actor with dual critic
Original
Bo Dai, Albert Shaw, Niao He, Lihong Li, and Le Song · 2017
Later among the works it cites.
Stochastic variance reduction methods for policy evaluation
Simon S Du, Jianshu Chen, Lihong Li, Lin Xiao, and Dengyong Zhou · 2017
Later among the works it cites.
Learning robust rewards with adversarial inverse reinforcement learning
Original
Justin Fu, Katie Luo, and Sergey Levine · 2017
Later among the works it cites.
A unified view of entropy-regularized markov decision processes
Original
Gergely Neu, Anders Jonsson, and Vicenç Gómez · 2017
Later among the works it cites.
Randomized linear programming solves the discounted markov decision problem in nearly-linear (sometimes sublinear) running time
Original
Mengdi Wang · 2017
Later among the works it cites.
Maximum a posteriori policy optimisation
Original
Abbas Abdolmaleki, Jost Tobias Springenberg, Yuval Tassa, Remi Munos, Nicolas Heess, and Martin Riedmiller · 2018
Later among the works it cites.
Scalable bilinear π \pi learning using state and action features
Original
Yichen Chen, Lihong Li, and Mengdi Wang · 2018
Later among the works it cites.
Discriminator-actor-critic: Addressing sample inefficiency and reward bias in adversarial imitation learning
Original
Ilya Kostrikov, Kumar Krishna Agrawal, Debidatta Dwibedi, Sergey Levine, and Jonathan Tompson · 2018
Later among the works it cites.
Faster saddle-point optimization for solving large-scale markov decision processes
Original
Joan Bas-Serrano and Gergely Neu · 2019
Later among the works it cites.
Exponential family estimation via adversarial dynamics embedding
Original
Bo Dai, Zhen Liu, Hanjun Dai, Niao He, Arthur Gretton, Le Song, and Dale Schuurmans · 2019
Later among the works it cites.
A divergence minimization perspective on imitation learning methods
Original
Seyed Kamyar Seyed Ghasemipour, Richard Zemel, and Shixiang Gu · 2019
Later among the works it cites.
Imitation learning as f f -divergence minimization
Original
Liyiming Ke, Matt Barnes, Wen Sun, Gilwoo Lee, Sanjiban Choudhury, and Siddhartha Srinivasa · 2019
Later among the works it cites.
Imitation learning via off-policy distribution matching, 2019
Ilya Kostrikov, Ofir Nachum, and Jonathan Tompson · 2019
Later among the works it cites.
Advantage-weighted regression: Simple and scalable off-policy reinforcement learning, 2019
Xue Bin Peng, Aviral Kumar, Grace Zhang, and Sergey Levine · 2019
Later among the works it cites.
V-mpo: On-policy maximum a posteriori policy optimization for discrete and continuous control
Original
H Francis Song, Abbas Abdolmaleki, Jost Tobias Springenberg, Aidan Clark, Hubert Soyer, Jack W Rae, Seb Noury, Arun Ahuja, Siqi Liu, Dhruva Tirumala, et al · 2019
Later among the works it cites.
Doubly robust bias reduction in infinite horizon off-policy estimation
Original
Ziyang Tang, Yihao Feng, Lihong Li, Dengyong Zhou, and Qiang Liu · 2019
Later among the works it cites.
Minimax weight and q-function learning for off-policy evaluation
Original
Masatoshi Uehara and Nan Jiang · 2019
Later among the works it cites.
Conjugate and lagrangian duality for convex programs
Arthur F. Jr Veinott · 2019
Later among the works it cites.
GenDICE: Generalized offline estimation of stationary values
Ruiyi Zhang, Bo Dai, Li Lihong, and Dale Schuurmans · 2020
Closest in time.