Fetching the paper…
Reading the bibliography…
Reward function is essential in reinforcement learning (RL), serving as the guiding signal to incentivize agents to solve given tasks, however, is also notoriously difficult to design.
Algaedice: Policy gradient from arbitrary experience
Ofir Nachum, Bo Dai, Ilya Kostrikov, Yinlam Chow, Lihong Li, and Dale Schuurmans · 1912
Earlier work this paper cites.
Alvinn: An autonomous land vehicle in a neural network
Dean A Pomerleau · 1988
Earlier work this paper cites.
Robot shaping: Developing autonomous agents through learning
Marco Dorigo and Marco Colombetti · 1994
Earlier work this paper cites.
Stochastic approximation with two time scales
Vivek S Borkar · 1997
Earlier work this paper cites.
Learning to drive a bicycle using reinforcement learning and shaping
Jette Randløv and Preben Alstrøm · 1998
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Andrew Y Ng, Daishi Harada, and Stuart Russell · 1999
Earlier work this paper cites.
Convex optimization
Stephen Boyd, Stephen P Boyd, and Lieven Vandenberghe · 2004
Earlier work this paper cites.
Intrinsically motivated reinforcement learning
Nuttapong Chentanez, Andrew Barto, and Satinder Singh · 2004
Earlier work this paper cites.
Automatic shaping and decomposition of reward functions
Bhaskara Marthi · 2007
Earlier work this paper cites.
Variational analysis , volume 317
R Tyrrell Rockafellar and Roger J-B Wets · 2009
Earlier work this paper cites.
Where do rewards come from
Satinder Singh, Richard L Lewis, and Andrew G Barto · 2009
Earlier work this paper cites.
Reward design via online gradient ascent
Jonathan Sorg, Richard L Lewis, and Satinder Singh · 2010
Earlier work this paper cites.
The optimal reward problem: Designing effective reward for bounded agents
Jonathan Daniel Sorg · 2011
Earlier work this paper cites.
Monte Carlo theory, methods and examples
Art B. Owen · 2013
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
Martin L Puterman · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Learning from conditional distributions via dual embeddings
Bo Dai, Niao He, Yunpeng Pan, Byron Boots, and Le Song · 2017
Earlier work this paper cites.
Inverse reward design
Dylan Hadfield-Menell, Smitha Milli, Pieter Abbeel, Stuart J Russell, and Anca Dragan · 2017
Earlier work this paper cites.
Addressing function approximation error in actor-critic methods
Scott Fujimoto, Herke Hoof, and David Meger · 2018
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Earlier work this paper cites.
Belief reward shaping in reinforcement learning
Ofir Marom and Benjamin Rosman · 2018
Cited alongside, same era.
Yuval Tassa, Yotam Doron, Alistair Muldal, Tom Erez, Yazhe Li, Diego de Las Casas, David Budden, Abbas Abdolmaleki, Josh Merel, Andrew Lefrancq, et al · 2018
Cited alongside, same era.
A low-cost ethics shaping approach for designing reinforcement learning agents
Yueh-Hua Wu and Shou-De Lin · 2018
Cited alongside, same era.
On learning intrinsic rewards for policy gradient methods
Zeyu Zheng, Junhyuk Oh, and Satinder Singh · 2018
Cited alongside, same era.
Off-policy deep reinforcement learning without exploration
Scott Fujimoto, David Meger, and Doina Precup · 2019
Cited alongside, same era.
Human-level performance in 3d multiplayer games with population-based reinforcement learning
Explicable reward design for reinforcement learning agents
Rati Devidze, Goran Radanovic, Parameswaran Kamalaruban, and Adish Singla · 2021
Later among the works it cites.
A minimalist approach to offline reinforcement learning
Scott Fujimoto and Shixiang Shane Gu · 2021
Later among the works it cites.
Offline inverse reinforcement learning
Firas Jarboui and Vianney Perchet · 2021
Later among the works it cites.
Demodice: Offline imitation learning with supplementary imperfect demonstrations
Geon-Hyeong Kim, Seokin Seo, Jongmin Lee, Wonseok Jeon, HyeongJoo Hwang, Hongseok Yang, and Kee-Eung Kim · 2021
Later among the works it cites.
Optidice: Offline policy optimization via stationary distribution correction estimation
Jongmin Lee, Wonseok Jeon, Byungjun Lee, Joelle Pineau, and Kee-Eung Kim · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Max Jaderberg, Wojciech M Czarnecki, Iain Dunning, Luke Marris, Guy Lever, Antonio Garcia Castaneda, Charles Beattie, Neil C Rabinowitz, Ari S Morcos, Avraham Ruderman, et al · 2019
Cited alongside, same era.
Imitation learning via off-policy distribution matching
Ilya Kostrikov, Ofir Nachum, and Jonathan Tompson · 2019
Cited alongside, same era.
Stabilizing off-policy q-learning via bootstrapping error reduction
Aviral Kumar, Justin Fu, Matthew Soh, George Tucker, and Sergey Levine · 2019
Cited alongside, same era.
A formal methods approach to interpretable reinforcement learning for robotic planning
Xiao Li, Zachary Serlin, Guang Yang, and Calin Belta · 2019
Cited alongside, same era.
D4rl: Datasets for deep data-driven reinforcement learning
Justin Fu, Aviral Kumar, Ofir Nachum, George Tucker, and Sergey Levine · 2020
Cited alongside, same era.
Generative adversarial networks
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2020
Cited alongside, same era.
Mingyi Hong, Hoi-To Wai, Zhaoran Wang, and Zhuoran Yang · 2020
Cited alongside, same era.
What matters in learning from offline human demonstrations for robot manipulation
Ajay Mandlekar, Danfei Xu, Josiah Wong, Soroush Nasiriany, Chen Wang, Rohun Kulkarni, Li Fei-Fei, Silvio Savarese, Yuke Zhu, and Roberto Martín-Martín · 2021
Later among the works it cites.
How to train your energy-based models
Yang Song and Diederik P Kingma · 2021
Later among the works it cites.
Offline reinforcement learning with soft behavior regularization
Haoran Xu, Xianyuan Zhan, Jianxiong Li, and Honglei Yin · 2021
Later among the works it cites.
Conservative data sharing for multi-task offline reinforcement learning
Tianhe Yu, Aviral Kumar, Yevgen Chebotar, Karol Hausman, Sergey Levine, and Chelsea Finn · 2021
Later among the works it cites.
Adversarially trained actor critic for offline reinforcement learning
Ching-An Cheng, Tengyang Xie, Nan Jiang, and Alekh Agarwal · 2022
Later among the works it cites.
Implicit behavioral cloning
Pete Florence, Corey Lynch, Andy Zeng, Oscar A Ramirez, Ayzaan Wahid, Laura Downs, Adrian Wong, Johnny Lee, Igor Mordatch, and Jonathan Tompson · 2022
Later among the works it cites.
Versatile offline imitation from observations and examples via regularized state-occupancy matching
Yecheng Ma, Andrew Shen, Dinesh Jayaraman, and Osbert Bastani · 2022
Later among the works it cites.
Exploiting reward shifting in value-based deep rl
Hao Sun, Lei Han, Rui Yang, Xiaoteng Ma, Jian Guo, and Bolei Zhou · 2022
Later among the works it cites.
How to leverage unlabeled data in offline reinforcement learning
Tianhe Yu, Aviral Kumar, Yevgen Chebotar, Karol Hausman, Chelsea Finn, and Sergey Levine · 2022
Later among the works it cites.
Deepthermal: Combustion optimization for thermal power generating units using offline reinforcement learning
Xianyuan Zhan, Haoran Xu, Yue Zhang, Xiangyu Zhu, Honglei Yin, and Yu Zheng · 2022
Later among the works it cites.
Discriminator-guided model-based offline imitation learning
Wenjia Zhang, Haoran Xu, Haoyi Niu, Peng Cheng, Ming Li, Heming Zhang, Guyue Zhou, and Xianyuan Zhan · 2022
Later among the works it cites.
Uncertainty-based multi-task data sharing for offline reinforcement learning, 2023
Chenjia Bai, Lingxiao Wang, Zhuoran Yang, Zhen Wang, Zhaoran Wang, Bin Zhao, and Xuelong Li · 2023
Closest in time.
When data geometry meets deep function: Generalizing offline reinforcement learning
Jianxiong Li, Xianyuan Zhan, Haoran Xu, Xiangyu Zhu, Jingjing Liu, and Ya-Qin Zhang · 2023
Closest in time.
Sparse q-learning: Offline reinforcement learning with implicit value regularization
Haoran Xu, Li Jiang, Jianxiong Li, Zhuoran Yang, Zhaoran Wang, Victor Wai Kin Chan, and Xianyuan Zhan · 2023
Closest in time.