Fetching the paper…
Reading the bibliography…
The design of a reward function often poses a major practical challenge to real-world applications of reinforcement learning.
A new approach to linear filtering and prediction problems
Kalman, Rudolf · 1960
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Ng, Andrew and Russell, Stuart · 2000
Earlier work this paper cites.
Artificial Intelligence: A Modern Approach
Russell, Stuart J. and Norvig, Peter · 2003
Earlier work this paper cites.
Linearly-solvable markov decision problems
Todorov, Emo · 2007
Earlier work this paper cites.
General duality between optimal control and estimation
Todorov, Emo · 2008
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Ziebart, Brian, Maas, Andrew, Bagnell, Andrew, and Dey, Anind · 2008
Earlier work this paper cites.
A survey of robot learning from demonstration
Argall, Brenna D., Chernova, Sonia, Veloso, Manuela, and Browning, Brett · 2009
Earlier work this paper cites.
Optimal control as a graphical model inference problem
Kappen, Hilbert J., Gomez, Vicenc, and Opper, Manfred · 2009
Earlier work this paper cites.
Robot trajectory optimization using approximate inference
Toussaint, Marc · 2009
Earlier work this paper cites.
Where do rewards come from?
Singh, S., Lewis, R., and Barto, A · 2010
Earlier work this paper cites.
Reward design via online gradient ascent
Sorg, Jonathan, Singh, Satinder P., and Lewis, Richard L · 2010
Cited alongside, same era.
Modeling purposeful adaptive behavior with the principle of maximum causal entropy
Ziebart, Brian · 2010
Cited alongside, same era.
On stochastic optimal control and reinforcement learning by approximate inference
Rawlik, Konrad, Toussaint, Marc, and Vijayakumar, Sethu · 2012
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, Volodymyr, Kavukcuoglu, Koray, Silver, David, Rusu, Andrei A, Veness, Joel, Bellemare, Marc G, Graves, Alex, Riedmiller, Martin, Fidjeland, Andreas K, Ostrovski, Georg, Petersen, Stig, Beattie, Charles, Sadik, Amir, Antonoglou, Ioannis, King, Helen, Kumaran, Dharshan, Wierstra, Daan, Legg, Shane, and Hassabis, Demis · 2015
Cited alongside, same era.
Trust Region Policy Optimization
Schulman, John, Levine, Sergey, Moritz, Philipp, Jordan, Michael I., and Abbeel, Pieter · 2015
Cited alongside, same era.
Supersizing self-supervision: Learning to grasp from 50k tries and 700 robot hours
Pinto, Lerrel and Gupta, Abhinav · 2016
Later among the works it cites.
Reinforcement learning with deep energy-based policies
Haarnoja, Tuomas, Tang, Haoran, Abbeel, Pieter, and Levine, Sergey · 2017
Later among the works it cites.
Inverse reward design
Hadfield-Menell, Dylan, Milli, Smitha, Abbeel, Pieter, Russell, Stuart J., and Dragan, Anca D · 2017
Later among the works it cites.
Bridging the gap between value and policy based reinforcement learning
Nachum, Ofir, Norouzi, Mohammad, Xu, Kelvin, and Schuurmans, Dale · 2017
Later among the works it cites.
Sim-to-real transfer of robotic control with dynamics randomization
Peng, Xue Bin, Andrychowicz, Marcin, Zaremba, Wojciech, and Abbeel, Pieter · 2017
Later among the works it cites.
Sim-to-real robot learning from pixels with progressive nets
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Amodei, Dario, Olah, Chris, Steinhardt, Jacob, Christiano, Paul, Schulman, John, and Mané, Dan · 2016
Cited alongside, same era.
Deep spatial autoencoders for visuomotor learning
Finn, C., Tan, X., Duan, Y., Darrell, T., Levine, S., and Abbeel, P · 2016
Cited alongside, same era.
Generative adversarial imitation learning
Ho, Jonathan and Ermon, Stefano · 2016
Cited alongside, same era.
End-to-end training of deep visuomotor policies
Levine, Sergey, Finn, Chelsea, Darrell, Trevor, and Abbeel, Pieter · 2016
Cited alongside, same era.
Combining policy gradient and q-learning
O’Donoghue, Brendan, Munos, Remi, Kavukcuoglu, Koray, and Mnih, Volodymyr · 2016
Cited alongside, same era.
Rusu, Andrei A., Vecerik, Matej, Rothörl, Thomas, Heess, Nicolas, Pascanu, Razvan, and Hadsell, Raia · 2017
Later among the works it cites.
Equivalence between policy gradients and soft q-learning
Schulman, John, Chen, Xi, and Abbeel, Pieter · 2017
Later among the works it cites.
Learning robust rewards with adversarial inverse reinforcement learning
Fu, Justin, Luo, Katie, and Levine, Sergey · 2018
Closest in time.
Reward learning from narrated demonstrations
Tung, Hsiao-Yu Fish, Harley, Adam W., Huang, Liang-Kang, and Fragkiadaki, Katerina · 2018
Closest in time.