Fetching the paper…
Reading the bibliography…
Robust real-world learning should benefit from both demonstrations and interactions with the environment.
Markov decision processes
Thie, Paul R · 1983
Earlier work this paper cites.
Q-learning
Watkins, Christopher JCH and Dayan, Peter · 1992
Earlier work this paper cites.
Reinforcement learning: An introduction , volume 1
Sutton, Richard S and Barto, Andrew G · 1998
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Ng, Andrew Y, Russell, Stuart J, et al · 2000
Earlier work this paper cites.
On the sample complexity of reinforcement learning
Kakade, Sham Machandranath et al · 2003
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Abbeel, Pieter and Ng, Andrew Y · 2004
Earlier work this paper cites.
General duality between optimal control and estimation
Todorov, Emanuel · 2008
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Ziebart, Brian D, Maas, Andrew L, Bagnell, J Andrew, and Dey, Anind K · 2008
Earlier work this paper cites.
Robot trajectory optimization using approximate inference
Toussaint, Marc · 2009
Earlier work this paper cites.
Modeling purposeful adaptive behavior with the principle of maximum causal entropy
Ziebart, Brian D · 2010
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Ross, Stéphane, Gordon, Geoffrey J, and Bagnell, Drew · 2011
Earlier work this paper cites.
Degris, Thomas, White, Martha, and Sutton, Richard S · 2012
Cited alongside, same era.
Learning from limited demonstrations
Kim, Beomjoon, massoud Farahmand, Amir, Pineau, Joelle, and Precup, Doina · 2013
Cited alongside, same era.
Boosted bellman residual minimization handling expert demonstrations
Piot, Bilal, Geist, Matthieu, and Pietquin, Olivier · 2014
Cited alongside, same era.
Direct policy iteration with demonstrations
Chemali, Jessica and Lazaric, Alessandro · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, Volodymyr, Kavukcuoglu, Koray, Silver, David, Rusu, Andrei A, Veness, Joel, Bellemare, Marc G, Graves, Alex, Riedmiller, Martin, Fidjeland, Andreas K, Ostrovski, Georg, et al · 2015
Cited alongside, same era.
Safe and efficient off-policy reinforcement learning
Munos, Rémi, Stepleton, Tom, Harutyunyan, Anna, and Bellemare, Marc · 2016
Later among the works it cites.
Loss is its own reward: Self-supervision for reinforcement learning
Shelhamer, Evan, Mahmoudieh, Parsa, Argus, Max, and Darrell, Trevor · 2016
Later among the works it cites.
Sample efficient actor-critic with experience replay
Wang, Ziyu, Bapst, Victor, Heess, Nicolas, Mnih, Volodymyr, Munos, Remi, Kavukcuoglu, Koray, and de Freitas, Nando · 2016
Later among the works it cites.
End-to-end learning of driving models from large-scale video datasets
Xu, Huazhe, Gao, Yang, Yu, Fisher, and Darrell, Trevor · 2016
Later among the works it cites.
Gradient-free policy architecture search and adaptation
Ebrahimi, Sayna, Rohrbach, Anna, and Darrell, Trevor · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bojarski, Mariusz, Del Testa, Davide, Dworakowski, Daniel, Firner, Bernhard, Flepp, Beat, Goyal, Prasoon, Jackel, Lawrence D, Monfort, Mathew, Muller, Urs, Zhang, Jiakai, et al · 2016
Cited alongside, same era.
Brockman, Greg, Cheung, Vicki, Pettersson, Ludwig, Schneider, Jonas, Schulman, John, Tang, Jie, and Zaremba, Wojciech · 2016
Cited alongside, same era.
Q-prop: Sample-efficient policy gradient with an off-policy critic
Gu, Shixiang, Lillicrap, Timothy, Ghahramani, Zoubin, Turner, Richard E, and Levine, Sergey · 2016
Cited alongside, same era.
Generative adversarial imitation learning
Ho, Jonathan and Ermon, Stefano · 2016
Cited alongside, same era.
Reinforcement learning with unsupervised auxiliary tasks
Jaderberg, Max, Mnih, Volodymyr, Czarnecki, Wojciech Marian, Schaul, Tom, Leibo, Joel Z, Silver, David, and Kavukcuoglu, Koray · 2016
Cited alongside, same era.
Reinforcement learning with deep energy-based policies
Haarnoja, Tuomas, Tang, Haoran, Abbeel, Pieter, and Levine, Sergey
Cited in the paper.
Reinforcement learning with deep energy-based policies
Haarnoja, Tuomas, Tang, Haoran, Abbeel, Pieter, and Levine, Sergey
Cited in the paper.
Later among the works it cites.
Gu, Shixiang, Lillicrap, Timothy, Ghahramani, Zoubin, Turner, Richard E, Schölkopf, Bernhard, and Levine, Sergey · 2017
Later among the works it cites.
Learning from demonstrations for real world reinforcement learning
Hester, Todd, Vecerik, Matej, Pietquin, Olivier, Lanctot, Marc, Schaul, Tom, Piot, Bilal, Sendonaris, Andrew, Dulac-Arnold, Gabriel, Osband, Ian, Agapiou, John, et al · 2017
Later among the works it cites.
Bridging the gap between value and policy based reinforcement learning
Nachum, Ofir, Norouzi, Mohammad, Xu, Kelvin, and Schuurmans, Dale · 2017
Later among the works it cites.
Equivalence between policy gradients and soft q-learning
Schulman, John, Abbeel, Pieter, and Chen, Xi · 2017
Later among the works it cites.
Robust imitation of diverse behaviors
Wang, Ziyu, Merel, Josh, Reed, Scott, Wayne, Greg, de Freitas, Nando, and Heess, Nicolas · 2017
Later among the works it cites.