Fetching the paper…
Reading the bibliography…
A critical flaw of existing inverse reinforcement learning (IRL) methods is their inability to significantly outperform the demonstrator.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Bradley, R. A. and Terry, M. E · 1952
Earlier work this paper cites.
Efficient training of artificial neural networks for autonomous navigation
Pomerleau, D. A · 1991
Earlier work this paper cites.
Introduction to reinforcement learning , volume 135
Sutton, R. S. and Barto, A. G · 1998
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Ng, A. Y., Harada, D., and Russell, S · 1999
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Abbeel, P. and Ng, A. Y · 2004
Earlier work this paper cites.
Bayesian inverse reinforcement learning
Ramachandran, D. and Amir, E · 2007
Earlier work this paper cites.
A game-theoretic approach to apprenticeship learning
Syed, U. and Schapire, R. E · 2008
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Ziebart, B. D., Maas, A. L., Bagnell, J. A., and Dey, A. K · 2008
Earlier work this paper cites.
A survey of robot learning from demonstration
Argall, B. D., Chernova, S., Veloso, M., and Browning, B · 2009
Earlier work this paper cites.
Policy search for motor primitives in robotics
Kober, J. and Peters, J. R · 2009
Earlier work this paper cites.
Preference-based policy learning
Akrour, R., Schoenauer, M., and Sebag, M · 2011
Earlier work this paper cites.
Relative entropy inverse reinforcement learning
Boularias, A., Kober, J., and Peters, J · 2011
Earlier work this paper cites.
Donut as i do: Learning from failed demonstrations
Grollman, D. H. and Billard, A · 2011
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Ross, S., Gordon, G., and Bagnell, D · 2011
Earlier work this paper cites.
Integrating reinforcement learning with human demonstrations of varying ability
Taylor, M. E., Suay, H. B., and Chernova, S · 2011
Earlier work this paper cites.
A survey of inverse reinforcement learning techniques
Gao, Y., Peters, J., Tsourdos, A., Zhifei, S., and Meng Joo, E · 2012
Earlier work this paper cites.
Individual choice behavior: A theoretical analysis
Luce, R. D · 2012
Earlier work this paper cites.
Preference-learning based inverse reinforcement learning for dialog control
Sugiyama, H., Meguro, T., and Minami, Y · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Todorov, E., Erez, T., and Tassa, Y · 2012
Cited alongside, same era.
Generative adversarial nets
Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Cited alongside, same era.
Markov decision processes: discrete stochastic dynamic programming
Puterman, M. L · 2014
Cited alongside, same era.
Robust bayesian inverse reinforcement learning with sparse behavior noise
Zheng, J., Liu, S., and Ni, L. M · 2014
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Deep q-learning from demonstrations
Hester, T., Vecerik, M., Pietquin, O., Lanctot, M., Schaul, T., Piot, B., Horgan, D., Quan, J., Sendonaris, A., Dulac-Arnold, G., et al · 2017
Later among the works it cites.
The atari grand challenge dataset
Kurin, V., Nowozin, S., Hofmann, K., Beyer, L., and Leibe, B · 2017
Later among the works it cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Later among the works it cites.
A survey of preference-based reinforcement learning methods
Wirth, C., Akrour, R., Neumann, G., and Fürnkranz, J · 2017
Later among the works it cites.
A survey of inverse reinforcement learning: Challenges, methods and progress
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Maximum entropy deep inverse reinforcement learning
Wulfmeier, M., Ondruska, P., and Posner, I · 2015
Cited alongside, same era.
Concrete problems in ai safety
Amodei, D., Olah, C., Steinhardt, J., Christiano, P., Schulman, J., and Mané, D · 2016
Cited alongside, same era.
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Cited alongside, same era.
Distance minimization for reward learning from scored trajectories
Burchfiel, B., Tomasi, C., and Parr, R · 2016
Cited alongside, same era.
Score-based inverse reinforcement learning
El Asri, L., Piot, B., Geist, M., Laroche, R., and Pietquin, O · 2016
Cited alongside, same era.
Guided cost learning: Deep inverse optimal control via policy optimization
Finn, C., Levine, S., and Abbeel, P · 2016
Cited alongside, same era.
Arora, S. and Doshi, P · 2018
Later among the works it cites.
Playing hard exploration games by watching youtube
Aytar, Y., Pfaff, T., Budden, D., Paine, T. L., Wang, Z., and de Freitas, N · 2018
Later among the works it cites.
Reinforcement learning from imperfect demonstrations
Gao, Y., Lin, J., Yu, F., Levine, S., Darrell, T., et al · 2018
Later among the works it cites.
Visualizing and understanding atari agents
Greydanus, S., Koul, A., Dodge, J., and Fern, A · 2018
Later among the works it cites.
Optiongan: Learning joint reward-policy options using generative adversarial inverse reinforcement learning
Henderson, P., Chang, W.-D., Bacon, P.-L., Meger, D., Pineau, J., and Precup, D · 2018
Later among the works it cites.
Reward learning from human preferences and demonstrations in atari
Ibarz, B., Leike, J., Pohlen, T., Irving, G., Legg, S., and Amodei, D · 2018
Later among the works it cites.
Imitation from observation: Learning to imitate behaviors from raw video via context translation
Liu, Y., Gupta, A., Abbeel, P., and Levine, S · 2018
Later among the works it cites.
An algorithmic perspective on imitation learning
Osa, T., Pajarinen, J., Neumann, G., Bagnell, J. A., Abbeel, P., Peters, J., et al · 2018
Later among the works it cites.
Adversarial imitation via variational inverse reinforcement learning
Qureshi, A. H. and Yip, M. C · 2018
Later among the works it cites.
Time-contrastive networks: Self-supervised learning from video
Sermanet, P., Lynch, C., Chebotar, Y., Hsu, J., Jang, E., Schaal, S., Levine, S., and Brain, G · 2018
Later among the works it cites.
Inverse reinforcement learning for video games
Tucker, A., Gleave, A., and Russell, S · 2018
Later among the works it cites.
One-shot imitation from observing humans via domain-adaptive meta-learning
Yu, T., Finn, C., Xie, A., Dasari, S., Zhang, T., Abbeel, P., and Levine, S · 2018
Later among the works it cites.
Robust learning from demonstrations with mixed qualities using leveraged gaussian processes
Choi, S., Lee, K., and Oh, S · 2019
Closest in time.
One-shot learning of multi-step tasks from observation via activity localization in auxiliary video
Goo, W. and Niekum, S · 2019
Closest in time.