Fetching the paper…
Reading the bibliography…
In practice, imitation learning is preferred over pure reinforcement learning whenever it is possible to design a teaching agent to provide expert supervision.
Efficient training of artificial neural networks for autonomous navigation
D. A. Pomerleau · 1991
Earlier work this paper cites.
Learning to fly
C. Sammut, S. Hurst, D. Kedzier, and D. Michie · 1992
Earlier work this paper cites.
A framework for behavioural cloning
M. Bain and C. Sammut · 1995
Earlier work this paper cites.
Asymptotic Statistics
A. van der Vaart · 2000
Earlier work this paper cites.
All of Nonparametric Statistics (Springer Texts in Statistics)
L. Wasserman · 2006
Earlier work this paper cites.
Reinforcement learning of motor skills with policy gradients
J. Peters and S. Schaal · 2008
Earlier work this paper cites.
Policy search for motor primitives in robotics
J. Kober and J. R. Peters · 2009
Earlier work this paper cites.
Efficient reductions for imitation learning
S. Ross and D. Bagnell · 2010
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
S. Ross, G. Gordon, and D. Bagnell · 2011
Earlier work this paper cites.
Robust asymmetric learning in pomdps
A. Warrington, J. W. Lavington, A. Ścibior, M. Schmidt, and F. Wood · 2012
Earlier work this paper cites.
Weighted importance sampling for off-policy learning with linear function approximation
A. R. Mahmood, H. P. van Hasselt, and R. S. Sutton · 2014
Earlier work this paper cites.
Reinforcement learning from demonstration through shaping
T. Brys, A. Harutyunyan, H. B. Suay, S. Chernova, M. E. Taylor, and A. Nowé · 2015
Earlier work this paper cites.
Learning to search better than your teacher
K.-W. Chang, A. Krishnamurthy, A. Agarwal, H. Daume, and J. Langford · 2015
Earlier work this paper cites.
Direct policy iteration with demonstrations
J. Chemali and A. Lazaric · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, S. Petersen, C. Beattie, A. Sadik, I. Antonoglou, H. King, D. Kumaran, D. Wierstra, S. Legg, and D. Hassabis · 2015
Earlier work this paper cites.
Increasing the action gap: New operators for reinforcement learning
M. G. Bellemare, G. Ostrovski, A. Guez, P. S. Thomas, and R. Munos · 2016
Earlier work this paper cites.
Dual learning for machine translation
D. He, Y. Xia, T. Qin, L. Wang, N. Yu, T.-Y. Liu, and W.-Y. Ma · 2016
Earlier work this paper cites.
Generative adversarial imitation learning
J. Ho and S. Ermon · 2016
Earlier work this paper cites.
End-to-end training of deep visuomotor policies
S. Levine, C. Finn, T. Darrell, and P. Abbeel · 2016
Earlier work this paper cites.
Continuous control with deep reinforcement learning
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra · 2016
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu · 2016
Earlier work this paper cites.
Prioritized Experience Replay
T. Schaul, J. Quan, I. Antonoglou, and D. Silver · 2016
Earlier work this paper cites.
Mastering the game of Go with deep neural networks and tree search
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. van den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, S. Dieleman, D. Grewe, J. Nham, N. Kalchbrenner, I. Sutskever, T. Lillicrap, M. Leach, K. Kavukcuoglu, T. Graepel, and D. Hassabis · 2016
Earlier work this paper cites.
One-step targeted minimum loss-based estimation based on universal least favorable one-dimensional submodels
M. van der Laan and S. Gruber · 2016
Cited alongside, same era.
Deep reinforcement learning with double q-learning
H. van Hasselt, A. Guez, and D. Silver · 2016
Cited alongside, same era.
Learning cooperative visual dialog agents with deep reinforcement learning
A. Das, S. Kottur, J. M. Moura, S. Lee, and D. Batra · 2017
Cited alongside, same era.
Emergence of locomotion behaviours in rich environments
N. Heess, D. TB, S. Sriram, J. Lemmon, J. Merel, G. Wayne, Y. Tassa, T. Erez, Z. Wang, S. Eslami, et al · 2017
Cited alongside, same era.
Reinforcement learning with unsupervised auxiliary tasks
M. Jaderberg, V. Mnih, W. M. Czarnecki, T. Schaul, J. Z. Leibo, D. Silver, and K. Kavukcuoglu · 2017
Cited alongside, same era.
Learning complex dexterous manipulation with deep reinforcement learning and demonstrations
A. Rajeswaran, V. Kumar, A. Gupta, G. Vezzani, J. Schulman, E. Todorov, and S. Levine · 2018
Later among the works it cites.
Truncated horizon policy search: Combining reinforcement learning & imitation learning
W. Sun, J. A. Bagnell, and B. Boots · 2018
Later among the works it cites.
Video captioning via hierarchical reinforcement learning
X. Wang, W. Chen, J. Wu, Y.-F. Wang, and W. Yang Wang · 2018
Later among the works it cites.
Gibson env: Real-world perception for embodied agents
F. Xia, A. R. Zamir, Z. He, A. Sax, J. Malik, and S. Savarese · 2018
Later among the works it cites.
Reinforcement and imitation learning for diverse visuomotor skills
Y. Zhu, Z. Wang, J. Merel, A. Rusu, T. Erez, S. Cabi, S. Tunyasuvunakool, J. Kramár, R. Hadsell, N. de Freitas, and N. Heess · 2018
Later among the works it cites.
Show your work: Improved reporting of experimental results
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
D. Kingma and J. Ba · 2017
Cited alongside, same era.
Multi-Agent Actor-Critic for Mixed Cooperative-Competitive Environments
R. Lowe, Y. Wu, A. Tamar, J. Harb, P. Abbeel, and I. Mordatch · 2017
Cited alongside, same era.
Learning to navigate in complex environments
P. Mirowski, R. Pascanu, F. Viola, H. Soyer, A. Ballard, A. Banino, M. Denil, R. Goroshin, L. Sifre, K. Kavukcuoglu, et al · 2017
Cited alongside, same era.
Curiosity-driven exploration by self-supervised prediction
D. Pathak, P. Agrawal, A. A. Efros, and T. Darrell · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Cited alongside, same era.
World of bits: An open-domain platform for web-based agents
T. T. Shi, A. Karpathy, L. J. Fan, J. Hernandez, and P. Liang · 2017
Cited alongside, same era.
Leveraging demonstrations for deep reinforcement learning on robotics problems with sparse rewards
M. Vecerik, T. Hester, J. Scholz, F. Wang, O. Pietquin, B. Piot, N. Heess, T. Rothörl, T. Lampe, and M. Riedmiller · 2017
Cited alongside, same era.
J. Dodge, S. Gururangan, D. Card, R. Schwartz, and N. A. Smith · 2019
Later among the works it cites.
Learning belief representations for imitation learning in pomdps
T. Gangwani, J. Lehman, Q. Liu, and J. Peng · 2019
Later among the works it cites.
Two body problem: Collaborative visual task completion
U. Jain, L. Weihs, E. Kolve, M. Rastegari, S. Lazebnik, A. Farhadi, A. G. Schwing, and A. Kembhavi · 2019
Later among the works it cites.
On value functions and the agent-environment boundary
N. Jiang · 2019
Later among the works it cites.
AI2-THOR: an interactive 3d environment for visual AI
E. Kolve, R. Mottaghi, W. Han, E. VanderBilt, L. Weihs, A. Herrasti, D. Gordon, Y. Zhu, A. Gupta, and A. Farhadi · 2019
Later among the works it cites.
PIC: Permutation Invariant Critic for Multi-Agent Deep Reinforcement Learning
I.-J. Liu, R. Yeh, and A. G. Schwing · 2019
Later among the works it cites.
Habitat: A Platform for Embodied AI Research
M. Savva, A. Kadian, O. Maksymets, Y. Zhao, E. Wijmans, B. Jain, J. Straub, J. Liu, V. Koltun, J. Malik, D. Parikh, and D. Batra · 2019
Later among the works it cites.
RoboTHOR: An Open Simulation-to-Real Embodied AI Platform
M. Deitke, W. Han, A. Herrasti, A. Kembhavi, E. Kolve, R. Mottaghi, J. Salvador, D. Schwenk, E. VanderBilt, M. Wallingford, L. Weihs, M. Yatskar, and A. Farhadi · 2020
Closest in time.
State-only imitation with transition dynamics mismatch
T. Gangwani and J. Peng · 2020
Closest in time.
Urban driving with conditional imitation learning
J. Hawke, R. Shen, C. Gurau, S. Sharma, D. Reda, N. Nikolov, P. Mazur, S. Micklethwaite, N. Griffiths, A. Shah, and A. Kendall · 2020
Closest in time.
A cordial sync: Going beyond marginal policies for multi-agent embodied tasks
U. Jain, L. Weihs, E. Kolve, A. Farhadi, S. Lazebnik, A. Kembhavi, and A. G. Schwing · 2020
Closest in time.
Reinforcement learning from imperfect demonstrations under soft expert guidance
M. Jing, X. Ma, W. Huang, F. Sun, C. Yang, B. Fang, and H. Liu · 2020
Closest in time.
On the interaction between supervision and self-play in emergent communication
R. Lowe, A. Gupta, J. N. Foerster, D. Kiela, and J. Pineau · 2020
Closest in time.
Alfred: A benchmark for interpreting grounded instructions for everyday tasks
M. Shridhar, J. Thomason, D. Gordon, Y. Bisk, W. Han, R. Mottaghi, L. Zettlemoyer, and D. Fox · 2020
Closest in time.
Multi-on: Benchmarking semantic map memory using multi-object navigation
S. Wani, S. Patel, U. Jain, A. X. Chang, and M. Savva · 2020
Closest in time.
Allenact: A framework for embodied ai research
L. Weihs, J. Salvador, K. Kotar, U. Jain, K.-H. Zeng, R. Mottaghi, and A. Kembhavi · 2020
Closest in time.
Gridtopix: Training embodied agents with minimal supervision
U. Jain, I.-J. Liu, S. Lazebnik, A. Kembhavi, L. Weihs, and A. G. Schwing · 2021
Closest in time.