Fetching the paper…
Reading the bibliography…
Real-world tasks of interest are generally poorly defined by human-readable descriptions and have no pre-defined reward signals unless it is defined by a human designer.
Asynchronous methods for deep reinforcement learning,
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. Harley, T. P. Lillicrap, D. Silver, K. Kavukcuoglu, · 1937
Earlier work this paper cites.
ALVINN: An Autonomous Land Vehicle in a Neural Network,
D. A. Pomerleau, · 1989
Earlier work this paper cites.
Feudal reinforcement learning,
P. Dayan, G. E. Hinton, · 1992
Earlier work this paper cites.
Algorithms for inverse reinforcement learning,
A. Y. Ng, S. J. Russell, · 2000
Earlier work this paper cites.
Trueskill™: A Bayesian skill rating system,
R. Herbrich, T. Minka, T. Graepel, · 2006
Earlier work this paper cites.
A survey of robot learning from demonstration,
B. D. Argall, S. Chernova, M. Veloso, B. Browning, · 2008
Earlier work this paper cites.
Interactively shaping agents via human reinforcement: The TAMER framework,
W. B. Knox, P. Stone, · 2009
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning,
S. Ross, G. Gordon, D. Bagnell, · 2011
Earlier work this paper cites.
Keyframe-based learning from demonstration,
B. Akgun, M. Cakmak, K. Jiang, A. L. Thomaz, · 2012
Earlier work this paper cites.
Trajectories and keyframes for kinesthetic teaching: A human-robot interaction perspective,
B. Akgun, M. Cakmak, J. W. Yoo, A. L. Thomaz, · 2012
Earlier work this paper cites.
Playing Atari with deep reinforcement learning,
V. Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, M. A. Riedmiller, · 2013
Earlier work this paper cites.
Learning monocular reactive UAV control in cluttered natural environments,
S. Ross, N. Melik-Barkhudarov, K. S. Shankar, A. Wendel, D. Dey, J. A. Bagnell, M. Hebert, · 2013
Earlier work this paper cites.
A machine learning approach to visual perception of forest trails for mobile robots,
A. Giusti, J. Guzzi, D. C. Cireşan, F.-L. He, J. P. Rodríguez, F. Fontana, M. Faessler, C. Forster, J. Schmidhuber, G. Di Caro, et al., · 2015
Earlier work this paper cites.
The Malmo platform for artificial intelligence experimentation,
M. Johnson, K. Hofmann, T. Hutton, D. Bignell, · 2016
Earlier work this paper cites.
End to end learning for self-driving cars,
M. Bojarski, D. D. Testa, D. Dworakowski, B. Firner, B. Flepp, P. Goyal, L. D. Jackel, M. Monfort, U. Muller, J. Zhang, X. Zhang, J. Zhao, K. Zieba, · 2016
Earlier work this paper cites.
Directions in hybrid intelligence: Complementing AI systems with human intelligence.,
E. Kamar, · 2016
Cited alongside, same era.
Learning manipulation trajectories using recurrent neural networks,
R. Rahmatizadeh, P. Abolghasemi, L. Bölöni, · 2016
Cited alongside, same era.
Guided cost learning: Deep inverse optimal control via policy optimization,
C. Finn, S. Levine, P. Abbeel, · 2016
Cited alongside, same era.
Continuous control with deep reinforcement learning,
T. Lillicrap, J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, D. Wierstra, · 2016
Cited alongside, same era.
Generative adversarial imitation learning,
J. Ho, S. Ermon, · 2016
Cited alongside, same era.
Hybrid intelligence,
D. Dellermann, P. Ebel, M. Söllner, J. M. Leimeister, · 2019
Later among the works it cites.
Extrapolating beyond suboptimal demonstrations via inverse reinforcement learning from observations,
D. Brown, W. Goo, P. Nagarajan, S. Niekum, · 2019
Later among the works it cites.
Efficiently combining human demonstrations and interventions for safe training of autonomous systems in real-time,
V. G. Goecks, G. M. Gremillion, V. J. Lawhern, J. Valasek, N. R. Waytowich, · 2019
Later among the works it cites.
MineRL: A large-scale dataset of Minecraft demonstrations,
W. H. Guss, B. Houghton, N. Topin, P. Wang, C. Codel, M. Veloso, R. Salakhutdinov, · 2019
Later among the works it cites.
Learning your way without map or compass: Panoramic target driven visual navigation,
D. Watkins-Valls, J. Xu, N. Waytowich, P. Allen, · 2020
Later among the works it cites.
Human-in-the-loop methods for data-driven and reinforcement learning systems,
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Redmon, S. Divvala, R. Girshick, A. Farhadi, · 2016
Cited alongside, same era.
Combining self-supervised learning and imitation for vision-based rope manipulation,
A. Nair, D. Chen, P. Agrawal, P. Isola, P. Abbeel, J. Malik, S. Levine, · 2017
Cited alongside, same era.
Feudal networks for hierarchical reinforcement learning,
A. S. Vezhnevets, S. Osindero, T. Schaul, N. Heess, M. Jaderberg, D. Silver, K. Kavukcuoglu, · 2017
Cited alongside, same era.
Interactive learning from policy-dependent human feedback,
J. MacGlashan, M. K. Ho, R. Loftin, B. Peng, G. Wang, D. L. Roberts, M. E. Taylor, M. L. Littman, · 2017
Cited alongside, same era.
Deep reinforcement learning from human preferences,
P. F. Christiano, J. Leike, T. B. Brown, M. Martic, S. Legg, D. Amodei, · 2017
Cited alongside, same era.
IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures,
L. Espeholt, H. Soyer, R. Munos, K. Simonyan, V. Mnih, T. Ward, Y. Doron, V. Firoiu, T. Harley, I. Dunning, S. Legg, K. Kavukcuoglu, · 2018
Cited alongside, same era.
Cycle-of-learning for autonomous systems from human interaction,
N. R. Waytowich, V. G. Goecks, V. J. Lawhern, · 2018
Cited alongside, same era.
V. G. Goecks, · 2020
Later among the works it cites.
Integrating behavior cloning and reinforcement learning for improved performance in dense and sparse reward environments,
V. G. Goecks, G. M. Gremillion, V. J. Lawhern, J. Valasek, N. R. Waytowich, · 2020
Later among the works it cites.
SQIL: imitation learning via reinforcement learning with sparse rewards,
S. Reddy, A. D. Dragan, S. Levine, · 2020
Later among the works it cites.
Scaling imitation learning in Minecraft,
A. Amiranashvili, N. Dorka, W. Burgard, V. Koltun, T. Brox, · 2020
Later among the works it cites.
Learning to walk in minutes using massively parallel deep reinforcement learning,
N. Rudin, D. Hoeller, P. Reist, M. Hutter, · 2021
Closest in time.
Replacing rewards with examples: Example-based policy search via recursive classification,
B. Eysenbach, S. Levine, R. Salakhutdinov, · 2021
Closest in time.
Inverse reinforcement learning with natural language goals,
L. Zhou, K. Small, · 2021
Closest in time.
GANcraft: Unsupervised 3D neural rendering of Minecraft worlds,
Z. Hao, A. Mallya, S. Belongie, M.-Y. Liu, · 2021
Closest in time.
NeurIPS 2021 competition proposal: The MineRL BASALT competition on learning from human feedback,
R. Shah, C. Wild, S. H. Wang, N. Alex, B. Houghton, W. Guss, S. Mohanty, A. Kanervisto, S. Milani, N. Topin, P. Abbeel, S. Russell, A. Dragan, · 2021
Closest in time.
Trial without error: Towards safe reinforcement learning via human intervention,
W. Saunders, G. Sastry, A. Stuhlmüller, O. Evans, · 2069
Closest in time.