Fetching the paper…
Reading the bibliography…
Neural policy learning methods have achieved remarkable results in various control problems, ranging from Atari games to simulated locomotion.
Dynamics-aware unsupervised discovery of skills
A. Sharma, S. Gu, S. Levine, V. Kumar, and K. Hausman · 1907
Earlier work this paper cites.
Alvinn: An autonomous land vehicle in a neural network
D. A. Pomerleau · 1988
Earlier work this paper cites.
Is imitation learning the route to humanoid robots?
S. Schaal · 1999
Earlier work this paper cites.
Recent advances in hierarchical reinforcement learning
A. G. Barto and S. Mahadevan · 2003
Earlier work this paper cites.
Survey: Robot programming by demonstration
A. Billard, S. Calinon, R. Dillmann, and S. Schaal · 2008
Earlier work this paper cites.
A survey of robot learning from demonstration
B. D. Argall, S. Chernova, M. Veloso, and B. Browning · 2009
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. Kingma and J. Ba · 2014
Earlier work this paper cites.
Modular multitask reinforcement learning with policy sketches, 2016
J. Andreas, D. Klein, and S. Levine · 2016
Earlier work this paper cites.
G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, S. Sidor, I. Sutskever, and R. S. Zemel · 2016
Earlier work this paper cites.
Gaussian error linear units (gelus)
D. Hendrycks and K. Gimpel · 2016
Earlier work this paper cites.
Generative adversarial imitation learning
J. Ho and S. Ermon · 2016
Earlier work this paper cites.
The option-critic architecture
P.-L. Bacon, J. Harb, and D. Precup · 2017
Earlier work this paper cites.
Stochastic neural networks for hierarchical reinforcement learning, 2017
C. Florensa, Y. Duan, and P. Abbeel · 2017
Earlier work this paper cites.
Categorical reparameterization with gumbel-softmax
E. Jang, S. Gu, and B. Poole · 2017
Earlier work this paper cites.
Learning multi-level hierarchies with hindsight, 2017
A. Levy, G. Konidaris, R. Platt, and K. Saenko · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Earlier work this paper cites.
Diversity is all you need: Learning skills without a reward function
B. Eysenbach, A. Gupta, J. Ibarz, and S. Levine · 2018
Earlier work this paper cites.
Addressing function approximation error in actor-critic methods
S. Fujimoto, H. Hoof, and D. Meger · 2018
Earlier work this paper cites.
Hierarchical imitation and reinforcement learning
H. Le, N. Jiang, A. Agarwal, M. Dudík, Y. Yue, and H. Daumé III · 2018
Earlier work this paper cites.
Data-efficient hierarchical reinforcement learning
O. Nachum, S. Gu, H. Lee, and S. Levine · 2018
Cited alongside, same era.
An algorithmic perspective on imitation learning
T. Osa, J. Pajarinen, G. Neumann, J. A. Bagnell, P. Abbeel, and J. Peters · 2018
Cited alongside, same era.
Deepmimic: Example-guided deep reinforcement learning of physics-based character skills
X. B. Peng, P. Abbeel, S. Levine, and M. van de Panne · 2018
Cited alongside, same era.
Vision-based multi-task manipulation for inexpensive robots using end-to-end learning from demonstration
R. Rahmatizadeh, P. Abolghasemi, L. Bölöni, and S. Levine · 2018
Cited alongside, same era.
Y. Tassa, D. Silver, J. Schrittwieser, A. Guez, L. Sifre, G. v. d. Driessche, S. Dieleman, K. Greff, T. Erez, and S. Petersen · 2018
Conservative q-learning for offline reinforcement learning
A. Kumar, A. Zhou, G. Tucker, and S. Levine · 2020
Later among the works it cites.
The nethack learning environment
H. Kuttler, N. Nardelli, T. Lavril, M. Selvatici, V. Sivakumar, M. G. Bellemare, R. Munos, A. Graves, and M. G. Bellemare · 2020
Later among the works it cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
S. Levine, A. Kumar, G. Tucker, and J. Fu · 2020
Later among the works it cites.
Sample factory: Egocentric 3d control from pixels at 100000 fps with asynchronous reinforcement learning
A. Petrenko, Z. Huang, T. Kumar, G. Sukhatme, and V. Koltun · 2020
Later among the works it cites.
Discovering motor programs by recomposing demonstrations
T. Shankar, S. Tulsiani, L. Pinto, and A. Gupta · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Deep imitation learning for complex manipulation tasks from virtual reality teleoperation
T. Zhang, Z. McCarthy, O. Jow, D. Lee, X. Chen, K. Goldberg, and P. Abbeel · 2018
Cited alongside, same era.
Reinforcement and imitation learning for diverse visuomotor skills
Y. Zhu, Z. Wang, J. Merel, A. Rusu, T. Erez, S. Cabi, S. Tunyasuvunakool, J. Kramár, R. Hadsell, N. de Freitas, et al · 2018
Cited alongside, same era.
Transformer-xl: Attentive language models beyond a fixed-length context
Z. Dai, Z. Yang, Y. Yang, J. Carbonell, Q. V. Le, and R. Salakhutdinov · 2019
Cited alongside, same era.
Self-supervised correspondence in visuomotor policy learning
P. Florence, L. Manuelli, and R. Tedrake · 2019
Cited alongside, same era.
Stabilizing off-policy q-learning via bootstrapping error reduction
A. Kumar, J. Fu, M. Soh, G. Tucker, and S. Levine · 2019
Cited alongside, same era.
Near-optimal representation learning for hierarchical reinforcement learning
O. Nachum, S. Gu, H. Lee, and S. Levine · 2019
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, et al · 2019
Cited alongside, same era.
A. Zeng, P. Florence, J. Tompson, S. Welker, J. Chien, M. Attarian, T. Armstrong, I. Krasin, D. Duong, V. Sindhwani, et al · 2020
Later among the works it cites.
Decision transformer: Reinforcement learning via sequence modeling
L. Chen, K. Lu, A. Rajeswaran, K. Lee, A. Grover, M. Laskin, P. Abbeel, A. Srinivas, and I. Mordatch · 2021
Later among the works it cites.
Assistive tele-op: Leveraging transformers to collect robotic task demonstrations
H. M. Clever, A. Handa, H. Mazhar, K. Parker, O. Shapira, Q. Wan, Y. Narang, I. Akinola, M. Cakmak, and D. Fox · 2021
Later among the works it cites.
Insights from the neurips 2021 nethack challenge
E. Hambro, S. Mohanty, D. Babaev, M. Byeon, D. Chakraborty, E. Grefenstette, M. Jiang, J. Daejin, A. Kanervisto, J. Kim, S. Kim, R. Kirk, V. Kurin, H. Küttler, T. Kwon, D. Lee, V. Mella, N. Nardelli, I. Nazarov, N. Ovsov, J. Holder, R. Raileanu, K. Ramanauskas, T. Rocktäschel, D. Rothermel, M. Samvelyan, D. Sorokin, M. Sypetkowski, and M. Sypetkowski · 2021
Later among the works it cites.
Perceiver: General perception with iterative attention
A. Jaegle, F. Gimeno, A. Brock, O. Vinyals, A. Zisserman, and J. Carreira · 2021
Later among the works it cites.
Offline reinforcement learning as one big sequence modeling problem
M. Janner, Q. Li, and S. Levine · 2021
Later among the works it cites.
Towards more generalizable one-shot visual imitation learning
Z. Mandi, F. Liu, K. Lee, and P. Abbeel · 2021
Later among the works it cites.
Amp: Adversarial motion priors for stylized physics-based character control
X. B. Peng, Z. Ma, P. Abbeel, S. Levine, and A. Kanazawa · 2021
Later among the works it cites.
Recurrent memory transformer
A. Bulatov, Y. Kuratov, and M. Burtsev · 2022
Later among the works it cites.
moolib: A Platform for Distributed RL
V. Mella, E. Hambro, D. Rothermel, and H. Küttler · 2022
Later among the works it cites.
S. Reed, K. Zolna, E. Parisotto, S. G. Colmenarejo, A. Novikov, G. Barth-Maron, M. Gimenez, Y. Sulsky, J. Kay, J. T. Springenberg, et al · 2022
Later among the works it cites.
Behavior transformers: Cloning k k modes with one stone
N. M. Shafiullah, Z. Cui, A. A. Altanzaya, and L. Pinto · 2022
Later among the works it cites.
Improving policy learning via language dynamics distillation
Y. Zhong, E. Hambro, E. Grefenstette, T. Park, M. Azar, and M. G. Bellemare · 2022
Later among the works it cites.
Learning about progress from experts
D. Bruce, E. Hambro, M. Azar, and M. G. Bellemare · 2023
Closest in time.
Human-timescale adaptation in an open-ended task space
A. A. Team, J. Bauer, K. Baumli, S. Baveja, F. Behbahani, A. Bhoopchand, N. Bradley-Schmieg, M. Chang, N. Clay, A. Collister, et al · 2023
Closest in time.