Fetching the paper…
Reading the bibliography…
This work considers two distinct settings: imitation learning and goal-conditioned reinforcement learning.
Sqil: Imitation learning via regularized behavioral cloning
Reddy, S., Dragan, A. D., and Levine, S. (2019) · 1905
Earlier work this paper cites.
Is the policy gradient a gradient?
Nota, C. and Thomas, P. S. (2019) · 1906
Earlier work this paper cites.
Alvinn: An autonomous land vehicle in a neural network
Pomerleau, D. A. (1989) · 1989
Earlier work this paper cites.
Policy Gradient Methods for Reinforcement Learning with Function Approximation
Sutton, R. S., Mcallester, D., Singh, S., and Mansour, Y. (1999) · 1999
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Abbeel, P. and Ng, A. Y. (2004) · 2004
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Ziebart, B. D., Maas, A. L., Bagnell, J. A., and Dey, A. K. (2008) · 2008
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Ross, S., Gordon, G., and Bagnell, D. (2011) · 2011
Earlier work this paper cites.
Horde: A scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction
Sutton, R. S., Modayil, J., Delp, M., Degris, T., Pilarski, P. M., White, A., and Precup, D. (2011) · 2011
Earlier work this paper cites.
Generative adversarial nets
Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y. (2014) · 2014
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. P. and Welling, M. (2014) · 2014
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Puterman, M. L. (2014) · 2014
Earlier work this paper cites.
Deterministic policy gradient algorithms
Silver, D., Lever, G., Heess, N., Degris, T., Wierstra, D., and Riedmiller, M. (2014) · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al. (2015) · 2015
Cited alongside, same era.
Universal Value Function Approximators
Schaul, T., Horgan, D., Gregor, K., and Silver, D. (2015) · 2015
Cited alongside, same era.
OpenAI Gym
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W. (2016) · 2016
Cited alongside, same era.
Density estimation using Real NVP
Dinh, L., Sohl-Dickstein, J., and Bengio, S. (2016) · 2016
Cited alongside, same era.
Generative adversarial imitation learning
Ho, J. and Ermon, S. (2016) · 2016
Cited alongside, same era.
Safe and efficient off-policy reinforcement learning
Munos, R., Stepleton, T., Harutyunyan, A., and Bellemare, M. (2016) · 2016
Learning robust rewards with adverserial inverse reinforcement learning
Fu, J., Luo, K., and Levine, S. (2018) · 2018
Later among the works it cites.
Addressing function approximation error in actor-critic methods
Fujimoto, S., Hoof, H., and Meger, D. (2018) · 2018
Later among the works it cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S. (2018) · 2018
Later among the works it cites.
ARCHER: Aggressive Rewards to Counter bias in Hindsight Experience Replay
Lanka, S. and Wu, T. (2018) · 2018
Later among the works it cites.
Visual reinforcement learning with imagined goals
Nair, A. V., Pong, V., Dalal, M., Bahl, S., Lin, S., and Levine, S. (2018) · 2018
Later among the works it cites.
Zero-shot visual imitation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Wavenet: A generative model for raw audio
van den Oord, A., Dieleman, S., Zen, H., Simonyan, K., Vinyals, O., Graves, A., Kalchbrenner, N., Senior, A., and Kavukcuoglu, K. (2016) · 2016
Cited alongside, same era.
Pixel Recurrent Neural Networks
van den Oord, A., Kalchbrenner, N., and Kavukcuoglu, K. (2016) · 2016
Cited alongside, same era.
Hindsight experience replay
Andrychowicz, M., Wolski, F., Ray, A., Schneider, J., Fong, R., Welinder, P., McGrew, B., Tobin, J., Abbeel, O. P., and Zaremba, W. (2017) · 2017
Cited alongside, same era.
State Aware Imitation Learning
Schroecker, Y. and Isbell, C. L. (2017) · 2017
Cited alongside, same era.
Parallel wavenet: Fast high-fidelity speech synthesis
van den Oord, A., Li, Y., Babuschkin, I., Simonyan, K., Vinyals, O., Kavukcuoglu, K., van den Driessche, G., Lockhart, E., Cobo, L. C., Stimberg, F., et al. (2017) · 2017
Cited alongside, same era.
Playing hard exploration games by watching youtube
Aytar, Y., Pfaff, T., Budden, D., Paine, T., Wang, Z., and de Freitas, N. (2018) · 2018
Cited alongside, same era.
Pathak, D., Mahmoudieh, P., Luo, G., Agrawal, P., Chen, D., Shentu, Y., Shelhamer, E., Malik, J., Efros, A. A., and Darrell, T. (2018) · 2018
Later among the works it cites.
Deepmimic: Example-guided deep reinforcement learning of physics-based character skills
Peng, X. B., Abbeel, P., Levine, S., and van de Panne, M. (2018) · 2018
Later among the works it cites.
Temporal Difference Models: Model-Free Deep RL for Model-Based Control
Pong, V., Gu, S., Dalal, M., and Levine, S. (2018) · 2018
Later among the works it cites.
Discriminator-actor-critic: Addressing sample inefficiency and reward bias in adversarial imitation learning
Kostrikov, I., Agrawal, K. K., Dwibedi, D., Levine, S., and Tompson, J. (2019) · 2019
Later among the works it cites.
Visual Hindsight Experience Replay
Sahni, H., Buckley, T., Abbeel, P., and Kuzovkin, I. (2019) · 2019
Later among the works it cites.
Generative predecessor models for sample-efficient imitation learning
Schroecker, Y., Vecerik, M., and Scholz, J. (2019) · 2019
Later among the works it cites.
Random expert distillation: Imitation learning via expert policy support estimation
Wang, R., Ciliberto, C., Amadori, P. V., and Demiris, Y. (2019) · 2019
Later among the works it cites.