Fetching the paper…
Reading the bibliography…
Imitation learning addresses the challenge of learning by observing an expert's demonstrations without access to reward signals from environments.
Alvinn: An autonomous land vehicle in a neural network
Pomerleau, D. A · 1989
Earlier work this paper cites.
A framework for behavioural cloning
Bain, M. and Sammut, C · 1995
Earlier work this paper cites.
Learning from demonstration
Schaal, S · 1997
Earlier work this paper cites.
Algorithms for inverse reinforcement learning
Ng, A. Y. and Russell, S. J · 2000
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
Abbeel, P. and Ng, A. Y · 2004
Earlier work this paper cites.
Pattern recognition and machine learning
Bishop, C. M. and Nasrabadi, N. M · 2006
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
Ziebart, B. D., Maas, A. L., Bagnell, J. A., Dey, A. K., et al · 2008
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Ross, S., Gordon, G., and Bagnell, D · 2011
Earlier work this paper cites.
Knowledge fusion for probabilistic generative classifiers with data mining applications
Fisch, D., Kalkowski, E., and Sick, B · 2013
Earlier work this paper cites.
Generative adversarial nets
Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y · 2014
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. P. and Welling, M · 2014
Earlier work this paper cites.
A comprehensive survey on safe reinforcement learning
Garcıa, J. and Fernández, F · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al · 2015
Earlier work this paper cites.
Variational inference with normalizing flows
Rezende, D. and Mohamed, S · 2015
Earlier work this paper cites.
Deep unsupervised learning using nonequilibrium thermodynamics
Sohl-Dickstein, J., Weiss, E., Maheswaranathan, N., and Ganguli, S · 2015
Earlier work this paper cites.
Openai gym, 2016
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W · 2016
Earlier work this paper cites.
Deep learning
Goodfellow, I., Bengio, Y., and Courville, A · 2016
Earlier work this paper cites.
Generative adversarial imitation learning
Ho, J. and Ermon, S · 2016
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D · 2016
Earlier work this paper cites.
Deep reinforcement learning: A brief survey
Arulkumaran, K., Deisenroth, M. P., Brundage, M., and Bharath, A. A · 2017
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D · 2017
Earlier work this paper cites.
Density estimation using real NVP
Dinh, L., Sohl-Dickstein, J., and Bengio, S · 2017
Earlier work this paper cites.
Learning robust rewards with adverserial inverse reinforcement learning
Fu, J., Luo, K., and Levine, S · 2018
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S · 2018
Cited alongside, same era.
Imitation learning with concurrent actions in 3d games
Harmer, J., Gisslén, L., del Val, J., Holst, H., Bergdahl, J., Olsson, T., Sjöö, K., and Nordin, M · 2018
Cited alongside, same era.
Pytorch implementations of reinforcement learning algorithms
Kostrikov, I · 2018
Cited alongside, same era.
Scalable agent alignment via reward modeling: a research direction
Leike, J., Krueger, D., Everitt, T., Martic, M., Maini, V., and Legg, S · 2018
Cited alongside, same era.
Representation learning with contrastive predictive coding
Oord, A. v. d., Li, Y., and Vinyals, O · 2018
Cited alongside, same era.
Augmenting gail with bc for sample efficient imitation learning
Jena, R., Liu, C., and Sycara, K · 2021
Later among the works it cites.
Generalizable imitation learning from observation via inferring goal proximity
Lee, Y., Szot, A., Sun, S.-H., and Lim, J. J · 2021
Later among the works it cites.
Improved denoising diffusion probabilistic models
Nichol, A. Q. and Dhariwal, P · 2021
Later among the works it cites.
How to train your energy-based models
Song, Y. and Kingma, D. P · 2021
Later among the works it cites.
Score-based generative modeling through stochastic differential equations
Song, Y., Sohl-Dickstein, J., Kingma, D. P., Kumar, A., Ermon, S., and Poole, B · 2021
Later among the works it cites.
Bridging explicit and implicit deep generative models via neural stein estimators
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
An algorithmic perspective on imitation learning
Osa, T., Pajarinen, J., Neumann, G., Bagnell, J. A., Abbeel, P., Peters, J., et al · 2018
Cited alongside, same era.
Multi-goal reinforcement learning: Challenging robotics environments and request for research
Plappert, M., Andrychowicz, M., Ray, A., McGrew, B., Baker, B., Powell, G., Schneider, J., Tobin, J., Chociej, M., Welinder, P., et al · 2018
Cited alongside, same era.
Neural program synthesis from diverse demonstration videos
Sun, S.-H., Noh, H., Somasundaram, S., and Lim, J · 2018
Cited alongside, same era.
Behavioral cloning from observation
Torabi, F., Warnell, G., and Stone, P · 2018
Cited alongside, same era.
Exploring the limitations of behavior cloning for autonomous driving
Codevilla, F., Santana, E., López, A. M., and Gaidon, A · 2019
Cited alongside, same era.
Implicit generation and modeling with energy based models
Du, Y. and Mordatch, I · 2019
Cited alongside, same era.
Discriminator-actor-critic: Addressing sample inefficiency and reward bias in adversarial imitation learning
Kostrikov, I., Agrawal, K. K., Dwibedi, D., Levine, S., and Tompson, J · 2019
Cited alongside, same era.
Wu, Q., Gao, R., and Zha, H · 2021
Later among the works it cites.
Task-relevant adversarial imitation learning
Zolna, K., Reed, S., Novikov, A., Colmenarejo, S. G., Budden, D., Cabi, S., Denil, M., de Freitas, N., and Wang, Z · 2021
Later among the works it cites.
Implicit behavioral cloning
Florence, P., Lynch, C., Zeng, A., Ramirez, O. A., Wahid, A., Downs, L., Wong, A., Lee, J., Mordatch, I., and Tompson, J · 2022
Later among the works it cites.
Implicit kinematic policies: Unifying joint and cartesian action spaces in end-to-end robot learning
Ganapathi, A., Florence, P., Varley, J., Burns, K., Goldberg, K., and Zeng, A · 2022
Later among the works it cites.
Diagnosing and fixing manifold overfitting in deep generative models
Loaiza-Ganem, G., Ross, B. L., Cresswell, J. C., and Caterini, A. L · 2022
Later among the works it cites.
Understanding diffusion models: A unified perspective
Luo, C · 2022
Later among the works it cites.
Diffusion-based voice conversion with fast maximum likelihood sampling scheme
Popov, V., Vovk, I., Gogoryan, V., Sadekova, T., Kudinov, M., and Wei, J · 2022
Later among the works it cites.
Behavior transformers: Cloning k k modes with one stone
Shafiullah, N. M. M., Cui, Z. J., Altanzaya, A., and Pinto, L · 2022
Later among the works it cites.
High-level decision making for automated highway driving via behavior cloning
Wang, L., Fernandez, C., and Stiller, C · 2022
Later among the works it cites.
Diffusion policy: Visuomotor policy learning via action diffusion
Chi, C., Feng, S., Du, Y., Xu, Z., Cousineau, E., Burchfiel, B., and Song, S · 2023
Closest in time.
Learning to act from actionless videos through dense correspondences
Ko, P.-C., Mao, J., Du, Y., Sun, S.-H., and Tenenbaum, J. B · 2023
Closest in time.
Reliable conditioning of behavioral cloning for offline reinforcement learning
Nguyen, T., Zheng, Q., and Grover, A · 2023
Closest in time.
Imitating human behaviour with diffusion models
Pearce, T., Rashid, T., Kanervisto, A., Bignell, D., Sun, M., Georgescu, R., Macua, S. V., Tan, S. Z., Momennejad, I., Hofmann, K., and Devlin, S · 2023
Closest in time.
Dreamfusion: Text-to-3d using 2d diffusion
Poole, B., Jain, A., Barron, J. T., and Mildenhall, B · 2023
Closest in time.
Goal conditioned imitation learning using score-based diffusion policies
Reuss, M., Li, M., Jia, X., and Lioutikov, R · 2023
Closest in time.
Inverse reinforcement learning without reinforcement learning
Swamy, G., Wu, D., Choudhury, S., Bagnell, D., and Wu, S · 2023
Closest in time.
Learning fine-grained bimanual manipulation with low-cost hardware
Zhao, T. Z., Kumar, V., Levine, S., and Finn, C · 2023
Closest in time.
Diffusion-reward adversarial imitation learning
Lai, C.-M., Wang, H.-C., Hsieh, P.-C., Wang, Y.-C. F., Chen, M.-H., and Sun, S.-H · 2024
Closest in time.