Fetching the paper…
Reading the bibliography…
Imitation learning aims to solve the problem of defining reward functions in real-world decision-making tasks.
Imitation learning via off-policy distribution matching
Kostrikov, I.; Nachum, O.; and Tompson, J. 2019 · 1912
Earlier work this paper cites.
Efficient training of artificial neural networks for autonomous navigation
Pomerleau, D. A. 1991 · 1991
Earlier work this paper cites.
Denoising diffusion implicit models
Song, J.; Meng, C.; and Ermon, S. 2020 · 2010
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Ross, S.; Gordon, G.; and Bagnell, D. 2011 · 2011
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Todorov, E.; Erez, T.; and Tassa, Y. 2012 · 2012
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. P.; and Welling, M. 2013 · 2013
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Mnih, V.; Kavukcuoglu, K.; Silver, D.; Graves, A.; Antonoglou, I.; Wierstra, D.; and Riedmiller, M. 2013 · 2013
Earlier work this paper cites.
Generative adversarial nets
Goodfellow, I.; Pouget-Abadie, J.; Mirza, M.; Xu, B.; Warde-Farley, D.; Ozair, S.; Courville, A.; and Bengio, Y. 2014 · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P.; and Ba, J. 2014 · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Mnih, V.; Kavukcuoglu, K.; Silver, D.; Rusu, A. A.; Veness, J.; Bellemare, M. G.; Graves, A.; Riedmiller, M.; Fidjeland, A. K.; Ostrovski, G.; et al. 2015 · 2015
Earlier work this paper cites.
Trust region policy optimization
Schulman, J.; Levine, S.; Abbeel, P.; Jordan, M.; and Moritz, P. 2015 · 2015
Earlier work this paper cites.
Generative adversarial imitation learning
Ho, J.; and Ermon, S. 2016 · 2016
Earlier work this paper cites.
f-gan: Training generative neural samplers using variational divergence minimization
Nowozin, S.; Cseke, B.; and Tomioka, R. 2016 · 2016
Earlier work this paper cites.
Mastering the game of Go with deep neural networks and tree search
Silver, D.; Huang, A.; Maddison, C. J.; Guez, A.; Sifre, L.; Van Den Driessche, G.; Schrittwieser, J.; Antonoglou, I.; Panneershelvam, V.; Lanctot, M.; et al. 2016 · 2016
Earlier work this paper cites.
Learning robust rewards with adversarial inverse reinforcement learning
Fu, J.; Luo, K.; and Levine, S. 2017 · 2017
Cited alongside, same era.
Improved training of wasserstein gans
Gulrajani, I.; Ahmed, F.; Arjovsky, M.; Dumoulin, V.; and Courville, A. C. 2017 · 2017
Cited alongside, same era.
Proximal policy optimization algorithms
Schulman, J.; Wolski, F.; Dhariwal, P.; Radford, A.; and Klimov, O. 2017 · 2017
Cited alongside, same era.
Addressing function approximation error in actor-critic methods
Fujimoto, S.; Hoof, H.; and Meger, D. 2018 · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T.; Zhou, A.; Abbeel, P.; and Levine, S. 2018 · 2018
Cited alongside, same era.
Analytic-dpm: an analytic estimate of the optimal reverse variance in diffusion probabilistic models
Bao, F.; Li, C.; Zhu, J.; and Zhang, B. 2022 · 2022
Later among the works it cites.
Planning with diffusion for flexible behavior synthesis
Janner, M.; Du, Y.; Tenenbaum, J. B.; and Levine, S. 2022 · 2022
Later among the works it cites.
Rethinking ValueDice: Does It Really Improve Performance?
Li, Z.; Xu, T.; Yu, Y.; and Luo, Z.-Q. 2022 · 2022
Later among the works it cites.
Pseudo numerical methods for diffusion models on manifolds
Liu, L.; Ren, Y.; Lin, Z.; and Zhao, Z. 2022 · 2022
Later among the works it cites.
Flow straight and fast: Learning to generate and transfer data with rectified flow
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kostrikov, I.; Agrawal, K. K.; Dwibedi, D.; Levine, S.; and Tompson, J. 2018 · 2018
Cited alongside, same era.
Dualdice: Behavior-agnostic estimation of discounted stationary distribution corrections
Nachum, O.; Chow, Y.; Dai, B.; and Li, L. 2019 · 2019
Cited alongside, same era.
A divergence minimization perspective on imitation learning methods
Ghasemipour, S. K. S.; Zemel, R.; and Gu, S. 2020 · 2020
Cited alongside, same era.
Denoising diffusion probabilistic models
Ho, J.; Jain, A.; and Abbeel, P. 2020 · 2020
Cited alongside, same era.
f-gail: Learning f-divergence for generative adversarial imitation learning
Zhang, X.; Li, Y.; Zhang, Z.; and Zhang, Z.-L. 2020 · 2020
Cited alongside, same era.
Off-policy imitation learning from observations
Zhu, Z.; Lin, K.; Dai, B.; and Zhou, J. 2020 · 2020
Cited alongside, same era.
Iq-learn: Inverse soft-q learning for imitation
Garg, D.; Chakraborty, S.; Cundy, C.; Song, J.; and Ermon, S. 2021 · 2021
Cited alongside, same era.
Liu, X.; Gong, C.; and Liu, Q. 2022 · 2022
Later among the works it cites.
Diffusion policies as an expressive policy class for offline reinforcement learning
Wang, Z.; Hunt, J. J.; and Zhou, M. 2022 · 2022
Later among the works it cites.
Auto-Encoding Adversarial Imitation Learning
Zhang, K.; Zhao, R.; Zhang, Z.; and Gao, Y. 2022 · 2022
Later among the works it cites.
A Coupled Flow Approach to Imitation Learning
Freund, G.; Sarafian, E.; and Kraus, S. 2023 · 2023
Closest in time.
How To Guide Your Learner: Imitation Learning with Active Adaptive Expert Involvement
Liu, X.-H.; Xu, F.; Zhang, X.; Liu, T.; Jiang, S.; Chen, R.; Zhang, Z.; and Yu, Y. 2023 · 2023
Closest in time.
Imitating human behaviour with diffusion models
Pearce, T.; Rashid, T.; Kanervisto, A.; Bignell, D.; Sun, M.; Georgescu, R.; Macua, S. V.; Tan, S. Z.; Momennejad, I.; Hofmann, K.; et al. 2023 · 2023
Closest in time.
Goal-conditioned imitation learning using score-based diffusion policies
Reuss, M.; Li, M.; Jia, X.; and Lioutikov, R. 2023 · 2023
Closest in time.
Song, Y.; Dhariwal, P.; Chen, M.; and Sutskever, I. 2023 · 2023
Closest in time.
DPM-Solver-v3: Improved Diffusion ODE Solver with Empirical Model Statistics
Zheng, K.; Lu, C.; Chen, J.; and Zhu, J. 2023 · 2023
Closest in time.