Fetching the paper…
Reading the bibliography…
Offline imitation learning (IL) is a powerful method to solve decision-making problems from expert demonstrations without reward labels.
Rings of real-valued continuous functions
E. Hewitt · 1948
Earlier work this paper cites.
Alvinn: An autonomous land vehicle in a neural network
D. A. Pomerleau · 1988
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
A. Y. Ng, D. Harada, and S. Russell · 1999
Earlier work this paper cites.
Calculus of variations
I. M. Gelfand, R. A. Silverman, et al · 2000
Earlier work this paper cites.
Maximum entropy inverse reinforcement learning
B. D. Ziebart, A. L. Maas, J. A. Bagnell, A. K. Dey, et al · 2008
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
S. Ross, G. J. Gordon, and J. A. Bagnell · 2011
Earlier work this paper cites.
Probabilistic model-based imitation learning
P. Englert, A. Paraschos, M. P. Deisenroth, and J. Peters · 2013
Earlier work this paper cites.
Generative adversarial nets
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio · 2014
Earlier work this paper cites.
Explaining and harnessing adversarial examples
I. J. Goodfellow, J. Shlens, and C. Szegedy · 2015
Earlier work this paper cites.
Generative adversarial imitation learning
J. Ho and S. Ermon · 2016
Earlier work this paper cites.
Query-efficient imitation learning for end-to-end simulated driving
J. Zhang and K. Cho · 2017
Earlier work this paper cites.
Wasserstein generative adversarial networks
M. Arjovsky, S. Chintala, and L. Bottou · 2017
Earlier work this paper cites.
Deep reinforcement learning in a handful of trials using probabilistic dynamics models
K. Chua, R. Calandra, R. McAllister, and S. Levine · 2018
Earlier work this paper cites.
Survey of imitation learning for robotic manipulation
B. Fang, S. Jia, D. Guo, M. Xu, S. Wen, and F. Sun · 2019
Earlier work this paper cites.
Discriminator-actor-critic: Addressing sample inefficiency and reward bias in adversarial imitation learning
I. Kostrikov, K. K. Agrawal, D. Dwibedi, S. Levine, and J. Tompson · 2019
Earlier work this paper cites.
Hg-dagger: Interactive imitation learning with human experts
M. Kelly, C. Sidrane, K. Driggs-Campbell, and M. Kochenderfer · 2019
Earlier work this paper cites.
Variational discriminator bottleneck: Improving imitation learning, inverse rl, and gans by constraining information flow
X. B. Peng, A. Kanazawa, S. Toyer, P. Abbeel, and S. Levine · 2019
Cited alongside, same era.
Stabilizing off-policy q-learning via bootstrapping error reduction
A. Kumar, J. Fu, M. Soh, G. Tucker, and S. Levine · 2019
Cited alongside, same era.
When to trust your model: Model-based policy optimization
M. Janner, J. Fu, M. Zhang, and S. Levine · 2019
Cited alongside, same era.
On evaluating adversarial robustness
N. Carlini, A. Athalye, N. Papernot, W. Brendel, J. Rauber, D. Tsipras, I. Goodfellow, A. Madry, and A. Kurakin · 2019
Cited alongside, same era.
Dualdice: Behavior-agnostic estimation of discounted stationary distribution corrections
O. Nachum, Y. Chow, B. Dai, and L. Li · 2019
Cited alongside, same era.
Feedback in imitation learning: The three regimes of covariate shift
J. C. Spencer, S. Choudhury, A. Venkatraman, B. D. Ziebart, and J. A. Bagnell · 2021
Later among the works it cites.
Mitigating covariate shift in imitation learning via offline data with partial coverage
J. Chang, M. Uehara, D. Sreenivas, R. Kidambi, and W. Sun · 2021
Later among the works it cites.
No need for interactions: Robust model-based imitation learning using neural ode
H. Lin, B. Li, X. Zhou, J. Wang, and M. Q.-H. Meng · 2021
Later among the works it cites.
Iq-learn: Inverse soft-q learning for imitation
D. Garg, S. Chakraborty, C. Cundy, J. Song, and S. Ermon · 2021
Later among the works it cites.
Visual adversarial imitation learning using variational models
R. Rafailov, T. Yu, A. Rajeswaran, and C. Finn · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Imitation learning via off-policy distribution matching
I. Kostrikov, O. Nachum, and J. Tompson · 2020
Cited alongside, same era.
Toward the fundamental limits of imitation learning
N. Rajaraman, L. F. Yang, J. Jiao, and K. Ramchandran · 2020
Cited alongside, same era.
Offline learning from demonstrations and unlabeled experience
K. Zolna, A. Novikov, K. Konyushkova, Ç. Gülçehre, Z. Wang, Y. Aytar, M. Denil, N. de Freitas, and S. E. Reed · 2020
Cited alongside, same era.
Semi-supervised reward learning for offline reinforcement learning
K. Konyushkova, K. Zolna, Y. Aytar, A. Novikov, S. Reed, S. Cabi, and N. de Freitas · 2020
Cited alongside, same era.
Mopo: Model-based offline policy optimization
T. Yu, G. Thomas, L. Yu, S. Ermon, J. Y. Zou, S. Levine, C. Finn, and T. Ma · 2020
Cited alongside, same era.
Offline imitation learning with a misspecified simulator
S. Jiang, J. Pang, and Y. Yu · 2020
Cited alongside, same era.
D4rl: Datasets for deep data-driven reinforcement learning
J. Fu, A. Kumar, O. Nachum, G. Tucker, and S. Levine · 2020
Cited alongside, same era.
P. Florence, C. Lynch, A. Zeng, O. A. Ramirez, A. Wahid, L. Downs, A. Wong, J. Lee, I. Mordatch, and J. Tompson · 2021
Later among the works it cites.
Curriculum offline imitating learning
M. Liu, H. Zhao, Z. Yang, J. Shen, W. Zhang, L. Zhao, and T.-Y. Liu · 2021
Later among the works it cites.
Of moments and matching: A game-theoretic framework for closing the imitation gap
G. Swamy, S. Choudhury, J. A. Bagnell, and S. Wu · 2021
Later among the works it cites.
Robust maximum entropy behavior cloning
M. Hussein, B. Crowe, M. Petrik, and M. Begum · 2021
Later among the works it cites.
Behavioral cloning from noisy demonstrations
F. Sasaki and R. Yamashina · 2021
Later among the works it cites.
A survey on imitation learning techniques for end-to-end autonomous vehicles
L. Le Mero, D. Yi, M. Dianati, and A. Mouzakitis · 2022
Closest in time.
On covariate shift of latent confounders in imitation and reinforcement learning
G. Tennenholtz, A. Hallak, G. Dalal, S. Mannor, G. Chechik, and U. Shalit · 2022
Closest in time.
What matters in learning from offline human demonstrations for robot manipulation
A. Mandlekar, D. Xu, J. Wong, S. Nasiriany, C. Wang, R. Kulkarni, L. Fei-Fei, S. Savarese, Y. Zhu, and R. Martín-Martín · 2022
Closest in time.
Discriminator-weighted offline imitation learning from suboptimal demonstrations
H. Xu, X. Zhan, H. Yin, and H. Qin · 2022
Closest in time.
Model-based offline planning with trajectory pruning
X. Zhan, X. Zhu, and H. Xu · 2022
Closest in time.
DemoDICE: Offline imitation learning with supplementary imperfect demonstrations
G.-H. Kim, S. Seo, J. Lee, W. Jeon, H. Hwang, H. Yang, and K.-E. Kim · 2022
Closest in time.