Fetching the paper…
Reading the bibliography…
Imitation learning from human-provided demonstrations is a strong approach for learning policies for robot manipulation.
Mixture density networks
C. M. Bishop · 1994
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
NASA task load index (nasa-tlx); 20 years later
S. G. Hart · 2006
Earlier work this paper cites.
Active learning literature survey
B. Settles · 2009
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
S. Ross, G. Gordon, and A. Bagnell · 2011
Earlier work this paper cites.
Maximum mean discrepancy imitation learning
B. Kim and J. Pineau · 2013
Earlier work this paper cites.
Dropout: A simple way to prevent neural networks from overfitting
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. Kingma and J. Ba · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
J. L. Ba, J. R. Kiros, and G. E. Hinton · 2016
Earlier work this paper cites.
Gaussian error linear units (gelus)
D. Hendrycks and K. Gimpel · 2016
Earlier work this paper cites.
Multi-modal imitation learning from unstructured demonstrations using generative adversarial nets
K. Hausman, Y. Chebotar, S. Schaal, G. S. Sukhatme, and J. J. Lim · 2017
Earlier work this paper cites.
Query-efficient imitation learning for end-to-end autonomous driving
J. Zhang and K. Cho · 2017
Earlier work this paper cites.
Burn-in demonstrations for multi-modal imitation learning
A. Kuefler and M. J. Kochenderfer · 2018
Earlier work this paper cites.
HG-DAgger: Interactive imitation learning with human experts
M. Kelly, C. Sidrane, K. Driggs-Campbell, and M. J. Kochenderfer · 2019
Cited alongside, same era.
Learning latent plans from play
C. Lynch, M. Khansari, T. Xiao, V. Kumar, J. Tompson, S. Levine, and P. Sermanet · 2019
Cited alongside, same era.
Advantage-weighted regression: Simple and scalable off-policy reinforcement learning
X. B. Peng, A. Kumar, G. H. Zhang, and S. Levine · 2019
Cited alongside, same era.
Extrapolating beyond suboptimal demonstrations via inverse reinforcement learning from observations
D. S. Brown, W. Goo, P. Nagarajan, and S. Niekum · 2019
Cited alongside, same era.
Ensembledagger: A Bayesian approach to safe imitation learning
K. Menda, K. Driggs-Campbell, and M. J. Kochenderfer · 2019
Cited alongside, same era.
Confidence-aware imitation learning from demonstrations with varying optimality
S. Zhang, Z. Cao, D. Sadigh, and Y. Sui · 2021
Later among the works it cites.
Mind your outliers! investigating the negative impact of outliers on active learning for visual question answering
S. Karamcheti, R. Krishna, L. Fei-Fei, and C. D. Manning · 2021
Later among the works it cites.
ThriftyDAgger: Budget-aware novelty and risk gating for interactive imitation learning
R. Hoque, A. Balakrishna, E. R. Novoseller, A. Wilcox, D. S. Brown, and K. Goldberg · 2021
Later among the works it cites.
Learning multimodal rewards from rankings
V. Myers, E. Biyik, N. Anari, and D. Sadigh · 2021
Later among the works it cites.
LazyDAgger: Reducing context switching in interactive imitation learning
R. Hoque, A. Balakrishna, C. Putterman, M. Luo, D. S. Brown, D. Seita, B. Thananjeyan, E. R. Novoseller, and K. Goldberg · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Decoupled weight decay regularization
I. Loshchilov and F. Hutter · 2019
Cited alongside, same era.
Learning from interventions: Human-robot interaction as both explicit and implicit feedback
J. Spencer, S. Choudhury, M. Barnes, M. Schmittle, M. Chiang, P. J. Ramadge, and S. S. Srinivasa · 2020
Cited alongside, same era.
Critic regularized regression
Z. Wang, A. Novikov, K. Zolna, J. T. Springenberg, S. E. Reed, B. Shahriari, N. Siegel, J. Merel, C. Gulcehre, N. M. O. Heess, and N. de Freitas · 2020
Cited alongside, same era.
Dataset cartography: Mapping and diagnosing datasets with training dynamics
S. Swayamdipta, R. Schwartz, N. Lourie, Y. Wang, H. Hajishirzi, N. A. Smith, and Y. Choi · 2020
Cited alongside, same era.
Learning from suboptimal demonstration via self-supervised reward regression
L. Chen, R. R. Paleja, and M. C. Gombolay · 2020
Cited alongside, same era.
BC-Z: Zero-shot task generalization with robotic imitation learning
E. Jang, A. Irpan, M. Khansari, D. Kappler, F. Ebert, C. Lynch, S. Levine, and C. Finn · 2021
Cited alongside, same era.
Implicit behavioral cloning
P. R. Florence, C. Lynch, A. Zeng, O. Ramirez, A. Wahid, L. Downs, A. S. Wong, J. Lee, I. Mordatch, and J. Tompson · 2021
Cited alongside, same era.
Y. Lin, A. S. Wang, G. Sutanto, A. Rai, and F. Meier · 2021
Later among the works it cites.
Bridge data: Boosting generalization of robotic skills with cross-domain datasets
F. Ebert, Y. Yang, K. Schmeckpeper, B. Bucher, G. Georgakis, K. Daniilidis, C. Finn, and S. Levine · 2022
Closest in time.
Do as I can, not as I say: Grounding language in robotic affordances
M. Ahn, A. Brohan, N. Brown, Y. Chebotar, O. Cortes, B. David, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, D. Ho, J. Hsu, J. Ibarz, B. Ichter, A. Irpan, E. Jang, R. J. Ruano, K. Jeffrey, S. Jesmonth, N. J. Joshi, R. C. Julian, D. Kalashnikov, Y. Kuang, K.-H. Lee, S. Levine, Y. Lu, L. Luu, C. Parada, P. Pastor, J. Quiambao, K. Rao, J. Rettinghouse, D. M. Reyes, P. Sermanet, N. Sievers, C. Tan, A. Toshev, V. Vanhoucke, F. Xia, T. Xiao, P. Xu, S. Xu, and M. Yan · 2022
Closest in time.
Imitation learning by estimating expertise of demonstrators
M. Beliaev, A. Shih, S. Ermon, D. Sadigh, and R. Pedarsani · 2022
Closest in time.
Trail: Near-optimal imitation learning with suboptimal data
M. Yang, S. Levine, and O. Nachum · 2022
Closest in time.
Rvs: What is essential for offline RL via supervised learning?
S. Emmons, B. Eysenbach, I. Kostrikov, and S. Levine · 2022
Closest in time.
Negative result for learning from demonstration: Challenges for end-users teaching robots with task and motion planning abstractions
N. Gopalan, N. Moorman, M. Natarajan, M. C. Gombolay, and Georgia · 2022
Closest in time.
Bottom-up skill discovery from unsegmented demonstrations for long-horizon robot manipulation
Y. Zhu, P. Stone, and Y. Zhu · 2022
Closest in time.