Fetching the paper…
Reading the bibliography…
Imitation learning from human demonstrations enables robots to perform complex manipulation tasks and has recently witnessed huge success.
R. A. Bradley and M. E. Terry, “Rank analysis of incomplete block designs: I. the method of paired comparisons,” Biometrika , vol. 39, no. 3/4, pp. 324–345, 1952
1952
Earlier work this paper cites.
R. J. Williams, “Simple statistical gradient-following algorithms for connectionist reinforcement learning,” Machine learning , vol. 8, pp. 229–256, 1992
1992
Earlier work this paper cites.
R. S. Sutton, D. McAllester, S. Singh, and Y. Mansour, “Policy gradient methods for reinforcement learning with function approximation,” Advances in neural information processing systems , vol. 12, 1999
1999
Earlier work this paper cites.
A. Y. Ng, S. Russell et al. , “Algorithms for inverse reinforcement learning.” in Icml , vol. 1, no. 2, 2000, p. 2
2000
Earlier work this paper cites.
D. Silver, G. Lever, N. Heess, T. Degris, D. Wierstra, and M. Riedmiller, “Deterministic policy gradient algorithms,” in International conference on machine learning . Pmlr, 2014, pp. 387–395
2014
Earlier work this paper cites.
J. Sohl-Dickstein, E. Weiss, N. Maheswaranathan, and S. Ganguli, “Deep unsupervised learning using nonequilibrium thermodynamics,” in International conference on machine learning . PMLR, 2015, pp. 2256–2265
2015
Earlier work this paper cites.
P. F. Christiano, J. Leike, T. Brown, M. Martic, S. Legg, and D. Amodei, “Deep reinforcement learning from human preferences,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
D. Sadigh, A. Dragan, S. Sastry, and S. Seshia, Active preference-based learning of reward functions , 2017
2017
Earlier work this paper cites.
C. Wirth, R. Akrour, G. Neumann, and J. Fürnkranz, “A survey of preference-based reinforcement learning methods,” Journal of Machine Learning Research , vol. 18, no. 136, pp. 1–46, 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
E. Biyik and D. Sadigh, “Batch active preference-based learning of reward functions,” in Conference on robot learning . PMLR, 2018, pp. 519–528
2018
Earlier work this paper cites.
D. Brown, W. Goo, P. Nagarajan, and S. Niekum, “Extrapolating beyond suboptimal demonstrations via inverse reinforcement learning from observations,” in International conference on machine learning . PMLR, 2019, pp. 783–792
2019
Earlier work this paper cites.
2020
Cited alongside, same era.
J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” Advances in neural information processing systems , vol. 33, pp. 6840–6851, 2020
2020
Cited alongside, same era.
2020
Cited alongside, same era.
N. Stiennon, L. Ouyang, J. Wu, D. Ziegler, R. Lowe, C. Voss, A. Radford, D. Amodei, and P. F. Christiano, “Learning to summarize with human feedback,” Advances in Neural Information Processing Systems , vol. 33, pp. 3008–3021, 2020
2020
Cited alongside, same era.
2023
Later among the works it cites.
S. P. Arunachalam, S. Silwal, B. Evans, and L. Pinto, “Dexterous imitation made easy: A learning-based framework for efficient dexterous manipulation,” in 2023 ieee international conference on robotics and automation (icra) . IEEE, 2023, pp. 5954–5961
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial networks,” Communications of the ACM , vol. 63, no. 11, pp. 139–144, 2020
2020
Cited alongside, same era.
2020
Cited alongside, same era.
2021
Cited alongside, same era.
Y. Qin, Y.-H. Wu, S. Liu, H. Jiang, R. Yang, Y. Fu, and X. Wang, “Dexmv: Imitation learning for dexterous manipulation from human videos,” in European Conference on Computer Vision . Springer, 2022, pp. 570–587
2022
Cited alongside, same era.
2022
Cited alongside, same era.
P. Florence, C. Lynch, A. Zeng, O. A. Ramirez, A. Wahid, L. Downs, A. Wong, J. Lee, I. Mordatch, and J. Tompson, “Implicit behavioral cloning,” in Conference on Robot Learning . PMLR, 2022, pp. 158–168
2022
Cited alongside, same era.
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray et al. , “Training language models to follow instructions with human feedback,” Advances in neural information processing systems , vol. 35, pp. 27 730–27 744, 2022
2022
Cited alongside, same era.
Y. Ze, G. Yan, Y.-H. Wu, A. Macaluso, Y. Ge, J. Ye, N. Hansen, L. E. Li, and X. Wang, “Gnfactor: Multi-task real robot learning with generalizable neural feature fields,” in Conference on Robot Learning . PMLR, 2023, pp. 284–301
2023
Cited alongside, same era.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2023
Later among the works it cites.
2024
Later among the works it cites.
Y. Fan, O. Watkins, Y. Du, H. Liu, M. Ryu, C. Boutilier, P. Abbeel, M. Ghavamzadeh, K. Lee, and K. Lee, “Reinforcement learning for fine-tuning text-to-image diffusion models,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Later among the works it cites.
2024
Later among the works it cites.