Fetching the paper…
Reading the bibliography…
Behavior cloning (BC) optimizes policies by treating human demonstrations as pointwise action labels.
arXiv preprint arXiv:2009.12293
Zhu Y, Wong J, Mandlekar A, Martín-Martín R, Joshi A, Lin K, Maddukuri A, Nasiriany S and Zhu Y (2020) robosuite: A modular simulation framework and benchmark for robot learning · 2009
Earlier work this paper cites.
In: Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics , Proceedings of Machine Learning Research , volume 15. Fort Lauderdale, FL, USA: PMLR, pp. 627–635
Ross S, Gordon G and Bagnell D (2011) A reduction of imitation learning and structured prediction to no-regret online learning · 2011
Earlier work this paper cites.
In: Proceedings of the 28th International Conference on Machine Learning (ICML-11) . Citeseer, pp. 681–688
Welling M and Teh YW (2011) Bayesian learning via stochastic gradient langevin dynamics · 2011
Earlier work this paper cites.
The International Journal of Robotics Research 34(10): 1296–1313
Jain A, Sharma S, Joachims T and Saxena A (2015) Learning preferences for manipulation tasks from online coactive feedback · 2015
Earlier work this paper cites.
In: International Conference on Machine Learning . PMLR, pp. 2256–2265
Sohl-Dickstein J, Weiss E, Maheswaranathan N and Ganguli S (2015) Deep unsupervised learning using nonequilibrium thermodynamics · 2015
Earlier work this paper cites.
In: Proceedings of the 34th International Conference on Machine Learning , Proceedings of Machine Learning Research , volume 70. PMLR, pp. 233–242
Arpit D, Jastrzębski S, Ballas N, Krueger D, Bengio E, Kanwal MS, Maharaj T, Fischer A, Courville A, Bengio Y and Lacoste-Julien S (2017) A closer look at memorization in deep networks · 2017
Earlier work this paper cites.
In: Proceedings of the 1st Annual Conference on Robot Learning , Proceedings of Machine Learning Research , volume 78. PMLR, pp. 217–226
Bajcsy A, Losey DP, O’Malley MK and Dragan AD (2017) Learning robot objectives from physical human interaction · 2017
Earlier work this paper cites.
In: Advances in Neural Information Processing Systems , volume 30. pp. 4302–4310
Christiano PF, Leike J, Brown T, Martic M, Legg S and Amodei D (2017) Deep reinforcement learning from human preferences · 2017
Earlier work this paper cites.
Journal of Machine Learning Research 18(136): 1–46
Wirth C, Akrour R, Neumann G and Fürnkranz J (2017) A survey of preference-based reinforcement learning methods · 2017
Earlier work this paper cites.
In: Conference on Robot Learning . PMLR, pp. 123–132
Losey DP and O’Malley MK (2018) Including uncertainty when learning from human corrections · 2018
Earlier work this paper cites.
arXiv preprint arXiv:1807.03748
Oord Avd, Li Y and Vinyals O (2018) Representation learning with contrastive predictive coding · 2018
Earlier work this paper cites.
Foundations and Trends in Robotics 7(1-2): 1–179
Osa T, Pajarinen J, Neumann G, Bagnell JA, Abbeel P, Peters J et al. (2018) An algorithmic perspective on imitation learning · 2018
Earlier work this paper cites.
In: Proceedings of the 2018 International Symposium on Experimental Robotics . Springer, pp. 353–363
Pérez-Dattari R, Celemin C, Ruiz-del Solar J and Kober J (2020b) Interactive learning with corrective feedback for policies based on deep neural networks · 2018
Earlier work this paper cites.
In: International Conference on Machine Learning . PMLR, pp. 783–792
Brown D, Goo W, Nagarajan P and Niekum S (2019) Extrapolating beyond suboptimal demonstrations via inverse reinforcement learning from observations · 2019
Earlier work this paper cites.
The International Journal of Robotics Research 38(14): 1560–1580
Celemin C, Maeda G, Ruiz-del Solar J, Peters J and Kober J (2019) Reinforcement learning of motor skills using policy search and human corrective advice · 2019
Earlier work this paper cites.
Journal of Intelligent & Robotic Systems 95: 77–97
Celemin C and Ruiz-del Solar J (2019) An interactive framework for learning continuous actions policies based on corrective feedback · 2019
Earlier work this paper cites.
In: Advances in Neural Information Processing Systems , volume 32
Du Y and Mordatch I (2019) Implicit generation and modeling with energy based models · 2019
Earlier work this paper cites.
In: 2019 International Conference on Robotics and Automation (ICRA) . IEEE, pp. 8077–8083
Kelly M, Sidrane C, Driggs-Campbell K and Kochenderfer MJ (2019) Hg-dagger: Interactive imitation learning with human experts · 2019
Earlier work this paper cites.
In: 2019 International Conference on Robotics and Automation (ICRA) . IEEE, pp. 7611–7617
Pérez-Dattari R, Celemin C, Ruiz-del Solar J and Kober J (2019) Continuous control for high-dimensional state spaces: An interactive learning approach · 2019
Earlier work this paper cites.
In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . pp. 7518–7528
Gao R, Nijkamp E, Kingma DP, Xu Z, Dai AM and Wu YN (2020) Flow contrastive estimation of energy-based models · 2020
Cited alongside, same era.
In: Advances in Neural Information Processing Systems , volume 33. pp. 6840–6851
Ho J, Jain A and Abbeel P (2020) Denoising diffusion probabilistic models · 2020
Cited alongside, same era.
In: Advances in Neural Information Processing Systems , volume 33. pp. 21798–21809
Kalantidis Y, Sariyildiz MB, Pion N, Weinzaepfel P and Larlus D (2020) Hard negative mixing for contrastive learning · 2020
Cited alongside, same era.
Annual Review of Control, Robotics, and Autonomous Systems 3(1): 297–330
Ravichandar H, Polydoros AS, Chernova S and Billard A (2020) Recent advances in robot learning from demonstration · 2020
Cited alongside, same era.
In: 16th Robotics: Science and Systems, RSS 2020 . MIT Press Journals
Spencer J, Choudhury S, Barnes M, Schmittle M, Chiang M, Ramadge P and Srinivasa S (2020) Learning from interventions: Human-robot interaction as both explicit and implicit feedback · 2020
In: Conference on Robot Learning . PMLR, pp. 2340–2356
Datta G, Hoque R, Gu A, Solowjow E and Goldberg K (2023) Iifl: Implicit interactive fleet learning from heterogeneous human supervisors · 2023
Later among the works it cites.
In: Advances in Neural Information Processing Systems . pp. 77969 – 77992
Peng Z, Mo W, Duan C, Li Q and Zhou B (2023) Learning from active human involvement through proxy value propagation · 2023
Later among the works it cites.
In: Robotics: Science and Systems
Reuss M, Li M, Jia X and Lioutikov R (2023) Goal conditioned imitation learning using score-based diffusion policies · 2023
Later among the works it cites.
arXiv preprint arXiv:2305.10425
Zhao Y, Joshi R, Liu T, Khalman M, Saleh M and Liu PJ (2023) Slic-hf: Sequence likelihood calibration with human feedback · 2023
Later among the works it cites.
In: The Twelfth International Conference on Learning Representations , volume 2024. pp. 18770–18798
Hejna J, Rafailov R, Sikchi H, Finn C, Niekum S, Knox WB and Sadigh D (2024) Contrastive preference learning: Learning from human feedback without reinforcement learning · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
In: Advances in Neural Information Processing Systems , volume 33. pp. 3008–3021
Stiennon N, Ouyang L, Wu J, Ziegler D, Lowe R, Voss C, Radford A, Amodei D and Christiano PF (2020) Learning to summarize with human feedback · 2020
Cited alongside, same era.
In: Conference on Robot Learning . PMLR, pp. 682–692
Jauhri S, Celemin C and Kober J (2021) Interactive imitation learning in state-space · 2021
Cited alongside, same era.
In: International Conference on Machine Learning . PMLR, pp. 6152–6163
Lee K, Smith LM and Abbeel P (2021) Pebble: Feedback-efficient interactive reinforcement learning via relabeling experience and unsupervised pre-training · 2021
Cited alongside, same era.
Master’s thesis, Delft University of Technology
Lopez Bosque I (2021) Towards Corrective Deep Imitation Learning in Data Intensive Environments: Helping robots to learn faster by leveraging human knowledge · 2021
Cited alongside, same era.
arXiv preprint arXiv:2101.03288
Song Y and Kingma DP (2021) How to train your energy-based models · 2021
Cited alongside, same era.
In: International Conference on Learning Representations , volume 34. pp. 37799 – 37812
Song Y, Sohl-Dickstein J, Kingma DP, Kumar A, Ermon S and Poole B (2021) Score-based generative modeling through stochastic differential equations · 2021
Cited alongside, same era.
Communications of the ACM 64(3): 107–115
Zhang C, Bengio S, Hardt M, Recht B and Vinyals O (2021) Understanding deep learning (still) requires rethinking generalization · 2021
Cited alongside, same era.
Later among the works it cites.
arXiv preprint arXiv:2406.18629
Lai X, Tian Z, Chen Y, Yang S, Peng X and Jia J (2024) Step-dpo: Step-wise preference optimization for long-chain reasoning of llms · 2024
Later among the works it cites.
In: Advances in Neural Information Processing Systems , volume 36. pp. 53728 – 53741
Rafailov R, Sharma A, Mitchell E, Manning CD, Ermon S and Finn C (2024) Direct preference optimization: Your language model is secretly a reward model · 2024
Later among the works it cites.
Transactions on Machine Learning Research URL https://openreview.net/forum?id=JmKAYb7I00
Singh S, Tu S and Sindhwani V (2024) Revisiting energy based models as policies: Ranking noise contrastive estimation and interpolating energy models · 2024
Later among the works it cites.
arXiv preprint arXiv:2408.04380
Urain J, Mandlekar A, Du Y, Shafiullah M, Xu D, Fragkiadaki K, Chalvatzaki G and Peters J (2024) Deep generative models in robotics: A survey on learning from multimodal demonstrations · 2024
Later among the works it cites.
In: 8th Annual Conference on Robot Learning
Yang Z, Jun M, Tien J, Russell S, Dragan A and Biyik E (2024) Trajectory improvement and reward learning from comparative language feedback · 2024
Later among the works it cites.
IEEE Transactions on Cybernetics
Zare M, Kebria PM, Khosravi A and Nahavandi S (2024) A survey of imitation learning: Algorithms, recent developments, and challenges · 2024
Later among the works it cites.
IEEE Transactions on Robotics
Zhang Z, Hong J, Enayati AMS and Najjaran H (2024) Using implicit behavior cloning and dynamic movement primitive to facilitate reinforcement learning for robot motion planning · 2024
Later among the works it cites.
In: The Thirty-ninth Annual Conference on Neural Information Processing Systems
Cai H, Peng Z and Zhou B (2025) Predictive preference learning from human interventions · 2025
Closest in time.
In: 2025 IEEE International Conference on Robotics and Automation (ICRA) . pp. 6896–6903
Chen W, Xue H, Zhou F, Fang Y and Lu C (2025) Deformpam: Data-Efficient Learning for Long-Horizon Deformable Object Manipulation Via Preference-Based Action Alignment · 2025
Closest in time.
The International Journal of Robotics Research 44(10-11): 1684–1704
Chi C, Xu Z, Feng S, Cousineau E, Du Y, Burchfiel B, Tedrake R and Song S (2025) Diffusion policy: Visuomotor policy learning via action diffusion · 2025
Closest in time.
The International Journal of Robotics Research 44(4): 665–698
Habibian S, Valdivia AA, Blumenschein LH and Losey DP (2025) A survey of communicating robot learning during human-robot interaction · 2025
Closest in time.
In: Proceedings of Robotics: Science and Systems . LosAngeles, CA, USA
Jankowski J, Marić A, Liu P, Tateo D, Peters J and Calinon S (2025) Distilling Contact Planning for Fast Trajectory Optimization in Robot Air Hockey · 2025
Closest in time.
In: 2025 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, pp. 4845–4852
Lee SW, Kang X and Kuo YL (2025) Diff-dagger: Uncertainty estimation with diffusion policy for robotic manipulation · 2025
Closest in time.