Fetching the paper…
Reading the bibliography…
Reinforcement learning (RL) algorithms face significant challenges when dealing with long-horizon robot manipulation tasks in real-world environments due to sample inefficiency and safety issues.
A unified approach for motion and force control of robot manipulators: The operational space formulation
Oussama Khatib · 1987
Earlier work this paper cites.
A social reinforcement learning agent
Charles Isbell, Christian R Shelton, Michael Kearns, Satinder Singh, and Peter Stone · 2001
Earlier work this paper cites.
Smote: Synthetic minority over-sampling technique
Nitesh V. Chawla, Kevin W. Bowyer, Lawrence O. Hall, and W. Philip Kegelmeyer · 2002
Earlier work this paper cites.
Reinforcement learning with human teachers: Evidence of feedback and guidance with implications for learning performance
Andrea L. Thomaz and Cynthia Breazeal · 2006
Earlier work this paper cites.
Interactively shaping agents via human reinforcement: The tamer framework
W Bradley Knox and Peter Stone · 2009
Earlier work this paper cites.
Dynamic reward shaping: training a robot by voice
Ana C Tenorio-Gonzalez, Eduardo F Morales, and Luis Villaseñor-Pineda · 2010
Earlier work this paper cites.
Combining manual feedback with subsequent mdp reward signals for reinforcement learning
W Bradley Knox and Peter Stone · 2010
Earlier work this paper cites.
Reinforcement learning from simultaneous human and mdp reward
W Bradley Knox and Peter Stone · 2012
Earlier work this paper cites.
On developing robust models for favourability analysis: Model choice, feature sets and imbalanced data
Peter C.R. Lane, Daoud Clarke, and Paul Hender · 2012
Earlier work this paper cites.
Policy shaping: Integrating human feedback with reinforcement learning
Shane Griffith, Kaushik Subramanian, Jonathan Scholz, Charles L Isbell, and Andrea L Thomaz · 2013
Earlier work this paper cites.
Training a robot via human feedback: A case study
W Bradley Knox, Peter Stone, and Cynthia Breazeal · 2013
Earlier work this paper cites.
A constraint-based method for solving sequential manipulation planning problems
Tomás Lozano-Pérez and Leslie Pack Kaelbling · 2014
Earlier work this paper cites.
Reinforcement learning improves behaviour from evaluative feedback
Michael L Littman · 2015
Earlier work this paper cites.
Logic-geometric programming: An optimization-based approach to combined task and motion planning
Marc Toussaint · 2015
Earlier work this paper cites.
Training a robot with evaluative feedback and unlabeled guidance signals
Anis Najar, Olivier Sigaud, and Mohamed Chetouani · 2016
Earlier work this paper cites.
Prioritized experience replay
Tom Schaul, John Quan, Ioannis Antonoglou, and David Silver · 2016
Cited alongside, same era.
Interactive learning from policy-dependent human feedback
James MacGlashan, Mark K Ho, Robert Loftin, Bei Peng, Guan Wang, David L Roberts, Matthew E Taylor, and Michael L Littman · 2017
Cited alongside, same era.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Cited alongside, same era.
Active reward learning from critiques
Yuchen Cui and Scott Niekum · 2018
Cited alongside, same era.
Deep tamer: Interactive agent shaping in high-dimensional state spaces
Garrett Warnell, Nicholas Waytowich, Vernon Lawhern, and Peter Stone · 2018
Cited alongside, same era.
Dqn-tamer: Human-in-the-loop reinforcement learning with intractable feedback
Riku Arakawa, Sosuke Kobayashi, Yuya Unno, Yuta Tsuboi, and Shin-ichi Maeda · 2018
Apple: Adaptive planner parameter learning from evaluative feedback
Zizhao Wang, Xuesu Xiao, Garrett Warnell, and Peter Stone · 2021
Later among the works it cites.
Hierarchical planning for long-horizon manipulation with geometric and symbolic scene graphs
Yifeng Zhu, Jonathan Tremblay, Stan Birchfield, and Yuke Zhu · 2021
Later among the works it cites.
Deep affordance foresight: Planning through what can be done in the future
Danfei Xu, Ajay Mandlekar, Roberto Martín-Martín, Yuke Zhu, Silvio Savarese, and Li Fei-Fei · 2021
Later among the works it cites.
Integrated task and motion planning
Caelan Reed Garrett, Rohan Chitnis, Rachel Holladay, Beomjoon Kim, Tom Silver, Leslie Pack Kaelbling, and Tomás Lozano-Pérez · 2021
Later among the works it cites.
Augmenting reinforcement learning with behavior primitives for diverse manipulation tasks
Soroush Nasiriany, Huihan Liu, and Yuke Zhu · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Cited alongside, same era.
Soft actor-critic algorithms and applications
Tuomas Haarnoja, Aurick Zhou, Kristian Hartikainen, George Tucker, Sehoon Ha, Jie Tan, Vikash Kumar, Henry Zhu, Abhishek Gupta, Pieter Abbeel, et al · 2018
Cited alongside, same era.
Leveraging human guidance for deep reinforcement learning tasks
Ruohan Zhang, Faraz Torabi, Lin Guan, Dana H Ballard, and Peter Stone · 2019
Cited alongside, same era.
Deep reinforcement learning from policy-dependent human feedback
Dilip Arumugam, Jun Ki Lee, Sophie Saskin, and Michael L Littman · 2019
Cited alongside, same era.
A review on interactive reinforcement learning from human social feedback
Jinying Lin, Zhen Ma, Randy Gomez, Keisuke Nakamura, Bo He, and Guangliang Li · 2020
Cited alongside, same era.
robosuite: A modular simulation framework and benchmark for robot learning
Yuke Zhu, Josiah Wong, Ajay Mandlekar, and Roberto Martín-Martín · 2020
Cited alongside, same era.
Rohan Chitnis, Tom Silver, Joshua B Tenenbaum, Tomas Lozano-Perez, and Leslie Pack Kaelbling · 2022
Later among the works it cites.
Cliport: What and where pathways for robotic manipulation
Mohit Shridhar, Lucas Manuelli, and Dieter Fox · 2022
Later among the works it cites.
Generalizable task planning through representation pretraining
Chen Wang, Danfei Xu, and Li Fei-Fei · 2022
Later among the works it cites.
Guided skill learning and abstraction for long-horizon manipulation
Shuo Cheng and Danfei Xu · 2022
Later among the works it cites.
A dual representation framework for robot learning with human guidance
Ruohan Zhang, Dhruva Bansal, Yilun Hao, Ayano Hiranaka, Jialu Gao, Chen Wang, Roberto Martín-Martín, Li Fei-Fei, and Jiajun Wu · 2023
Closest in time.
Perceiver-actor: A multi-task transformer for robotic manipulation
Mohit Shridhar, Lucas Manuelli, and Dieter Fox · 2023
Closest in time.
Stap: Sequencing task-agnostic policies
Christopher Agia, Toki Migimatsu, Jiajun Wu, and Jeannette Bohg · 2023
Closest in time.
Behavior-1k: A benchmark for embodied ai with 1,000 everyday activities and realistic simulation
Chengshu Li, Ruohan Zhang, Josiah Wong, Cem Gokmen, Sanjana Srivastava, Roberto Martín-Martín, Chen Wang, Gabrael Levine, Michael Lingelbach, Jiankai Sun, et al · 2023
Closest in time.
Viola: Imitation learning for vision-based manipulation with object proposal priors
Yifeng Zhu, Abhishek Joshi, Peter Stone, and Yuke Zhu · 2023
Closest in time.