Fetching the paper…
Reading the bibliography…
Learning in simulation and transferring the learned policy to the real world has the potential to enable generalist robots.
A unified approach for motion and force control of robot manipulators: The operational space formulation
O. Khatib · 1987
Earlier work this paper cites.
Alvinn: An autonomous land vehicle in a neural network
D. A. Pomerleau · 1988
Earlier work this paper cites.
Noise and the reality gap: The use of simulation in evolutionary robotics
N. Jakobi, P. Husbands, and I. Harvey · 1995
Earlier work this paper cites.
Biped dynamic walking using reinforcement learning
H. Benbrahim and J. A. Franklin · 1997
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
System Identification , pages 163–173
L. Ljung · 1998
Earlier work this paper cites.
Characterizing efficiency of human robot interaction: a case study of shared-control teleoperation
J. Crandall and M. Goodrich · 2002
Earlier work this paper cites.
Interactive machine learning
J. A. Fails and D. R. Olsen · 2003
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
P. Abbeel and A. Y. Ng · 2004
Earlier work this paper cites.
Operational space control: A theoretical and empirical comparison
J. Nakanishi, R. Cory, M. Mistry, J. Peters, and S. Schaal · 2008
Earlier work this paper cites.
Search-based structured prediction
H. D. III, J. Langford, and D. Marcu · 2009
Earlier work this paper cites.
Crossing the reality gap in evolutionary robotics by promoting transferable controllers
S. Koos, J.-B. Mouret, and S. Doncieux · 2010
Earlier work this paper cites.
Tactile guidance for policy refinement and reuse
B. D. Argall, E. L. Sauser, and A. G. Billard · 2010
Earlier work this paper cites.
Operational space control of constrained and underactuated systems
M. N. Mistry and L. Righetti · 2011
Earlier work this paper cites.
The k-armed dueling bandits problem
Y. Yue, J. Broder, R. Kleinberg, and T. Joachims · 2011
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
E. Todorov, T. Erez, and Y. Tassa · 2012
Earlier work this paper cites.
Formalizing assistive teleoperation , volume 376
A. D. Dragan and S. S. Srinivasa · 2012
Earlier work this paper cites.
Leveraging multiple simulators for crossing the reality gap
A. Boeing and T. Bräunl · 2012
Earlier work this paper cites.
Reinforcement learning from human reward: Discounting in episodic tasks
W. B. Knox and P. Stone · 2012
Earlier work this paper cites.
A policy-blending formalism for shared control
A. D. Dragan and S. S. Srinivasa · 2013
Earlier work this paper cites.
Learning trajectory preferences for manipulators via iterative improvement
A. Jain, B. Wojcik, T. Joachims, and A. Saxena · 2013
Earlier work this paper cites.
Power to the people: The role of humans in interactive machine learning
S. Amershi, M. Cakmak, W. B. Knox, and T. Kulesza · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Earlier work this paper cites.
Towards adapting deep visuomotor representations from simulated to real environments
E. Tzeng, C. Devin, J. Hoffman, C. Finn, X. Peng, S. Levine, K. Saenko, and T. Darrell · 2015
Earlier work this paper cites.
Shared autonomy via hindsight optimization
S. Javdani, S. S. Srinivasa, and J. A. Bagnell · 2015
Earlier work this paper cites.
Policy distillation
A. A. Rusu, S. G. Colmenarejo, C. Gulcehre, G. Desjardins, J. Kirkpatrick, R. Pascanu, V. Mnih, K. Kavukcuoglu, and R. Hadsell · 2015
Earlier work this paper cites.
Fast and accurate deep network learning by exponential linear units (elus)
D.-A. Clevert, T. Unterthiner, and S. Hochreiter · 2015
Earlier work this paper cites.
High-dimensional continuous control using generalized advantage estimation
J. Schulman, P. Moritz, S. Levine, M. Jordan, and P. Abbeel · 2015
Earlier work this paper cites.
Pointnet: Deep learning on point sets for 3d classification and segmentation
C. R. Qi, H. Su, K. Mo, and L. J. Guibas · 2016
Earlier work this paper cites.
Human-in-the-loop optimization of shared autonomy in assistive robotics
D. Gopinath, S. Jain, and B. D. Argall · 2016
Earlier work this paper cites.
Gaussian error linear units (gelus)
D. Hendrycks and K. Gimpel · 2016
Earlier work this paper cites.
Sgdr: Stochastic gradient descent with warm restarts
I. Loshchilov and F. Hutter · 2016
Earlier work this paper cites.
Query-efficient imitation learning for end-to-end autonomous driving
J. Zhang and K. Cho · 2016
Earlier work this paper cites.
Using simulation and domain adaptation to improve efficiency of deep robotic grasping
K. Bousmalis, A. Irpan, P. Wohlhart, Y. Bai, M. Kelcey, M. Kalakrishnan, L. Downs, J. Ibarz, P. Pastor, K. Konolige, S. Levine, and V. Vanhoucke · 2017
Earlier work this paper cites.
Sim-to-real transfer of robotic control with dynamics randomization
X. B. Peng, M. Andrychowicz, W. Zaremba, and P. Abbeel · 2017
Earlier work this paper cites.
Grounded action transformation for robot learning in simulation
J. P. Hanna and P. Stone · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Earlier work this paper cites.
Model-agnostic meta-learning for fast adaptation of deep networks
C. Finn, P. Abbeel, and S. Levine · 2017
Earlier work this paper cites.
Pointnet++: Deep hierarchical feature learning on point sets in a metric space
C. R. Qi, L. Yi, H. Su, and L. J. Guibas · 2017
Earlier work this paper cites.
Deep reinforcement learning from human preferences
P. Christiano, J. Leike, T. B. Brown, M. Martic, S. Legg, and D. Amodei · 2017
Earlier work this paper cites.
Learning robot objectives from physical human interaction
A. V. Bajcsy, D. P. Losey, M. K. O’Malley, and A. D. Dragan · 2017
Earlier work this paper cites.
Trial without error: Towards safe reinforcement learning via human intervention
W. Saunders, G. Sastry, A. Stuhlmueller, and O. Evans · 2017
Earlier work this paper cites.
Sim-to-real: Learning agile locomotion for quadruped robots
J. Tan, T. Zhang, E. Coumans, A. Iscen, Y. Bai, D. Hafner, S. Bohez, and V. Vanhoucke · 2018
Earlier work this paper cites.
Closing the sim-to-real loop: Adapting simulation randomization with real world experience
Y. Chebotar, A. Handa, V. Makoviychuk, M. Macklin, J. Issac, N. Ratliff, and D. Fox · 2018
Earlier work this paper cites.
Feedback control for cassie with deep reinforcement learning
Z. Xie, G. Berseth, P. Clary, J. Hurst, and M. van de Panne · 2018
Earlier work this paper cites.
Shared autonomy via deep reinforcement learning
S. Reddy, A. D. Dragan, and S. Levine · 2018
Earlier work this paper cites.
Hg-dagger: Interactive imitation learning with human experts
M. Kelly, C. Sidrane, K. Driggs-Campbell, and M. J. Kochenderfer · 2018
Earlier work this paper cites.
Reinforcement and imitation learning for diverse visuomotor skills
Y. Zhu, Z. Wang, J. Merel, A. Rusu, T. Erez, S. Cabi, S. Tunyasuvunakool, J. Kramár, R. Hadsell, N. de Freitas, and N. Heess · 2018
Earlier work this paper cites.
Residual reinforcement learning for robot control
T. Johannink, S. Bahl, A. Nair, J. Luo, A. Kumar, M. Loskyll, J. A. Ojea, E. Solowjow, and S. Levine · 2018
Earlier work this paper cites.
Residual policy learning
T. Silver, K. Allen, J. Tenenbaum, and L. Kaelbling · 2018
Earlier work this paper cites.
Set transformer: A framework for attention-based permutation-invariant neural networks
J. Lee, Y. Lee, J. Kim, A. R. Kosiorek, S. Choi, and Y. W. Teh · 2018
Earlier work this paper cites.
Maximum a posteriori policy optimisation
A. Abdolmaleki, J. T. Springenberg, Y. Tassa, R. Munos, N. Heess, and M. Riedmiller · 2018
Earlier work this paper cites.
An interactive framework for learning continuous actions policies based on corrective feedback
C. Celemin and J. R. del Solar · 2018
Earlier work this paper cites.
Solving rubik’s cube with a robot hand
OpenAI, I. Akkaya, M. Andrychowicz, M. Chociej, M. Litwin, B. McGrew, A. Petron, A. Paino, M. Plappert, G. Powell, R. Ribas, J. Schneider, N. Tezak, J. Tworek, P. Welinder, L. Weng, Q. Yuan, W. Zaremba, and L. Zhang · 2019
Cited alongside, same era.
Variable impedance control in end-effector space: An action space for reinforcement learning in contact-rich tasks
R. Martín-Martín, M. A. Lee, R. Gardner, S. Savarese, J. Bohg, and A. Garg · 2019
Cited alongside, same era.
Closing the sim-to-real loop: Adapting simulation randomization with real world experience
Y. Chebotar, A. Handa, V. Makoviychuk, M. Macklin, J. Issac, N. D. Ratliff, and D. Fox · 2019
Cited alongside, same era.
Leveraging human guidance for deep reinforcement learning tasks
R. Zhang, F. Torabi, L. Guan, D. H. Ballard, and P. Stone · 2019
Cited alongside, same era.
Tossingbot: Learning to throw arbitrary objects with residual physics
A. Zeng, S. Song, J. Lee, A. Rodriguez, and T. Funkhouser · 2019
Cited alongside, same era.
Reinforcement learning of impedance policies for peg-in-hole tasks: Role of asymmetric matrices
S. Kozlovsky, E. Newman, and M. Zacksenhouse · 2022
Later among the works it cites.
Factory: Fast contact for robotic assembly
Y. S. Narang, K. Storey, I. Akinola, M. Macklin, P. Reist, L. Wawrzyniak, Y. Guo, Á. Moravánszky, G. State, M. Lu, A. Handa, and D. Fox · 2022
Later among the works it cites.
Dexpoint: Generalizable point cloud reinforcement learning for sim-to-real dexterous manipulation
Y. Qin, B. Huang, Z.-H. Yin, H. Su, and X. Wang · 2022
Later among the works it cites.
Visual dexterity: In-hand dexterous manipulation from depth
T. Chen, M. Tippur, S. Wu, V. Kumar, E. Adelson, and P. Agrawal · 2022
Later among the works it cites.
Robot learning on the job: Human-in-the-loop autonomy and learning during deployment
H. Liu, S. Nasiriany, L. Zhang, Z. Bao, and Y. Zhu · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
On the similarities and differences among contact models in robot simulation
P. C. Horak and J. C. Trinkle · 2019
Cited alongside, same era.
Drake: Model-based design and verification for robotics, 2019
R. Tedrake and the Drake Development Team · 2019
Cited alongside, same era.
Learning to manipulate deformable objects without demonstrations
Y. Wu, W. Yan, T. Kurutach, L. Pinto, and P. Abbeel · 2019
Cited alongside, same era.
Learning agile and dynamic motor skills for legged robots
J. Hwangbo, J. Lee, A. Dosovitskiy, D. Bellicoso, V. Tsounis, V. Koltun, and M. Hutter · 2019
Cited alongside, same era.
Reinforcement learning on variable impedance controller for high-precision robotic assembly
J. Luo, E. Solowjow, C. Wen, J. A. Ojea, A. M. Agogino, A. Tamar, and P. Abbeel · 2019
Cited alongside, same era.
Human-guided trajectory adaptation for tool transfer
T. Fitzgerald, E. Short, A. Goel, and A. Thomaz · 2019
Cited alongside, same era.
Interactively shaping robot behaviour with unlabeled human instructions
A. Najar, O. Sigaud, and M. Chetouani · 2019
Cited alongside, same era.
Y. Zhu, A. Joshi, P. Stone, and Y. Zhu · 2022
Later among the works it cites.
Motion policy networks
A. Fishman, A. Murali, C. Eppner, B. Peele, B. Boots, and D. Fox · 2022
Later among the works it cites.
Vima: General robot manipulation with multimodal prompts
Y. Jiang, A. Gupta, Z. Zhang, G. Wang, Y. Dou, Y. Chen, L. Fei-Fei, A. Anandkumar, Y. Zhu, and L. Fan · 2022
Later among the works it cites.
Perceiver-actor: A multi-task transformer for robotic manipulation
M. Shridhar, L. Manuelli, and D. Fox · 2022
Later among the works it cites.
Multi-skill mobile manipulation for object rearrangement
J. Gu, D. S. Chaplot, H. Su, and J. Malik · 2022
Later among the works it cites.
Toolflownet: Robotic manipulation with tools via predicting tool flow from point clouds
D. Seita, Y. Wang, S. J. Shetty, E. Y. Li, Z. Erickson, and D. Held · 2022
Later among the works it cites.
Diffskill: Skill abstraction from differentiable physics for deformable object manipulations with tools
X. Lin, Z. Huang, Y. Li, J. B. Tenenbaum, D. Held, and C. Gan · 2022
Later among the works it cites.
Hierarchical reinforcement learning for precise soccer shooting skills using a quadrupedal robot
Y. Ji, Z. Li, Y. Sun, X. B. Peng, S. Levine, G. Berseth, and K. Sreenath · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. L. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, J. Schulman, J. Hilton, F. Kelton, L. Miller, M. Simens, A. Askell, P. Welinder, P. Christiano, J. Leike, and R. Lowe · 2022
Later among the works it cites.
Interactive imitation learning in robotics: A survey
C. Celemin, R. Pérez-Dattari, E. Chisari, G. Franzese, L. de Souza Rosa, R. Prakash, Z. Ajanović, M. Ferraz, A. Valada, and J. Kober · 2022
Later among the works it cites.
Foundation models for decision making: Problems, methods, and opportunities
S. Yang, O. Nachum, Y. Du, J. Wei, P. Abbeel, and D. Schuurmans · 2023
Later among the works it cites.
Robocat: A self-improving generalist agent for robotic manipulation
K. Bousmalis, G. Vezzani, D. Rao, C. Devin, A. X. Lee, M. Bauza, T. Davchev, Y. Zhou, A. Gupta, A. Raju, A. Laurens, C. Fantacci, V. Dalibard, M. Zambelli, M. Martins, R. Pevceviciute, M. Blokzijl, M. Denil, N. Batchelor, T. Lampe, E. Parisotto, K. Żołna, S. Reed, S. G. Colmenarejo, J. Scholz, A. Abdolmaleki, O. Groth, J.-B. Regli, O. Sushkov, T. Rothörl, J. E. Chen, Y. Aytar, D. Barker, J. Ortiz, M. Riedmiller, J. T. Springenberg, R. Hadsell, F. Nori, and N. Heess · 2023
Later among the works it cites.
General in-hand object rotation with vision and touch
H. Qi, B. Yi, S. Suresh, M. Lambeta, Y. Ma, R. Calandra, and J. Malik · 2023
Later among the works it cites.
Robot parkour learning
Z. Zhuang, Z. Fu, J. Wang, C. Atkeson, S. Schwertfeger, C. Finn, and H. Zhao · 2023
Later among the works it cites.
Neural volumetric memory for visual locomotion control
R. Yang, G. Yang, and X. Wang · 2023
Later among the works it cites.
Real-world humanoid locomotion with reinforcement learning
I. Radosavovic, T. Xiao, B. Zhang, T. Darrell, J. Malik, and K. Sreenath · 2023
Later among the works it cites.
Champion-level drone racing using deep reinforcement learning
E. Kaufmann, L. Bauersfeld, A. Loquercio, M. Müller, V. Koltun, and D. Scaramuzza · 2023
Later among the works it cites.
Reaching the limit in autonomous racing: Optimal control versus reinforcement learning
Y. Song, A. Romero, M. Müller, V. Koltun, and D. Scaramuzza · 2023
Later among the works it cites.
Pre- and post-contact policy decomposition for non-prehensile manipulation with zero-shot sim-to-real transfer
M. Kim, J. Han, J. Kim, and B. Kim · 2023
Later among the works it cites.
Learning generalizable pivoting skills
X. Zhang, S. Jain, B. Huang, M. Tomizuka, and D. Romeres · 2023
Later among the works it cites.
Industreal: Transferring contact-rich assembly tasks from simulation to reality
B. Tang, M. A. Lin, I. Akinola, A. Handa, G. S. Sukhatme, F. Ramos, D. Fox, and Y. S. Narang · 2023
Later among the works it cites.
On the role of the action space in robot manipulation learning and sim-to-real transfer
E. Aljalbout, F. Frank, M. Karl, and P. van der Smagt · 2023
Later among the works it cites.
Dextreme: Transfer of agile in-hand manipulation from simulation to reality
A. Handa, A. Allshire, V. Makoviychuk, A. Petrenko, R. Singh, J. Liu, D. Makoviichuk, K. V. Wyk, A. Zhurkevich, B. Sundaralingam, and Y. S. Narang · 2023
Later among the works it cites.
Cherry-Picking with Reinforcement Learning
Y. Zhang, L. Ke, A. Deshpande, A. Gupta, and S. Srinivasa · 2023
Later among the works it cites.
Decomposing the generalization gap in imitation learning for visual robotic manipulation
A. Xie, L. Lee, T. Xiao, and C. Finn · 2023
Later among the works it cites.
Furniturebench: Reproducible real-world benchmark for long-horizon complex manipulation
M. Heo, Y. Lee, D. Lee, and J. J. Lim · 2023
Later among the works it cites.
Interventional data generation for robust and data-efficient robot imitation learning
R. Hoque, A. Mandlekar, C. R. Garrett, K. Goldberg, and D. Fox · 2023
Later among the works it cites.
”no, to the right” – online language corrections for robotic manipulation via shared autonomy
Y. Cui, S. Karamcheti, R. Palleti, N. Shivakumar, P. Liang, and D. Sadigh · 2023
Later among the works it cites.
Gello: A general, low-cost, and intuitive teleoperation framework for robot manipulators
P. Wu, Y. Shentu, Z. Yi, X. Lin, and P. Abbeel · 2023
Later among the works it cites.
Mimicgen: A data generation system for scalable robot learning using human demonstrations
A. Mandlekar, S. Nasiriany, B. Wen, I. Akinola, Y. Narang, L. Fan, Y. Zhu, and D. Fox · 2023
Later among the works it cites.
Diffusion policy: Visuomotor policy learning via action diffusion
C. Chi, S. Feng, Y. Du, Z. Xu, E. Cousineau, B. Burchfiel, and S. Song · 2023
Later among the works it cites.
Orbit: A unified simulation framework for interactive robot learning environments
M. Mittal, C. Yu, Q. Yu, J. Liu, N. Rudin, D. Hoeller, J. L. Yuan, P. P. Tehrani, R. Singh, Y. Guo, H. Mazhar, A. Mandlekar, B. Babich, G. State, M. Hutter, and A. Garg · 2023
Later among the works it cites.
Homerobot: Open-vocabulary mobile manipulation
S. Yenamandra, A. Ramachandran, K. Yadav, A. Wang, M. Khanna, T. Gervet, T.-Y. Yang, V. Jain, A. W. Clegg, J. Turner, Z. Kira, M. Savva, A. Chang, D. S. Chaplot, D. Batra, R. Mottaghi, Y. Bisk, and C. Paxton · 2023
Later among the works it cites.
Imitating shortest paths in simulation enables effective navigation and manipulation in the real world
K. Ehsani, T. Gupta, R. Hendrix, J. Salvador, L. Weihs, K.-H. Zeng, K. P. Singh, Y. Kim, W. Han, A. Herrasti, R. Krishna, D. Schwenk, E. VanderBilt, and A. Kembhavi · 2023
Later among the works it cites.
Eureka: Human-level reward design via coding large language models
Y. J. Ma, W. Liang, G. Wang, D.-A. Huang, O. Bastani, D. Jayaraman, Y. Zhu, L. Fan, and A. Anandkumar · 2023
Later among the works it cites.
Active reward learning from online preferences
V. Myers, E. Bıyık, and D. Sadigh · 2023
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model
R. Rafailov, A. Sharma, E. Mitchell, S. Ermon, C. D. Manning, and C. Finn · 2023
Later among the works it cites.
Contrastive preference learning: Learning from human feedback without rl
J. Hejna, R. Rafailov, H. Sikchi, C. Finn, S. Niekum, W. B. Knox, and D. Sadigh · 2023
Later among the works it cites.
Learning from active human involvement through proxy value propagation
Z. Peng, W. Mo, C. Duan, Q. Li, and B. Zhou · 2023
Later among the works it cites.
Reinforcement learning for versatile, dynamic, and robust bipedal locomotion control
Z. Li, X. B. Peng, P. Abbeel, S. Levine, G. Berseth, and K. Sreenath · 2024
Closest in time.
Cyberdemo: Augmenting simulated human demonstration for real-world dexterous manipulation
J. Wang, Y. Qin, K. Kuang, Y. Korkmaz, A. Gurumoorthy, H. Su, and X. Wang · 2024
Closest in time.
ACDC: Automated creation of digital cousins for robust policy learning
T. Dai, J. Wong, Y. Jiang, C. Wang, C. Gokmen, R. Zhang, J. Wu, and L. Fei-Fei · 2024
Closest in time.
Dexcap: Scalable and portable mocap data collection system for dexterous manipulation
C. Wang, H. Shi, W. Wang, R. Zhang, L. Fei-Fei, and C. K. Liu · 2024
Closest in time.
Mobile aloha: Learning bimanual mobile manipulation with low-cost whole-body teleoperation
Z. Fu, T. Z. Zhao, and C. Finn · 2024
Closest in time.
Telemoma: A modular and versatile teleoperation system for mobile manipulation
S. Dass, W. Ai, Y. Jiang, S. Singh, J. Hu, R. Zhang, P. Stone, B. Abbatematteo, and R. Martín-Martín · 2024
Closest in time.
Learning human-to-humanoid real-time whole-body teleoperation
T. He, Z. Luo, W. Xiao, C. Zhang, K. Kitani, C. Liu, and G. Shi · 2024
Closest in time.
Juicer: Data-efficient imitation learning for robotic assembly
L. Ankile, A. Simeonov, I. Shenfeld, and P. Agrawal · 2024
Closest in time.