Fetching the paper…
Reading the bibliography…
In this work, we study how to build a robotic system that can solve multiple 3D manipulation tasks given language instructions.
V-REP: A versatile and scalable robot simulation framework
Eric Rohmer, Surya P. N. Singh, and Marc Freese · 2013
Earlier work this paper cites.
Multi-view convolutional neural networks for 3D shape recognition
Hang Su, Subhransu Maji, Evangelos Kalogerakis, and Erik Learned-Miller · 2015
Earlier work this paper cites.
Multi-view 3D object detection network for autonomous driving
Xiaozhi Chen, Huimin Ma, Ji Wan, Bo Li, and Tian Xia · 2017
Earlier work this paper cites.
Making sense of vision and touch: Self-supervised learning of multimodal representations for contact-rich tasks
Michelle A. Lee, Yuke Zhu, Krishnan Srinivasan, Parth Shah, Silvio Savarese, Li Fei-Fei, Animesh Garg, and Jeannette Bohg · 2019
Earlier work this paper cites.
Transformer-based meta-imitation learning for robotic manipulation
Théo Cachet, Julien Perez, and Seungsu Kim · 2020
Earlier work this paper cites.
Transformers for one-shot visual imitation
Sudeep Dasari and Abhinav Gupta · 2020
Earlier work this paper cites.
Packit: A virtual environment for geometric planning
Ankit Goyal and Jia Deng · 2020
Earlier work this paper cites.
Imitation learning for high precision peg-in-hole tasks
Sagar Gubbi, Shishir Kolathaya, and Bharadwaj Amrutur · 2020
Earlier work this paper cites.
RLBench: The robot learning benchmark & learning environment
Stephen James, Zicong Ma, David Rovick Arrojo, and Andrew J. Davison · 2020
Earlier work this paper cites.
Understanding multi-modal perception using behavioral cloning for peg-in-a-hole insertion tasks
Yifang Liu, Diego Romeres, Devesh K. Jha, and Daniel Nikovski · 2020
Earlier work this paper cites.
Accelerating 3D deep learning with PyTorch3D
Nikhila Ravi, Jeremy Reizenstein, David Novotny, Taylor Gordon, Wan-Yen Lo, Justin Johnson, and Georgia Gkioxari · 2020
Earlier work this paper cites.
Deep reinforcement learning for industrial insertion tasks with visual inputs and natural rewards
Gerrit Schoettler, Ashvin Nair, Jianlan Luo, Shikhar Bahl, Juan Aparicio Ojea, Eugen Solowjow, and Sergey Levine · 2020
Earlier work this paper cites.
RAFT: Recurrent all-pairs field transforms for optical flow
Zachary Teed and Jia Deng · 2020
Earlier work this paper cites.
Transporter networks: Rearranging the visual world for robotic manipulation
Andy Zeng, Pete Florence, Jonathan Tompson, Stefan Welker, Jonathan Chien, Maria Attarian, Travis Armstrong, Ivan Krasin, Dan Duong, Vikas Sindhwani, and Johnny Lee · 2020
Earlier work this paper cites.
Differentiable spatial planning using transformers
Devendra Singh Chaplot, Deepak Pathak, and Jitendra Malik · 2021
Earlier work this paper cites.
Assistive tele-op: Leveraging transformers to collect robotic task demonstrations
Henry M. Clever, Ankur Handa, Hammad Mazhar, Kevin Parker, Omer Shapira, Qian Wan, Yashraj Narang, Iretiayo Akinola, Maya Cakmak, and Dieter Fox · 2021
Earlier work this paper cites.
Tactile-RL for insertion: Generalization to objects of unknown geometry
Siyuan Dong, Devesh K. Jha, Diego Romeres, Sangwoon Kim, Daniel Nikovski, and Alberto Rodriguez · 2021
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Earlier work this paper cites.
Revisiting point cloud shape classification with a simple and effective baseline
Ankit Goyal, Hei Law, Bowei Liu, Alejandro Newell, and Jia Deng · 2021
Cited alongside, same era.
MVTN: Multi-view transformation network for 3D shape recognition
Abdullah Hamdi, Silvio Giancola, and Bernard Ghanem · 2021
Cited alongside, same era.
BC-Z: Zero-shot task generalization with robotic imitation learning
Eric Jang, Alex Irpan, Mohi Khansari, Daniel Kappler, Frederik Ebert, Corey Lynch, Sergey Levine, and Chelsea Finn · 2021
Cited alongside, same era.
Motion Planning Transformers: A motion planning framework for mobile robots
Jacob J. Johnson, Uday S. Kalra, Ankit Bhatia, Linjun Li, Ahmed H. Qureshi, and Michael C. Yip · 2021
Cited alongside, same era.
Transformer-based deep imitation learning for dual-arm robot manipulation
Heecheol Kim, Yoshiyuki Ohmura, and Yasuo Kuniyoshi · 2021
Cited alongside, same era.
Learning vision-guided quadrupedal locomotion end-to-end with cross-modal transformers
Ruihan Yang, Minghao Zhang, Nicklas Hansen, Huazhe Xu, and Xiaolong Wang · 2022
Later among the works it cites.
RT-2: Vision-language-action models transfer web knowledge to robotic control
Anthony Brohan, Noah Brown, Justice Carbajal, Yevgen Chebotar, Xi Chen, Krzysztof Choromanski, Tianli Ding, Danny Driess, Avinava Dubey, Chelsea Finn, Pete Florence, Chuyuan Fu, Montse Gonzalez Arenas, Keerthana Gopalakrishnan, Kehang Han, Karol Hausman, Alexander Herzog, Jasmine Hsu, Brian Ichter, Alex Irpan, Nikhil Joshi, Ryan Julian, Dmitry Kalashnikov, Yuheng Kuang, Isabel Leal, Lisa Lee, Tsang-Wei Edward Lee, Sergey Levine, Yao Lu, Henryk Michalewski, Igor Mordatch, Karl Pertsch, Kanishka Rao, Krista Reymann, Michael Ryoo, Grecia Salazar, Pannag Sanketi, Pierre Sermanet, Jaspiar Singh, Anikait Singh, Radu Soricut, Huong Tran, Vincent Vanhoucke, Quan Vuong, Ayzaan Wahid, Stefan Welker, Paul Wohlhart, Jialin Wu, Fei Xia, Ted Xiao, Peng Xu, Sichun Xu, Tianhe Yu, and Brianna Zitkovich · 2023
Later among the works it cites.
PolarNet: 3D point clouds for language-guided robotic manipulation
Shizhe Chen, Ricardo Garcia-Pinel, Cordelia Schmid, and Ivan Laptev · 2023
Later among the works it cites.
Diffusion Policy: Visuomotor Policy Learning via Action Diffusion
Cheng Chi, Siyuan Feng, Yilun Du, Zhenjia Xu, Eric Cousineau, Benjamin CM Burchfiel, and Shuran Song · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Rendering point clouds with compute shaders and vertex order optimization
Markus Schütz, Bernhard Kerbl, and Michael Wimmer · 2021
Cited alongside, same era.
RT-1: Robotics transformer for real-world control at scale
Anthony Brohan, Noah Brown, Justice Carbajal, Yevgen Chebotar, Joseph Dabis, Chelsea Finn, Keerthana Gopalakrishnan, Karol Hausman, Alex Herzog, Jasmine Hsu, Julian Ibarz, Brian Ichter, Alex Irpan, Tomas Jackson, Sally Jesmonth, Nikhil J Joshi, Ryan Julian, Dmitry Kalashnikov, Yuheng Kuang, Isabel Leal, Kuang-Huei Lee, Sergey Levine, Yao Lu, Utsav Malla, Deeksha Manjunath, Igor Mordatch, Ofir Nachum, Carolina Parada, Jodilyn Peralta, Emily Perez, Karl Pertsch, Jornell Quiambao, Kanishka Rao, Michael Ryoo, Grecia Salazar, Pannag Sanketi, Kevin Sayed, Jaspiar Singh, Sumedh Sontakke, Austin Stone, Clayton Tan, Huong Tran, Vincent Vanhoucke, Steve Vega, Quan Vuong, Fei Xia, Ted Xiao, Peng Xu, Sichun Xu, Tianhe Yu, and Brianna Zitkovich · 2022
Cited alongside, same era.
8-bit optimizers via block-wise quantization
Tim Dettmers, Mike Lewis, Sam Shleifer, and Luke Zettlemoyer · 2022
Cited alongside, same era.
IFOR: Iterative flow minimization for robotic object rearrangement
Ankit Goyal, Arsalan Mousavian, Chris Paxton, Yu-Wei Chao, Brian Okorn, Jia Deng, and Dieter Fox · 2022
Cited alongside, same era.
Instruction-driven history-aware policies for robotic manipulations
Pierre-Louis Guhur, Shizhe Chen, Ricardo Garcia Pinel, Makarand Tapaswi, Ivan Laptev, and Cordelia Schmid · 2022
Cited alongside, same era.
Multi-view transformer for 3D visual grounding
Shijia Huang, Yilun Chen, Jiaya Jia, and Liwei Wang · 2022
Cited alongside, same era.
Coarse-to-fine Q-attention: Efficient learning for visual robotic manipulation via discretisation
Stephen James, Kentaro Wada, Tristan Laidlow, and Andrew J. Davison · 2022
Cited alongside, same era.
Act3D: 3D feature field transformers for multi-task robotic manipulation
Theophile Gervet, Zhou Xian, Nikolaos Gkanatsios, and Katerina Fragkiadaki · 2023
Later among the works it cites.
RVT: Robotic view transformer for 3D object manipulation
Ankit Goyal, Jie Xu, Yijie Guo, Valts Blukis, Yu-Wei Chao, and Dieter Fox · 2023
Later among the works it cites.
Voint Cloud: Multi-view point cloud representation for 3D understanding
Abdullah Hamdi, Silvio Giancola, and Bernard Ghanem · 2023
Later among the works it cites.
Octo: An open-source generalist robot policy
Octo Model Team, Dibya Ghosh, Homer Walke, Karl Pertsch, Kevin Black, Oier Mees, Sudeep Dasari, Joey Hejna, Charles Xu, Jianlan Luo, Tobias Kreiman, You Liang Tan, Dorsa Sadigh, Chelsea Finn, and Sergey Levine · 2023
Later among the works it cites.
Waypoint-based imitation learning for robotic manipulation
Lucy Xiaoyang Shi, Archit Sharma, Tony Z. Zhao, and Chelsea Finn · 2023
Later among the works it cites.
Shelving, stacking, hanging: Relational pose diffusion for multi-modal rearrangement
Anthony Simeonov, Ankit Goyal, Lucas Manuelli, Lin Yen-Chen, Alina Sarmiento, Alberto Rodriguez, Pulkit Agrawal, and Dieter Fox · 2023
Later among the works it cites.
KITE: Keypoint-conditioned policies for semantic manipulation
Priya Sundaresan, Suneel Belkhale, Dorsa Sadigh, and Jeannette Bohg · 2023
Later among the works it cites.
IndustReal: Transferring Contact-Rich Assembly Tasks from Simulation to Reality
Bingjie Tang, Michael A. Lin, Iretiayo A. Akinola, Ankur Handa, Gaurav S. Sukhatme, Fabio Ramos, Dieter Fox, and Yashraj Narang · 2023
Later among the works it cites.
Mimicplay: Long-horizon imitation learning by watching human play
Chen Wang, Linxi Fan, Jiankai Sun, Ruohan Zhang, Li Fei-Fei, Danfei Xu, Yuke Zhu, and Anima Anandkumar · 2023
Later among the works it cites.
ChainedDiffuser: Unifying trajectory diffusion and keypose prediction for robotic manipulation
Zhou Xian, Nikolaos Gkanatsios, Theophile Gervet, Tsung-Wei Ke, and Katerina Fragkiadaki · 2023
Later among the works it cites.
M2T2: Multi-task masked transformer for object-centric pick and place
Wentao Yuan, Adithyavairavan Murali, Arsalan Mousavian, and Dieter Fox · 2023
Later among the works it cites.
Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware
Tony Z. Zhao, Vikash Kumar, Sergey Levine, and Chelsea Finn · 2023
Later among the works it cites.
Fourier Transporter: Bi-equivariant robotic manipulation in 3D
Haojie Huang, Owen Howell, Xupeng Zhu, Dian Wang, Robin Walters, and Robert Platt · 2024
Closest in time.