Fetching the paper…
Reading the bibliography…
The ability for robots to comprehend and execute manipulation tasks based on natural language instructions is a long-term goal in robotics.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Contact-based language for robotic manipulation planning
A. P. Shah, G. A. D. Lopes, and E. Najafi · 2016
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. u. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Pointnet: Deep learning on point sets for 3d classification and segmentation
C. R. Qi, H. Su, K. Mo, and L. J. Guibas · 2017
Earlier work this paper cites.
Compositional human pose regression
X. Sun, J. Shang, S. Liang, and Y. Wei · 2017
Earlier work this paper cites.
Voxelnet: End-to-end learning for point cloud based 3d object detection
Y. Zhou and O. Tuzel · 2018
Earlier work this paper cites.
Interactive visual grounding of referring expressions for human-robot interaction
M. Shridhar and D. Hsu · 2018
Earlier work this paper cites.
Pointnet++: Deep hierarchical feature learning on point sets in a metric space
C. R. Qi, L. Yi, H. Su, and L. J. Guibas · 2018
Earlier work this paper cites.
Open3D: A modern library for 3D data processing
Q.-Y. Zhou, J. Park, and V. Koltun · 2018
Earlier work this paper cites.
Integral human pose regression
X. Sun, B. Xiao, F. Wei, S. Liang, and Y. Wei · 2018
Earlier work this paper cites.
BERT: pre-training of deep bidirectional transformers for language understanding
J. Devlin, M. Chang, K. Lee, and K. Toutanova · 2019
Earlier work this paper cites.
Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning
T. Yu, D. Quillen, Z. He, R. Julian, K. Hausman, C. Finn, and S. Levine · 2019
Earlier work this paper cites.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
J. Lu, D. Batra, D. Parikh, and S. Lee · 2019
Earlier work this paper cites.
Scanrefer: 3d object localization in rgb-d scans using natural language
D. Z. Chen, A. X. Chang, and M. Nießner · 2020
Earlier work this paper cites.
RLBench: The robot learning benchmark & learning environment
S. James, Z. Ma, D. R. Arrojo, and A. J. Davison · 2020
Earlier work this paper cites.
Concept2robot: Learning manipulation concepts from instructions and human demonstrations
L. Shao, T. Migimatsu, Q. Zhang, K. Yang, and J. Bohg · 2020
Earlier work this paper cites.
Language-conditioned imitation learning for robot manipulation tasks
S. Stepputtis, J. Campbell, M. Phielipp, S. Lee, C. Baral, and H. Ben Amor · 2020
Earlier work this paper cites.
Grasping in the wild: Learning 6dof closed-loop grasping from low-cost demonstrations
S. Song, A. Zeng, J. Lee, and T. Funkhouser · 2020
Earlier work this paper cites.
Learning obstacle representations for neural motion planning
R. Strudel, R. Garcia, J. Carpentier, J. Laumond, I. Laptev, and C. Schmid · 2020
Earlier work this paper cites.
BC-z: Zero-shot task generalization with robotic imitation learning
E. Jang, A. Irpan, M. Khansari, D. Kappler, F. Ebert, C. Lynch, S. Levine, and C. Finn · 2021
Cited alongside, same era.
Cliport: What and where pathways for robotic manipulation
M. Shridhar, L. Manuelli, and D. Fox · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Cited alongside, same era.
Language conditioned imitation learning over unstructured data
C. Lynch and P. Sermanet · 2021
Cited alongside, same era.
Perceiver: General perception with iterative attention, 2021
A. Jaegle, F. Gimeno, A. Brock, A. Zisserman, O. Vinyals, and J. Carreira · 2021
Cited alongside, same era.
Learning visible connectivity dynamics for cloth smoothing
X. Lin, Y. Wang, Z. Huang, and D. Held · 2021
Interactive language: Talking to robots in real time
C. Lynch, A. Wahid, J. Tompson, T. Ding, J. Betker, R. Baruch, T. Armstrong, and P. Florence · 2022
Later among the works it cites.
VIOLA: Object-centric imitation learning for vision-based robot manipulation
Y. Zhu, A. Joshi, P. Stone, and Y. Zhu · 2022
Later among the works it cites.
BEHAVIOR-1k: A benchmark for embodied AI with 1,000 everyday activities and realistic simulation
C. Li, R. Zhang, J. Wong, C. Gokmen, S. Srivastava, R. Martín-Martín, C. Wang, G. Levine, M. Lingelbach, J. Sun, M. Anvari, M. Hwang, M. Sharma, A. Aydin, D. Bansal, S. Hunter, K.-Y. Kim, A. Lou, C. R. Matthews, I. Villa-Renteria, J. H. Tang, C. Tang, F. Xia, S. Savarese, H. Gweon, K. Liu, J. Wu, and L. Fei-Fei · 2022
Later among the works it cites.
VLMbench: A compositional benchmark for vision-and-language manipulation
K. Zheng, X. Chen, O. Jenkins, and X. E. Wang · 2022
Later among the works it cites.
Vima: General robot manipulation with multimodal prompts
Y. Jiang, A. Gupta, Z. Zhang, G. Wang, Y. Dou, Y. Chen, L. Fei-Fei, A. Anandkumar, Y. Zhu, and L. Fan · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Contact-graspnet: Efficient 6-dof grasp generation in cluttered scenes
M. Sundermeyer, A. Mousavian, R. Triebel, and D. Fox · 2021
Cited alongside, same era.
Point transformer
H. Zhao, L. Jiang, J. Jia, P. H. Torr, and V. Koltun · 2021
Cited alongside, same era.
Pct: Point cloud transformer
M.-H. Guo, J.-X. Cai, Z.-N. Liu, T.-J. Mu, R. R. Martin, and S.-M. Hu · 2021
Cited alongside, same era.
Auto-Lambda: Disentangling dynamic task relationships
S. Liu, S. James, A. J. Davison, and E. Johns · 2022
Cited alongside, same era.
Instruction-driven history-aware policies for robotic manipulations
P.-L. Guhur, S. Chen, R. Garcia-Pinel, M. Tapaswi, I. Laptev, and C. Schmid · 2022
Cited alongside, same era.
Instruction-following agents with jointly pre-trained vision-language models
H. Liu, L. Lee, K. Lee, and P. Abbeel · 2022
Cited alongside, same era.
Rt-1: Robotics transformer for real-world control at scale
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, J. Hsu, J. Ibarz, B. Ichter, A. Irpan, T. Jackson, S. Jesmonth, N. Joshi, R. Julian, D. Kalashnikov, Y. Kuang, I. Leal, K.-H. Lee, S. Levine, Y. Lu, U. Malla, D. Manjunath, I. Mordatch, O. Nachum, C. Parada, J. Peralta, E. Perez, K. Pertsch, J. Quiambao, K. Rao, M. Ryoo, G. Salazar, P. Sanketi, K. Sayed, J. Singh, S. Sontakke, A. Stone, C. Tan, H. Tran, V. Vanhoucke, S. Vega, Q. Vuong, F. Xia, T. Xiao, P. Xu, S. Xu, T. Yu, and B. Zitkovich · 2022
Later among the works it cites.
Q-Attention: Enabling efficient learning for vision-based robotic manipulation
S. James and A. J. Davison · 2022
Later among the works it cites.
Coarse-to-fine q-attention: Efficient learning for visual robotic manipulation via discretisation
S. James, K. Wada, T. Laidlow, and A. J. Davison · 2022
Later among the works it cites.
End-to-end learning to grasp via sampling from object point clouds
A. Alliegro, M. Rudorfer, F. Frattin, A. Leonardis, and T. Tommasi · 2022
Later among the works it cites.
Sagci-system: Towards sample-efficient, generalizable, compositional, and incremental robot learning
J. Lv, Q. Yu, L. Shao, W. Liu, W. Xu, and C. Lu · 2022
Later among the works it cites.
Q-attention: Enabling efficient learning for vision-based robotic manipulation
S. James and A. J. Davison · 2022
Later among the works it cites.
FlowBot3D: Learning 3D Articulation Flow to Manipulate Articulated Objects
B. Eisner, H. Zhang, and D. Held · 2022
Later among the works it cites.
Dexpoint: Generalizable point cloud reinforcement learning for sim-to-real dexterous manipulation
Y. Qin, B. Huang, Z.-H. Yin, H. Su, and X. Wang · 2022
Later among the works it cites.
Toolflownet: Robotic manipulation with tools via predicting tool flow from point clouds
D. Seita, Y. Wang, S. J. Shetty, E. Y. Li, Z. Erickson, and D. Held · 2022
Later among the works it cites.
Dexpoint: Generalizable point cloud reinforcement learning for sim-to-real dexterous manipulation
Y. Qin, B. Huang, Z.-H. Yin, H. Su, and X. Wang · 2022
Later among the works it cites.
Palm-e: An embodied multimodal language model
D. Driess, F. Xia, M. S. M. Sajjadi, C. Lynch, A. Chowdhery, B. Ichter, A. Wahid, J. Tompson, Q. Vuong, T. Yu, W. Huang, Y. Chebotar, P. Sermanet, D. Duckworth, S. Levine, V. Vanhoucke, K. Hausman, M. Toussaint, K. Greff, A. Zeng, I. Mordatch, and P. Florence · 2023
Closest in time.
Instruct2act: Mapping multi-modality instructions to robotic actions with large language model, 2023
S. Huang, Z. Jiang, H. Dong, Y. Qiao, P. Gao, and H. Li · 2023
Closest in time.
Ulip-2: Towards scalable multimodal pre-training for 3d understanding
L. Xue, N. Yu, S. Zhang, J. Li, R. Martín-Martín, J. Wu, C. Xiong, R. Xu, J. C. Niebles, and S. Savarese · 2023
Closest in time.
Polarnet project, 2023
S. Chen, R. Garcia, C. Schmid, and I. Laptev · 2023
Closest in time.