Fetching the paper…
Reading the bibliography…
3D perceptual representations are well suited for robot manipulation as they easily encode occlusions and simplify spatial reasoning.
Rrt-connect: An efficient approach to single-query path planning
J. J. Kuffner and S. M. LaValle · 2000
Earlier work this paper cites.
The open motion planning library
I. A. Sucan, M. Moll, and L. E. Kavraki · 2012
Earlier work this paper cites.
Reducing the barrier to entry of complex robotic software: a moveit! case study
D. Coleman, I. Sucan, S. Chitta, and N. Correll · 2014
Earlier work this paper cites.
Sparse 3d convolutional neural networks, 2015
B. Graham · 2015
Earlier work this paper cites.
Feature pyramid networks for object detection
T.-Y. Lin, P. Dollár, R. Girshick, K. He, B. Hariharan, and S. Belongie · 2017
Earlier work this paper cites.
Self-attention with relative position representations, 2018
P. Shaw, J. Uszkoreit, and A. Vaswani · 2018
Earlier work this paper cites.
4d spatio-temporal convnets: Minkowski convolutional neural networks, 2019
C. Choy, J. Gwak, and S. Savarese · 2019
Earlier work this paper cites.
Learning spatial common sense with geometry-aware recurrent networks
H.-Y. F. Tung, R. Cheng, and K. Fragkiadaki · 2019
Earlier work this paper cites.
Learning from unlabelled videos using contrastive predictive neural 3d mapping
A. W. Harley, S. K. Lakshmikanth, F. Li, X. Zhou, H.-Y. F. Tung, and K. Fragkiadaki · 2019
Earlier work this paper cites.
Graph-structured visual imitation
M. Sieb, Z. Xian, A. Huang, O. Kroemer, and K. Fragkiadaki · 2020
Earlier work this paper cites.
3d-oes: Viewpoint-invariant object-factorized environment simulators
H.-Y. F. Tung, Z. Xian, M. Prabhudesai, S. Lal, and K. Fragkiadaki · 2020
Earlier work this paper cites.
Rlbench: The robot learning benchmark & learning environment
S. James, Z. Ma, D. R. Arrojo, and A. J. Davison · 2020
Earlier work this paper cites.
Transporter networks: Rearranging the visual world for robotic manipulation
A. Zeng, P. Florence, J. Tompson, S. Welker, J. Chien, M. Attarian, T. Armstrong, I. Krasin, D. Duong, V. Sindhwani, et al · 2021
Earlier work this paper cites.
Perceiver: General perception with iterative attention, 2021
A. Jaegle, F. Gimeno, A. Brock, A. Zisserman, O. Vinyals, and J. Carreira · 2021
Earlier work this paper cites.
Roformer: Enhanced transformer with rotary position embedding
J. Su, Y. Lu, S. Pan, A. Murtadha, B. Wen, and Y. Liu · 2021
Earlier work this paper cites.
Learning to rearrange deformable cables, fabrics, and bags with goal-conditioned transporter networks, 2021
D. Seita, P. Florence, J. Tompson, E. Coumans, V. Sindhwani, K. Goldberg, and A. Zeng · 2021
Cited alongside, same era.
Learning to see before learning to act: Visual pre-training for manipulation, 2021
L. Yen-Chen, A. Zeng, S. Song, P. Isola, and T.-Y. Lin · 2021
Cited alongside, same era.
Hyperdynamics: Meta-learning object and agent dynamics with hypernetworks
Z. Xian, S. Lal, H.-Y. Tung, E. A. Platanios, and K. Fragkiadaki · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision, 2021
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever · 2021
Cited alongside, same era.
Instruction-following agents with jointly pre-trained vision-language models
H. Liu, L. Lee, K. Lee, and P. Abbeel · 2022
The unsurprising effectiveness of pre-trained vision models for control, 2022
S. Parisi, A. Rajeswaran, S. Purushwalkam, and A. Gupta · 2022
Later among the works it cites.
Swin transformer v2: Scaling up capacity and resolution, 2022
Z. Liu, H. Hu, Y. Lin, Z. Yao, Z. Xie, Y. Wei, J. Ning, Y. Cao, Z. Zhang, L. Dong, F. Wei, and B. Guo · 2022
Later among the works it cites.
Roformer: Enhanced transformer with rotary position embedding, 2022
J. Su, Y. Lu, S. Pan, A. Murtadha, B. Wen, and Y. Liu · 2022
Later among the works it cites.
Point transformer v2: Grouped vector attention and partition-based pooling, 2022
X. Wu, Y. Lao, L. Jiang, X. Liu, and H. Zhao · 2022
Later among the works it cites.
Q-attention: Enabling efficient learning for vision-based robotic manipulation
S. James and A. J. Davison · 2022
Later among the works it cites.
Do as i can, not as i say: Grounding language in robotic affordances
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Coarse-to-fine q-attention: Efficient learning for visual robotic manipulation via discretisation
S. James, K. Wada, T. Laidlow, and A. J. Davison · 2022
Cited alongside, same era.
Rt-1: Robotics transformer for real-world control at scale
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, J. Hsu, et al · 2022
Cited alongside, same era.
Sample efficient grasp learning using equivariant models, 2022
X. Zhu, D. Wang, O. Biza, G. Su, R. Walters, and R. Platt · 2022
Cited alongside, same era.
Auto-lambda: Disentangling dynamic task relationships
S. Liu, S. James, A. J. Davison, and E. Johns · 2022
Cited alongside, same era.
Behavior transformers: Cloning k k modes with one stone
N. M. Shafiullah, Z. Cui, A. A. Altanzaya, and L. Pinto · 2022
Cited alongside, same era.
S. Reed, K. Zolna, E. Parisotto, S. G. Colmenarejo, A. Novikov, G. Barth-Maron, M. Gimenez, Y. Sulsky, J. Kay, J. T. Springenberg, et al · 2022
Cited alongside, same era.
Bc-z: Zero-shot task generalization with robotic imitation learning
E. Jang, A. Irpan, M. Khansari, D. Kappler, F. Ebert, C. Lynch, S. Levine, and C. Finn · 2022
Cited alongside, same era.
M. Ahn, A. Brohan, N. Brown, Y. Chebotar, O. Cortes, B. David, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, et al · 2022
Later among the works it cites.
Perceiver-actor: A multi-task transformer for robotic manipulation
M. Shridhar, L. Manuelli, and D. Fox · 2023
Closest in time.
Instruction-driven history-aware policies for robotic manipulations
P.-L. Guhur, S. Chen, R. G. Pinel, M. Tapaswi, I. Laptev, and C. Schmid · 2023
Closest in time.
Spatial-language attention policies for efficient robot learning
P. Parashar, J. Vakil, S. Powers, and C. Paxton · 2023
Closest in time.
Energy-based models as zero-shot planners for compositional scene rearrangement
N. Gkanatsios, A. Jain, Z. Xian, Y. Zhang, C. Atkeson, and K. Fragkiadaki · 2023
Closest in time.
Open-world object manipulation using pre-trained vision-language models, 2023
A. Stone, T. Xiao, Y. Lu, K. Gopalakrishnan, K.-H. Lee, Q. Vuong, P. Wohlhart, B. Zitkovich, F. Xia, C. Finn, and K. Hausman · 2023
Closest in time.
Swin3d: A pretrained transformer backbone for 3d indoor scene understanding, 2023
Y.-Q. Yang, Y.-X. Guo, J.-Y. Xiong, Y. Liu, H. Pan, P.-S. Wang, X. Tong, and B. Guo · 2023
Closest in time.
Grounded decoding: Guiding text generation with grounded models for robot control
W. Huang, F. Xia, D. Shah, D. Driess, A. Zeng, Y. Lu, P. Florence, I. Mordatch, S. Levine, K. Hausman, et al · 2023
Closest in time.
Text2motion: From natural language instructions to feasible plans
K. Lin, C. Agia, T. Migimatsu, M. Pavone, and J. Bohg · 2023
Closest in time.