Fetching the paper…
Reading the bibliography…
Generalizing language-conditioned robotic policies to new tasks remains a significant challenge, hampered by the lack of suitable simulation benchmarks.
O. Ronneberger, P. Fischer, and T. Brox, “U-Net: Convolutional networks for biomedical image segmentation,” in MICCAI , 2015
2015
Earlier work this paper cites.
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in NeurIPS , 2017
2017
Earlier work this paper cites.
D. Kalashnikov, A. Irpan, P. Pastor, J. Ibarz, A. Herzog, E. Jang, D. Quillen, E. Holly, M. Kalakrishnan, V. Vanhoucke et al. , “Scalable deep reinforcement learning for vision-based robotic manipulation,” in CoRL , 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
S. Karamcheti, S. Nair, A. S. Chen, T. Kollar, C. Finn, D. Sadigh, and P. Liang, “Language-driven representation learning for robotics,” 2019
2019
Earlier work this paper cites.
M. Shridhar, J. Thomason, D. Gordon, Y. Bisk, W. Han, R. Mottaghi, L. Zettlemoyer, and D. Fox, “ALFRED: A benchmark for interpreting grounded instructions for everyday tasks,” in CVPR , 2020
2020
Earlier work this paper cites.
A. Zeng, P. Florence, J. Tompson, S. Welker, J. Chien, M. Attarian, T. Armstrong, I. Krasin, D. Duong, V. Sindhwani, and J. Lee, “Transporter networks: Rearranging the visual world for robotic manipulation,” in CoRL , 2020
2020
Earlier work this paper cites.
S. James, Z. Ma, D. R. Arrojo, and A. J. Davison, “RLBench: The robot learning benchmark & learning environment,” IEEE RA-L , 2020
2020
Earlier work this paper cites.
S. Stepputtis, J. Campbell, M. Phielipp, S. Lee, C. Baral, and H. Ben Amor, “Language-conditioned imitation learning for robot manipulation tasks,” in NeurIPS , 2020
2020
Earlier work this paper cites.
V. Makoviychuk, L. Wawrzyniak, Y. Guo, M. Lu, K. Storey, M. Macklin, D. Hoeller, N. Rudin, A. Allshire, A. Handa, and G. State, “Isaac Gym: High performance gpu based physics simulation for robot learning.” in NeurIPS Datasets and Benchmarks , 2021
2021
Earlier work this paper cites.
L. Shao, T. Migimatsu, Q. Zhang, K. Yang, and J. Bohg, “Concept2Robot: Learning manipulation concepts from instructions and human demonstrations,” IJRR , 2021
2021
Earlier work this paper cites.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al. , “Learning transferable visual models from natural language supervision,” in ICML , 2021
2021
Earlier work this paper cites.
S. Nair, A. Rajeswaran, V. Kumar, C. Finn, and A. Gupta, “R3M: A universal visual representation for robot manipulation,” in CoRL , 2022
2022
Earlier work this paper cites.
K. Zheng, X. Chen, O. C. Jenkins, and X. Wang, “VLMbench: A compositional benchmark for vision-and-language manipulation,” in NeurIPS , 2022
2022
Earlier work this paper cites.
O. Mees, L. Hermann, E. Rosete-Beas, and W. Burgard, “CALVIN: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks,” IEEE RA-L , 2022
2022
Earlier work this paper cites.
E. Jang, A. Irpan, M. Khansari, D. Kappler, F. Ebert, C. Lynch, S. Levine, and C. Finn, “BC-Z: Zero-shot task generalization with robotic imitation learning,” in CoRL , 2022
2022
Earlier work this paper cites.
S. James, K. Wada, T. Laidlow, and A. J. Davison, “Coarse-to-fine q-attention: Efficient learning for visual robotic manipulation via discretisation,” in CVPR , 2022
2022
Earlier work this paper cites.
W. Huang, P. Abbeel, D. Pathak, and I. Mordatch, “Language models as zero-shot planners: Extracting actionable knowledge for embodied agents,” in ICML , 2022
2022
Cited alongside, same era.
S. Liu, S. James, A. J. Davison, and E. Johns, “Auto-Lambda: Disentangling dynamic task relationships,” TMLR , 2022
2022
Cited alongside, same era.
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, J. Hsu et al. , “RT-1: Robotics transformer for real-world control at scale,” in RSS , 2023
2023
Cited alongside, same era.
S. Chen, R. Garcia, C. Schmid, and I. Laptev, “PolarNet: 3D point clouds for language-guided robotic manipulation,” in CoRL , 2023
2023
Cited alongside, same era.
W. Huang, C. Wang, R. Zhang, Y. Li, J. Wu, and L. Fei-Fei, “VoxPoser: Composable 3D value maps for robotic manipulation with language models,” in CoRL , 2023
A. Goyal, J. Xu, Y. Guo, V. Blukis, Y.-W. Chao, and D. Fox, “RVT: Robotic view transformer for 3D object manipulation,” in CoRL , 2023
2023
Later among the works it cites.
C. Chi, S. Feng, Y. Du, Z. Xu, E. Cousineau, B. Burchfiel, and S. Song, “Diffusion policy: Visuomotor policy learning via action diffusion,” RSS , 2023
2023
Later among the works it cites.
M. Minderer, A. A. Gritsenko, and N. Houlsby, “Scaling open-vocabulary object detection,” in NeurIPS , A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, Eds., 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2023
Cited alongside, same era.
T. Gervet, Z. Xian, N. Gkanatsios, and K. Fragkiadaki, “Act3D: 3D feature field transformers for multi-task robotic manipulation,” in CoRL , 2023
2023
Cited alongside, same era.
I. Radosavovic, T. Xiao, S. James, P. Abbeel, J. Malik, and T. Darrell, “Real-world robot learning with masked visual pre-training,” in CoRL , 2023
2023
Cited alongside, same era.
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, X. Chen, K. Choromanski, T. Ding, D. Driess, A. Dubey, C. Finn et al. , “RT-2: Vision-language-action models transfer web knowledge to robotic control,” in CoRL , 2023
2023
Cited alongside, same era.
D. Driess, F. Xia, M. S. Sajjadi, C. Lynch, A. Chowdhery, B. Ichter, A. Wahid, J. Tompson, Q. Vuong, T. Yu et al. , “PALM-E: An embodied multimodal language model,” in ICML , 2023
2023
Cited alongside, same era.
Q. Vuong, S. Levine, H. R. Walke, K. Pertsch, A. Singh, R. Doshi, C. Xu, J. Luo, L. Tan, D. Shah et al. , “Open X-Embodiment: Robotic learning datasets and RT-X models,” in CoRL , 2023
2023
Cited alongside, same era.
A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y. Lo, P. Dollár, and R. Girshick, “Segment anything,” in ICCV , 2023
2023
Cited alongside, same era.
J. Liang, W. Huang, F. Xia, P. Xu, K. Hausman, B. Ichter, P. Florence, and A. Zeng, “Code as policies: Language model programs for embodied control,” in ICRA , 2023
2023
Cited alongside, same era.
2023
Later among the works it cites.
A. Brohan, Y. Chebotar, C. Finn, K. Hausman, A. Herzog, D. Ho, J. Ibarz, A. Irpan, E. Jang, R. Julian et al. , “Do as I can, not as I say: Grounding language in robotic affordances,” in CoRL , 2023
2023
Later among the works it cites.
2023
Later among the works it cites.
G. OpenAI, “GPT-4V(ision) system card,” preprint , 2023
2023
Later among the works it cites.
S. Dubey, S. K. Singh, and B. Chaudhuri, “AdaNorm: Adaptive gradient norm correction based optimizer for CNNs,” in WACV , 2023
2023
Later among the works it cites.
S. Chen, R. Garcia, I. Laptev, and C. Schmid, “SUGAR: Pre-training 3D visual representations for robotics,” in CVPR , 2024
2024
Closest in time.
O. M. Team, D. Ghosh, H. Walke, K. Pertsch, K. Black, O. Mees, S. Dasari, J. Hejna, T. Kreiman, C. Xu et al. , “Octo: An open-source generalist robot policy,” 2024
2024
Closest in time.
M. J. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. Foster, G. Lam, P. Sanketi et al. , “OpenVLA: An open-source vision-language-action model,” CoRL , 2024
2024
Closest in time.
Meta, “Llama 3 model card,” 2024
2024
Closest in time.
2024
Closest in time.
T.-W. Ke, N. Gkanatsios, and K. Fragkiadaki, “3D Diffuser Actor: Policy diffusion with 3D scene representations,” in CoRL , 2024
2024
Closest in time.
A. Goyal, V. Blukis, J. Xu, Y. Guo, Y.-W. Chao, and D. Fox, “RVT2: Learning precise manipulation from few demonstrations,” in RSS , 2024
2024
Closest in time.
X. Wu, L. Jiang, P.-S. Wang, Z. Liu, X. Liu, Y. Qiao, W. Ouyang, T. He, and H. Zhao, “Point Transformer v3: Simpler, faster, stronger,” in CVPR , 2024
2024
Closest in time.