Fetching the paper…
Reading the bibliography…
Large-scale endeavors like and widespread community efforts such as Open-X-Embodiment have contributed to growing the scale of robot demonstration data.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
The open motion planning library
I. A. Sucan, M. Moll, and L. E. Kavraki · 2012
Earlier work this paper cites.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
Scaling laws for neural language models
J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei · 2020
Earlier work this paper cites.
Don’t stop pretraining: Adapt language models to domains and tasks
S. Gururangan, A. Marasović, S. Swayamdipta, K. Lo, I. Beltagy, D. Downey, and N. A. Smith · 2020
Earlier work this paper cites.
Rlbench: The robot learning benchmark & learning environment
S. James, Z. Ma, D. R. Arrojo, and A. J. Davison · 2020
Earlier work this paper cites.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Earlier work this paper cites.
Deduplicating training data makes language models better
K. Lee, D. Ippolito, A. Nystrom, C. Zhang, D. Eck, C. Callison-Burch, and N. Carlini · 2021
Earlier work this paper cites.
Integrated task and motion planning
C. R. Garrett, R. Chitnis, R. Holladay, B. Kim, T. Silver, L. P. Kaelbling, and T. Lozano-Pérez · 2021
Earlier work this paper cites.
Rt-1: Robotics transformer for real-world control at scale
A. Brohan, N. Brown, J. Carbajal, Y. Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Hausman, A. Herzog, J. Hsu, et al · 2022
Earlier work this paper cites.
Training compute-optimal large language models
J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. d. L. Casas, L. A. Hendricks, J. Welbl, A. Clark, et al · 2022
Earlier work this paper cites.
Laion-5b: An open large-scale dataset for training next generation image-text models
C. Schuhmann, R. Beaumont, R. Vencu, C. Gordon, R. Wightman, M. Cherti, T. Coombes, A. Katta, C. Mullis, M. Wortsman, et al · 2022
Earlier work this paper cites.
Quality not quantity: On the interaction between dataset design and robustness of clip
T. Nguyen, G. Ilharco, M. Wortsman, S. Oh, and L. Schmidt · 2022
Earlier work this paper cites.
Data determines distributional robustness in contrastive language image pre-training (clip)
A. Fang, G. Ilharco, M. Wortsman, Y. Wan, V. Shankar, A. Dave, and L. Schmidt · 2022
Earlier work this paper cites.
LAION-5b: An open large-scale dataset for training next generation image-text models
C. Schuhmann, R. Beaumont, R. Vencu, C. W. Gordon, R. Wightman, M. Cherti, T. Coombes, A. Katta, C. Mullis, M. Wortsman, P. Schramowski, S. R. Kundurthy, K. Crowson, L. Schmidt, R. Kaczmarczyk, and J. Jitsev · 2022
Earlier work this paper cites.
Cliport: What and where pathways for robotic manipulation
M. Shridhar, L. Manuelli, and D. Fox · 2022
Earlier work this paper cites.
Ego4d: Around the world in 3,000 hours of egocentric video
K. Grauman, A. Westbury, E. Byrne, Z. Chavis, A. Furnari, R. Girdhar, J. Hamburger, H. Jiang, M. Liu, X. Liu, et al · 2022
Earlier work this paper cites.
Sprint: Scalable semantic policy pre-training via language instruction relabeling
J. Zhang, K. Pertsch, J. Zhang, T. Nam, S. J. Hwang, X. Ren, and J. J. Lim · 2022
Earlier work this paper cites.
Inner monologue: Embodied reasoning through planning with language models
W. Huang, F. Xia, T. Xiao, H. Chan, J. Liang, P. Florence, A. Zeng, J. Tompson, I. Mordatch, Y. Chebotar, et al · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al · 2022
Earlier work this paper cites.
Self-instruct: Aligning language models with self-generated instructions
Y. Wang, Y. Kordi, S. Mishra, A. Liu, N. A. Smith, D. Khashabi, and H. Hajishirzi · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou, et al · 2022
Cited alongside, same era.
Open X-Embodiment: Robotic learning datasets and RT-X models
O. X.-E. Collaboration · 2023
Cited alongside, same era.
Voxposer: Composable 3d value maps for robotic manipulation with language models
W. Huang, C. Wang, R. Zhang, Y. Li, J. Wu, and L. Fei-Fei · 2023
Cited alongside, same era.
Scaling up and distilling down: Language-guided robot skill acquisition
3d-llm: Injecting the 3d world into large language models
Y. Hong, H. Zhen, P. Chen, S. Zheng, Y. Du, Z. Chen, and C. Gan · 2023
Later among the works it cites.
Rvt: Robotic view transformer for 3d object manipulation
A. Goyal, J. Xu, Y. Guo, V. Blukis, Y.-W. Chao, and D. Fox · 2023
Later among the works it cites.
Mvtrans: Multi-view perception of transparent objects
Y. R. Wang, Y. Zhao, H. Xu, S. Eppel, A. Aspuru-Guzik, F. Shkurti, and A. Garg · 2023
Later among the works it cites.
M2t2: Multi-task masked transformer for object-centric pick and place, 2023
W. Yuan, A. Murali, A. Mousavian, and D. Fox · 2023
Later among the works it cites.
Qwen-vl: A frontier large vision-language model with versatile abilities
J. Bai, S. Bai, S. Yang, S. Wang, S. Tan, P. Wang, J. Lin, C. Zhou, and J. Zhou · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
H. Ha, P. Florence, and S. Song · 2023
Cited alongside, same era.
Code as policies: Language model programs for embodied control
J. Liang, W. Huang, F. Xia, P. Xu, K. Hausman, B. Ichter, P. Florence, and A. Zeng · 2023
Cited alongside, same era.
Reproducible scaling laws for contrastive language-image learning
M. Cherti, R. Beaumont, R. Wightman, M. Wortsman, G. Ilharco, C. Gordon, C. Schuhmann, L. Schmidt, and J. Jitsev · 2023
Cited alongside, same era.
Learning vision from models rivals learning vision from data
Y. Tian, L. Fan, K. Chen, D. Katabi, D. Krishnan, and P. Isola · 2023
Cited alongside, same era.
H. Xu, S. Xie, X. E. Tan, P.-Y. Huang, R. Howes, V. Sharma, S.-W. Li, G. Ghosh, L. Zettlemoyer, and C. Feichtenhofer · 2023
Cited alongside, same era.
Meditron-70b: Scaling medical pretraining for large language models
Z. Chen, A. H. Cano, A. Romanou, A. Bonnet, K. Matoba, F. Salvi, M. Pagliardini, S. Fan, A. Köpf, A. Mohtashami, et al · 2023
Cited alongside, same era.
Ar2-d2: Training a robot without a robot
J. Duan, Y. R. Wang, M. Shridhar, D. Fox, and R. Krishna · 2023
Cited alongside, same era.
Progprompt: Generating situated robot task plans using large language models
I. Singh, V. Blukis, A. Mousavian, A. Goyal, D. Xu, J. Tremblay, D. Fox, J. Thomason, and A. Garg · 2023
Cited alongside, same era.
V. Ramanujan, T. Nguyen, S. Oh, A. Farhadi, and L. Schmidt · 2024
Closest in time.
V. Udandarao, A. Prabhu, A. Ghosh, Y. Sharma, P. H. Torr, A. Bibi, S. Albanie, and M. Bethge · 2024
Closest in time.
Datacomp: In search of the next generation of multimodal datasets
S. Y. Gadre, G. Ilharco, A. Fang, J. Hayase, G. Smyrnis, T. Nguyen, R. Marten, M. Wortsman, D. Ghosh, J. Zhang, et al · 2024
Closest in time.
Lima: Less is more for alignment
C. Zhou, P. Liu, P. Xu, S. Iyer, J. Sun, Y. Mao, X. Ma, A. Efrat, P. Yu, L. Yu, et al · 2024
Closest in time.
Improving multimodal datasets with image captioning
T. Nguyen, S. Y. Gadre, G. Ilharco, S. Oh, and L. Schmidt · 2024
Closest in time.
Universal manipulation interface: In-the-wild robot teaching without in-the-wild robots
C. Chi, Z. Xu, C. Pan, E. Cousineau, B. Burchfiel, S. Feng, R. Tedrake, and S. Song · 2024
Closest in time.
Dexcap: Scalable and portable mocap data collection system for dexterous manipulation
C. Wang, H. Shi, W. Wang, R. Zhang, L. Fei-Fei, and C. K. Liu · 2024
Closest in time.
Pivot: Iterative visual prompting elicits actionable knowledge for vlms, 2024
S. Nasiriany, F. Xia, W. Yu, T. Xiao, J. Liang, I. Dasgupta, A. Xie, D. Driess, A. Wahid, Z. Xu, Q. Vuong, T. Zhang, T.-W. E. Lee, K.-H. Lee, P. Xu, S. Kirmani, Y. Zhu, A. Zeng, K. Hausman, N. Heess, C. Finn, S. Levine, and B. Ichter · 2024
Closest in time.
Copa: General robotic manipulation through spatial constraints of parts with foundation models
H. Huang, F. Lin, Y. Hu, S. Wang, and Y. Gao · 2024
Closest in time.
Imitating shortest paths in simulation enables effective navigation and manipulation in the real world
K. Ehsani, T. Gupta, R. Hendrix, J. Salvador, L. Weihs, K.-H. Zeng, K. P. Singh, Y. Kim, W. Han, A. Herrasti, et al · 2024
Closest in time.
Autort: Embodied foundation models for large scale orchestration of robotic agents
M. Ahn, D. Dwibedi, C. Finn, M. G. Arenas, K. Gopalakrishnan, K. Hausman, B. Ichter, A. Irpan, N. Joshi, R. Julian, et al · 2024
Closest in time.
Rl-vlm-f: Reinforcement learning from vision language foundation model feedback
Y. Wang, Z. Sun, J. Zhang, Z. Xian, E. Biyik, D. Held, and Z. Erickson · 2024
Closest in time.
Robopoint: A vision-language model for spatial affordance prediction for robotics
W. Yuan, J. Duan, V. Blukis, W. Pumacay, R. Krishna, A. Murali, A. Mousavian, and D. Fox · 2024
Closest in time.
Rvt-2: Learning precise manipulation from few demonstrations
A. Goyal, V. Blukis, J. Xu, Y. Guo, Y.-W. Chao, and D. Fox · 2024
Closest in time.