Fetching the paper…
Reading the bibliography…
Embodied AI is widely recognized as a cornerstone of artificial general intelligence (AGI) because it involves controlling embodied agents to perform tasks in the physical world.
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. P. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu, “Asynchronous methods for deep reinforcement learning,” in ICML , ser. JMLR Workshop and Conference Proceedings, vol. 48. JMLR.org, 2016, pp. 1928–1937
1937
Earlier work this paper cites.
H. Arai, K. Miwa, K. Sasaki, K. Watanabe, Y. Yamaguchi, S. Aoki, and I. Yamamoto, “Covla: Comprehensive vision-language-action dataset for autonomous driving,” in WACV . IEEE, 2025, pp. 1933–1943
1943
Earlier work this paper cites.
N. G. Tsung-Wei Ke and and K. Fragkiadaki, “3d diffuser actor: Policy diffusion with 3d scene representations,” in CoRL , ser. Proceedings of Machine Learning Research, vol. 270. PMLR, 2024, pp. 1949–1974
1974
Earlier work this paper cites.
Y. LeCun, B. E. Boser, J. S. Denker, D. Henderson, R. E. Howard, W. E. Hubbard, and L. D. Jackel, “Backpropagation applied to handwritten zip code recognition,” Neural Comput. , vol. 1, no. 4, pp. 541–551, 1989
1989
Earlier work this paper cites.
J. L. Elman, “Finding structure in time,” Cogn. Sci. , vol. 14, no. 2, pp. 179–211, 1990
1990
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural Comput. , vol. 9, no. 8, pp. 1735–1780, 1997
1997
Earlier work this paper cites.
Y. Bengio, R. Ducharme, and P. Vincent, “A neural probabilistic language model,” in NIPS . MIT Press, 2000, pp. 932–938
2000
Earlier work this paper cites.
J. Tang, J. Zhang, L. Yao, J. Li, L. Zhang, and Z. Su, “Arnetminer: extraction and mining of academic social networks,” in KDD . ACM, 2008, pp. 990–998
2008
Earlier work this paper cites.
F. Scarselli, M. Gori, A. C. Tsoi, M. Hagenbuchner, and G. Monfardini, “The graph neural network model,” IEEE Trans. Neural Networks , vol. 20, no. 1, pp. 61–80, 2009
2009
Earlier work this paper cites.
A. Micheli, “Neural network for graphs: A contextual constructive approach,” IEEE Trans. Neural Networks , vol. 20, no. 3, pp. 498–511, 2009
2009
Earlier work this paper cites.
S. Yenamandra, A. Ramachandran, K. Yadav, A. S. Wang, M. Khanna, T. Gervet, T. Yang, V. Jain, A. Clegg, J. M. Turner, Z. Kira, M. Savva, A. X. Chang, D. S. Chaplot, D. Batra, R. Mottaghi, Y. Bisk, and C. Paxton, “Homerobot: Open-vocabulary mobile manipulation,” in CoRL , vol. 229. PMLR, 2023, pp. 1975–2011
2011
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” in NIPS , 2012, pp. 1106–1114
2012
Earlier work this paper cites.
T. Mikolov, I. Sutskever, K. Chen, G. S. Corrado, and J. Dean, “Distributed representations of words and phrases and their compositionality,” in NIPS , 2013, pp. 3111–3119
2013
Earlier work this paper cites.
T. Mikolov, K. Chen, G. Corrado, and J. Dean, “Efficient estimation of word representations in vector space,” in ICLR (Workshop Poster) , 2013
2013
Earlier work this paper cites.
S. Levine and V. Koltun, “Guided policy search,” in ICML (3) , ser. JMLR Workshop and Conference Proceedings, vol. 28. JMLR.org, 2013, pp. 1–9
2013
Earlier work this paper cites.
K. Cho, B. van Merrienboer, Ç. Gülçehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio, “Learning phrase representations using RNN encoder-decoder for statistical machine translation,” in EMNLP . ACL, 2014, pp. 1724–1734
2014
Earlier work this paper cites.
R. B. Girshick, J. Donahue, T. Darrell, and J. Malik, “Rich feature hierarchies for accurate object detection and semantic segmentation,” in CVPR . IEEE Computer Society, 2014, pp. 580–587
2014
Earlier work this paper cites.
J. Pennington, R. Socher, and C. D. Manning, “Glove: Global vectors for word representation,” in EMNLP . ACL, 2014, pp. 1532–1543
2014
Earlier work this paper cites.
Y. Kim, “Convolutional neural networks for sentence classification,” in EMNLP . ACL, 2014, pp. 1746–1751
2014
Earlier work this paper cites.
D. Silver, G. Lever, N. Heess, T. Degris, D. Wierstra, and M. A. Riedmiller, “Deterministic policy gradient algorithms,” in ICML , ser. JMLR Workshop and Conference Proceedings, vol. 32. JMLR.org, 2014, pp. 387–395
2014
Earlier work this paper cites.
J. Bruna, W. Zaremba, A. Szlam, and Y. LeCun, “Spectral networks and locally connected networks on graphs,” in ICLR , 2014
2014
Earlier work this paper cites.
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. A. Riedmiller, A. Fidjeland, G. Ostrovski, S. Petersen, C. Beattie, A. Sadik, I. Antonoglou, H. King, D. Kumaran, D. Wierstra, S. Legg, and D. Hassabis, “Human-level control through deep reinforcement learning,” Nat. , vol. 518, no. 7540, pp. 529–533, 2015
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh, “VQA: visual question answering,” in ICCV . IEEE Computer Society, 2015, pp. 2425–2433
2015
Earlier work this paper cites.
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan, “Show and tell: A neural image caption generator,” in CVPR . IEEE Computer Society, 2015, pp. 3156–3164
2015
Earlier work this paper cites.
K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,” in ICLR , 2015
2015
Earlier work this paper cites.
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. E. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich, “Going deeper with convolutions,” in CVPR . IEEE Computer Society, 2015, pp. 1–9
2015
Earlier work this paper cites.
R. B. Girshick, “Fast R-CNN,” in ICCV . IEEE Computer Society, 2015, pp. 1440–1448
2015
Earlier work this paper cites.
S. Ren, K. He, R. B. Girshick, and J. Sun, “Faster R-CNN: towards real-time object detection with region proposal networks,” in NIPS , 2015, pp. 91–99
2015
Earlier work this paper cites.
J. Long, E. Shelhamer, and T. Darrell, “Fully convolutional networks for semantic segmentation,” in CVPR . IEEE Computer Society, 2015, pp. 3431–3440
2015
Earlier work this paper cites.
O. Ronneberger, P. Fischer, and T. Brox, “U-net: Convolutional networks for biomedical image segmentation,” in MICCAI (3) , ser. Lecture Notes in Computer Science, vol. 9351. Springer, 2015, pp. 234–241
2015
Earlier work this paper cites.
D. Bahdanau, K. Cho, and Y. Bengio, “Neural machine translation by jointly learning to align and translate,” in ICLR , 2015
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
X. Zhang, J. J. Zhao, and Y. LeCun, “Character-level convolutional networks for text classification,” in NIPS , 2015, pp. 649–657
2015
Earlier work this paper cites.
J. Schulman, S. Levine, P. Abbeel, M. I. Jordan, and P. Moritz, “Trust region policy optimization,” in ICML , ser. JMLR Workshop and Conference Proceedings, vol. 37. JMLR.org, 2015, pp. 1889–1897
2015
Earlier work this paper cites.
S. Levine, P. Pastor, A. Krizhevsky, and D. Quillen, “Learning hand-eye coordination for robotic grasping with large-scale data collection,” in ISER , ser. Springer Proceedings in Advanced Robotics, vol. 1. Springer, 2016, pp. 173–184
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
H. M. Le, T. N. Do, and S. J. Phee, “A survey on actuators-driven surgical robots,” Sensors and Actuators A-physical , vol. 247, pp. 323–354, 2016
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in CVPR . IEEE Computer Society, 2016, pp. 770–778
2016
Earlier work this paper cites.
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. van den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, S. Dieleman, D. Grewe, J. Nham, N. Kalchbrenner, I. Sutskever, T. P. Lillicrap, M. Leach, K. Kavukcuoglu, T. Graepel, and D. Hassabis, “Mastering the game of go with deep neural networks and tree search,” Nat. , vol. 529, no. 7587, pp. 484–489, 2016
2016
Earlier work this paper cites.
L. Yu, P. Poirson, S. Yang, A. C. Berg, and T. L. Berg, “Modeling context in referring expressions,” in ECCV (2) , ser. Lecture Notes in Computer Science, vol. 9906. Springer, 2016, pp. 69–85
2016
Earlier work this paper cites.
J. Redmon, S. K. Divvala, R. B. Girshick, and A. Farhadi, “You only look once: Unified, real-time object detection,” in CVPR . IEEE Computer Society, 2016, pp. 779–788
2016
Earlier work this paper cites.
C. R. Qi, H. Su, M. Nießner, A. Dai, M. Yan, and L. J. Guibas, “Volumetric and multi-view cnns for object classification on 3d data,” in CVPR . IEEE Computer Society, 2016, pp. 5648–5656
2016
Earlier work this paper cites.
X. Ma and E. H. Hovy, “End-to-end sequence labeling via bi-directional lstm-cnns-crf,” in ACL (1) . The Association for Computer Linguistics, 2016
2016
Earlier work this paper cites.
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra, “Continuous control with deep reinforcement learning,” in ICLR (Poster) , 2016
2016
Earlier work this paper cites.
S. Gu, T. P. Lillicrap, I. Sutskever, and S. Levine, “Continuous deep q-learning with model-based acceleration,” in ICML , ser. JMLR Workshop and Conference Proceedings, vol. 48. JMLR.org, 2016, pp. 2829–2838
2016
Earlier work this paper cites.
J. Schulman, P. Moritz, S. Levine, M. I. Jordan, and P. Abbeel, “High-dimensional continuous control using generalized advantage estimation,” in ICLR (Poster) , 2016
2016
Earlier work this paper cites.
J. Ho and S. Ermon, “Generative adversarial imitation learning,” in NIPS , 2016, pp. 4565–4573
2016
Earlier work this paper cites.
S. Levine, C. Finn, T. Darrell, and P. Abbeel, “End-to-end training of deep visuomotor policies,” J. Mach. Learn. Res. , vol. 17, pp. 39:1–39:40, 2016
2016
Earlier work this paper cites.
M. Defferrard, X. Bresson, and P. Vandergheynst, “Convolutional neural networks on graphs with fast localized spectral filtering,” in NIPS , 2016, pp. 3837–3845
2016
Earlier work this paper cites.
S. Cao, W. Lu, and Q. Xu, “Deep neural networks for learning graph representations,” in AAAI . AAAI Press, 2016, pp. 1145–1152
2016
Earlier work this paper cites.
T. N. Kipf and M. Welling, “Variational graph auto-encoders,” CoRR , vol. abs/1611.07308, 2016
2016
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in NIPS , 2017, pp. 5998–6008
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
C. Szegedy, S. Ioffe, V. Vanhoucke, and A. A. Alemi, “Inception-v4, inception-resnet and the impact of residual connections on learning,” in AAAI . AAAI Press, 2017, pp. 4278–4284
2017
Earlier work this paper cites.
S. Xie, R. B. Girshick, P. Dollár, Z. Tu, and K. He, “Aggregated residual transformations for deep neural networks,” in CVPR . IEEE Computer Society, 2017, pp. 5987–5995
2017
Earlier work this paper cites.
K. He, G. Gkioxari, P. Dollár, and R. B. Girshick, “Mask R-CNN,” in ICCV . IEEE Computer Society, 2017, pp. 2980–2988
2017
Earlier work this paper cites.
T. Lin, P. Dollár, R. B. Girshick, K. He, B. Hariharan, and S. J. Belongie, “Feature pyramid networks for object detection,” in CVPR . IEEE Computer Society, 2017, pp. 936–944
2017
Earlier work this paper cites.
T. Lin, P. Goyal, R. B. Girshick, K. He, and P. Dollár, “Focal loss for dense object detection,” in ICCV . IEEE Computer Society, 2017, pp. 2999–3007
2017
Earlier work this paper cites.
V. Badrinarayanan, A. Kendall, and R. Cipolla, “Segnet: A deep convolutional encoder-decoder architecture for image segmentation,” IEEE Trans. Pattern Anal. Mach. Intell. , vol. 39, no. 12, pp. 2481–2495, 2017
2017
Earlier work this paper cites.
A. Ioannidou, E. Chatzilari, S. Nikolopoulos, and I. Kompatsiaris, “Deep learning advances in computer vision with 3d data: A survey,” ACM Comput. Surv. , vol. 50, no. 2, pp. 20:1–20:38, 2017
2017
Earlier work this paper cites.
Y. Li, “Deep reinforcement learning: An overview,” CoRR , vol. abs/1701.07274, 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
M. Andrychowicz, D. Crow, A. Ray, J. Schneider, R. Fong, P. Welinder, B. McGrew, J. Tobin, P. Abbeel, and W. Zaremba, “Hindsight experience replay,” in NIPS , 2017, pp. 5048–5058
2017
Earlier work this paper cites.
L. Pinto, J. Davidson, R. Sukthankar, and A. Gupta, “Robust adversarial reinforcement learning,” in ICML , vol. 70. PMLR, 2017, pp. 2817–2826
2017
Earlier work this paper cites.
P. F. Christiano, J. Leike, T. B. Brown, M. Martic, S. Legg, and D. Amodei, “Deep reinforcement learning from human preferences,” in NIPS , 2017, pp. 4299–4307
2017
Earlier work this paper cites.
A. S. Vezhnevets, S. Osindero, T. Schaul, N. Heess, M. Jaderberg, D. Silver, and K. Kavukcuoglu, “Feudal networks for hierarchical reinforcement learning,” in ICML , vol. 70. PMLR, 2017, pp. 3540–3549
2017
Earlier work this paper cites.
T. N. Kipf and M. Welling, “Semi-supervised classification with graph convolutional networks,” in ICLR (Poster) , 2017
2017
Earlier work this paper cites.
J. Gilmer, S. S. Schoenholz, P. F. Riley, O. Vinyals, and G. E. Dahl, “Neural message passing for quantum chemistry,” in ICML , vol. 70. PMLR, 2017, pp. 1263–1272
2017
Earlier work this paper cites.
W. L. Hamilton, Z. Ying, and J. Leskovec, “Inductive representation learning on large graphs,” in NIPS , 2017, pp. 1024–1034
2017
Earlier work this paper cites.
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L. Li, D. A. Shamma, M. S. Bernstein, and L. Fei-Fei, “Visual genome: Connecting language and vision using crowdsourced dense image annotations,” Int. J. Comput. Vis. , vol. 123, no. 1, pp. 32–73, 2017
2017
Earlier work this paper cites.
C. R. Qi, H. Su, K. Mo, and L. J. Guibas, “Pointnet: Deep learning on point sets for 3d classification and segmentation,” in CVPR . IEEE Computer Society, 2017, pp. 77–85
2017
Earlier work this paper cites.
C. R. Qi, L. Yi, H. Su, and L. J. Guibas, “Pointnet++: Deep hierarchical feature learning on point sets in a metric space,” in NIPS , 2017, pp. 5099–5108
2017
Earlier work this paper cites.
R. Goyal, S. E. Kahou, V. Michalski, J. Materzynska, S. Westphal, H. Kim, V. Haenel, I. Fründ, P. Yianilos, M. Mueller-Freitag, F. Hoppe, C. Thurau, I. Bax, and R. Memisevic, “The ”something something” video database for learning and evaluating visual common sense,” in ICCV . IEEE Computer Society, 2017, pp. 5843–5851
2017
Earlier work this paper cites.
P. Sharma, L. Mohan, L. Pinto, and A. Gupta, “Multiple interactions made easy (MIME): large scale demonstrations data for imitation,” in CoRL , vol. 87. PMLR, 2018, pp. 906–915
2018
Earlier work this paper cites.
A. Mandlekar, Y. Zhu, A. Garg, J. Booher, M. Spero, A. Tung, J. Gao, J. Emmons, A. Gupta, E. Orbay, S. Savarese, and L. Fei-Fei, “ROBOTURK: A crowdsourcing platform for robotic skill learning through imitation,” in CoRL , vol. 87. PMLR, 2018, pp. 879–893
2018
Earlier work this paper cites.
F. Xia, A. R. Zamir, Z. He, A. Sax, J. Malik, and S. Savarese, “Gibson env: Real-world perception for embodied agents,” in CVPR . Computer Vision Foundation / IEEE Computer Society, 2018, pp. 9068–9079
2018
Earlier work this paper cites.
X. Puig, K. Ra, M. Boben, J. Li, T. Wang, S. Fidler, and A. Torralba, “Virtualhome: Simulating household activities via programs,” in 2018 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2018, Salt Lake City, UT, USA, June 18-22, 2018 . Computer Vision Foundation / IEEE Computer Society, 2018, pp. 8494–8502
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
A. Das, S. Datta, G. Gkioxari, S. Lee, D. Parikh, and D. Batra, “Embodied question answering,” in CVPR . Computer Vision Foundation / IEEE Computer Society, 2018, pp. 1–10
2018
Earlier work this paper cites.
D. Gordon, A. Kembhavi, M. Rastegari, J. Redmon, D. Fox, and A. Farhadi, “IQA: visual question answering in interactive environments,” in CVPR . Computer Vision Foundation / IEEE Computer Society, 2018, pp. 4089–4098
2018
Earlier work this paper cites.
R. Zhang, Z. Liu, L. Zhang, J. A. Whritner, K. S. Muller, M. M. Hayhoe, and D. H. Ballard, “AGIL: learning attention from human for visuomotor tasks,” in ECCV (11) , vol. 11215. Springer, 2018, pp. 692–707
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
A. Radford, K. Narasimhan, T. Salimans, and I. Sutskever, “Improving language understanding by generative pre-training,” OpenAI blog , 2018
2018
Earlier work this paper cites.
J. Hu, L. Shen, and G. Sun, “Squeeze-and-excitation networks,” in CVPR . Computer Vision Foundation / IEEE Computer Society, 2018, pp. 7132–7141
2018
Earlier work this paper cites.
P. Anderson, X. He, C. Buehler, D. Teney, M. Johnson, S. Gould, and L. Zhang, “Bottom-up and top-down attention for image captioning and visual question answering,” in CVPR . Computer Vision Foundation / IEEE Computer Society, 2018, pp. 6077–6086
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
M. E. Peters, M. Neumann, M. Iyyer, M. Gardner, C. Clark, K. Lee, and L. Zettlemoyer, “Deep contextualized word representations,” in NAACL-HLT . Association for Computational Linguistics, 2018, pp. 2227–2237
2018
Earlier work this paper cites.
J. Howard and S. Ruder, “Universal language model fine-tuning for text classification,” in ACL (1) . Association for Computational Linguistics, 2018, pp. 328–339
2018
Earlier work this paper cites.
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine, “Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor,” in ICML , vol. 80. PMLR, 2018, pp. 1856–1865
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
P. Velickovic, G. Cucurull, A. Casanova, A. Romero, P. Liò, and Y. Bengio, “Graph attention networks,” in ICLR (Poster) , 2018
2018
Earlier work this paper cites.
Y. Seo, M. Defferrard, P. Vandergheynst, and X. Bresson, “Structured sequence modeling with graph convolutional recurrent networks,” in ICONIP (1) , ser. Lecture Notes in Computer Science, vol. 11301. Springer, 2018, pp. 362–373
2018
Earlier work this paper cites.
D. K. Misra, A. Bennett, V. Blukis, E. Niklasson, M. Shatkhin, and Y. Artzi, “Mapping instructions to actions in 3d environments with visual goal prediction,” in EMNLP , 2018, pp. 2667–2678
2018
Earlier work this paper cites.
J. Devlin, M. Chang, K. Lee, and K. Toutanova, “BERT: pre-training of deep bidirectional transformers for language understanding,” in NAACL-HLT (1) . Association for Computational Linguistics, 2019, pp. 4171–4186
2019
Earlier work this paper cites.
S. Dasari, F. Ebert, S. Tian, S. Nair, B. Bucher, K. Schmeckpeper, S. Singh, S. Levine, and C. Finn, “Robonet: Large-scale multi-robot learning,” in CoRL , vol. 100. PMLR, 2019, pp. 885–897
2019
Earlier work this paper cites.
T. Yu, D. Quillen, Z. He, R. Julian, K. Hausman, C. Finn, and S. Levine, “Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning,” in CoRL , vol. 100. PMLR, 2019, pp. 1094–1100
2019
Earlier work this paper cites.
A. Gupta, V. Kumar, C. Lynch, S. Levine, and K. Hausman, “Relay policy learning: Solving long-horizon tasks via imitation and reinforcement learning,” in CoRL , vol. 100. PMLR, 2019, pp. 1025–1037
2019
Earlier work this paper cites.
M. Savva, J. Malik, D. Parikh, D. Batra, A. Kadian, O. Maksymets, Y. Zhao, E. Wijmans, B. Jain, J. Straub, J. Liu, and V. Koltun, “Habitat: A platform for embodied AI research,” in ICCV . IEEE, 2019, pp. 9338–9346
2019
Earlier work this paper cites.
L. Yu, X. Chen, G. Gkioxari, M. Bansal, T. L. Berg, and D. Batra, “Multi-target embodied question answering,” in CVPR . Computer Vision Foundation / IEEE, 2019, pp. 6309–6318
2019
Earlier work this paper cites.
E. Wijmans, S. Datta, O. Maksymets, A. Das, G. Gkioxari, S. Lee, I. Essa, D. Parikh, and D. Batra, “Embodied question answering in photorealistic environments with point cloud perception,” in CVPR . Computer Vision Foundation / IEEE, 2019, pp. 6659–6668
2019
Earlier work this paper cites.
C. Fan, “Egovqa - an egocentric video question answering benchmark dataset,” in ICCV Workshops . IEEE, 2019, pp. 4359–4366
2019
Earlier work this paper cites.
J. Lu, D. Batra, D. Parikh, and S. Lee, “Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks,” in NeurIPS , 2019, pp. 13–23
2019
Earlier work this paper cites.
M. Tan and Q. V. Le, “Efficientnet: Rethinking model scaling for convolutional neural networks,” in ICML , vol. 97. PMLR, 2019, pp. 6105–6114
2019
Earlier work this paper cites.
Y. Feng, Y. Feng, H. You, X. Zhao, and Y. Gao, “Meshnet: Mesh neural network for 3d shape representation,” in AAAI . AAAI Press, 2019, pp. 8279–8286
2019
Earlier work this paper cites.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever et al. , “Language models are unsupervised multitask learners,” OpenAI blog , vol. 1, no. 8, p. 9, 2019
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
Z. Yang, Z. Dai, Y. Yang, J. G. Carbonell, R. Salakhutdinov, and Q. V. Le, “Xlnet: Generalized autoregressive pretraining for language understanding,” in NeurIPS , 2019, pp. 5754–5764
2019
Earlier work this paper cites.
A. Kumar, J. Fu, M. Soh, G. Tucker, and S. Levine, “Stabilizing off-policy q-learning via bootstrapping error reduction,” in NeurIPS , 2019, pp. 11 761–11 771
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
K. Xu, W. Hu, J. Leskovec, and S. Jegelka, “How powerful are graph neural networks?” in ICLR , 2019
2019
Earlier work this paper cites.
D. Ghosal, N. Majumder, S. Poria, N. Chhaya, and A. F. Gelbukh, “Dialoguegcn: A graph convolutional neural network for emotion recognition in conversation,” in EMNLP/IJCNLP (1) . Association for Computational Linguistics, 2019, pp. 154–164
2019
Earlier work this paper cites.
H. Tan and M. Bansal, “LXMERT: learning cross-modality encoder representations from transformers,” in EMNLP/IJCNLP (1) . Association for Computational Linguistics, 2019, pp. 5099–5110
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
C. Sun, A. Myers, C. Vondrick, K. Murphy, and C. Schmid, “Videobert: A joint model for video and language representation learning,” in ICCV . IEEE, 2019, pp. 7463–7472
2019
Earlier work this paper cites.
N. Reimers and I. Gurevych, “Sentence-bert: Sentence embeddings using siamese bert-networks,” in EMNLP/IJCNLP (1) . Association for Computational Linguistics, 2019, pp. 3980–3990
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
C. Lynch, M. Khansari, T. Xiao, V. Kumar, J. Tompson, S. Levine, and P. Sermanet, “Learning latent plans from play,” in CoRL , vol. 100. PMLR, 2019, pp. 1113–1132
2019
Earlier work this paper cites.
Z. Dai, Z. Yang, Y. Yang, J. G. Carbonell, Q. V. Le, and R. Salakhutdinov, “Transformer-xl: Attentive language models beyond a fixed-length context,” in ACL (1) . Association for Computational Linguistics, 2019, pp. 2978–2988
2019
Earlier work this paper cites.
D. Hafner, T. P. Lillicrap, J. Ba, and M. Norouzi, “Dream to control: Learning behaviors by latent imagination,” in ICLR , 2020
2020
Earlier work this paper cites.
F. Xia, W. B. Shen, C. Li, P. Kasimbeg, M. Tchapmi, A. Toshev, R. Martín-Martín, and S. Savarese, “Interactive gibson benchmark: A benchmark for interactive navigation in cluttered environments,” IEEE Robotics Autom. Lett. , vol. 5, no. 2, pp. 713–720, 2020
2020
Earlier work this paper cites.
F. Xiang, Y. Qin, K. Mo, Y. Xia, H. Zhu, F. Liu, M. Liu, H. Jiang, Y. Yuan, H. Wang, L. Yi, A. X. Chang, L. J. Guibas, and H. Su, “SAPIEN: A simulated part-based interactive environment,” in CVPR . Computer Vision Foundation / IEEE, 2020, pp. 11 094–11 104
2020
Earlier work this paper cites.
S. James, Z. Ma, D. R. Arrojo, and A. J. Davison, “Rlbench: The robot learning benchmark & learning environment,” IEEE Robotics Autom. Lett. , vol. 5, no. 2, pp. 3019–3026, 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
M. Shridhar, J. Thomason, D. Gordon, Y. Bisk, W. Han, R. Mottaghi, L. Zettlemoyer, and D. Fox, “ALFRED: A benchmark for interpreting grounded instructions for everyday tasks,” in CVPR . Computer Vision Foundation / IEEE, 2020, pp. 10 737–10 746
2020
Earlier work this paper cites.
R. Zhang, A. Saran, B. Liu, Y. Zhu, S. Guo, S. Niekum, D. Ballard, and M. Hayhoe, “Human gaze assisted artificial intelligence: A review,” in IJCAI-20 , 7 2020, pp. 4951–4958, survey track
2020
Earlier work this paper cites.
A. Khan, A. Sohail, U. Zahoora, and A. S. Qureshi, “A survey of the recent architectures of deep convolutional neural networks,” Artif. Intell. Rev. , vol. 53, no. 8, pp. 5455–5516, 2020
2020
Earlier work this paper cites.
N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko, “End-to-end object detection with transformers,” in ECCV (1) , ser. Lecture Notes in Computer Science, vol. 12346. Springer, 2020, pp. 213–229
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
Z. Lan, M. Chen, S. Goodman, K. Gimpel, P. Sharma, and R. Soricut, “ALBERT: A lite BERT for self-supervised learning of language representations,” in ICLR , 2020
2020
Earlier work this paper cites.
K. Clark, M. Luong, Q. V. Le, and C. D. Manning, “ELECTRA: pre-training text encoders as discriminators rather than generators,” in ICLR , 2020
2020
Earlier work this paper cites.
M. Lewis, Y. Liu, N. Goyal, M. Ghazvininejad, A. Mohamed, O. Levy, V. Stoyanov, and L. Zettlemoyer, “BART: denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension,” in ACL . Association for Computational Linguistics, 2020, pp. 7871–7880
2020
Earlier work this paper cites.
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu, “Exploring the limits of transfer learning with a unified text-to-text transformer,” J. Mach. Learn. Res. , vol. 21, pp. 140:1–140:67, 2020
2020
Earlier work this paper cites.
A. Kumar, A. Zhou, G. Tucker, and S. Levine, “Conservative q-learning for offline reinforcement learning,” in NeurIPS , 2020
2020
Earlier work this paper cites.
Y. Rong, Y. Bian, T. Xu, W. Xie, Y. Wei, W. Huang, and J. Huang, “Self-supervised graph transformer on large-scale molecular data,” in NeurIPS , 2020
2020
Earlier work this paper cites.
F. Fuchs, D. E. Worrall, V. Fischer, and M. Welling, “Se(3)-transformers: 3d roto-translation equivariant attention networks,” in NeurIPS , 2020
2020
Earlier work this paper cites.
W. Zhong, J. Xu, D. Tang, Z. Xu, N. Duan, M. Zhou, J. Wang, and J. Yin, “Reasoning over semantic-level graph for fact checking,” in ACL . Association for Computational Linguistics, 2020, pp. 6170–6180
2020
Earlier work this paper cites.
W. Su, X. Zhu, Y. Cao, B. Li, L. Lu, F. Wei, and J. Dai, “VL-BERT: pre-training of generic visual-linguistic representations,” in ICLR , 2020
2020
Earlier work this paper cites.
Y. Chen, L. Li, L. Yu, A. E. Kholy, F. Ahmed, Z. Gan, Y. Cheng, and J. Liu, “UNITER: universal image-text representation learning,” in ECCV (30) , ser. Lecture Notes in Computer Science, vol. 12375. Springer, 2020, pp. 104–120
2020
Earlier work this paper cites.
Q. Xie, M. Luong, E. H. Hovy, and Q. V. Le, “Self-training with noisy student improves imagenet classification,” in CVPR . Computer Vision Foundation / IEEE, 2020, pp. 10 684–10 695
2020
Earlier work this paper cites.
E. Wijmans, A. Kadian, A. Morcos, S. Lee, I. Essa, D. Parikh, M. Savva, and D. Batra, “DD-PPO: learning near-perfect pointgoal navigators from 2.5 billion frames,” in ICLR , 2020
2020
Earlier work this paper cites.
J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” in NeurIPS , 2020
2020
Earlier work this paper cites.
A. Zeng, P. Florence, J. Tompson, S. Welker, J. Chien, M. Attarian, T. Armstrong, I. Krasin, D. Duong, V. Sindhwani, and J. Lee, “Transporter networks: Rearranging the visual world for robotic manipulation,” in CoRL , vol. 155. PMLR, 2020, pp. 726–747
2020
Earlier work this paper cites.
L. Chen, K. Lu, A. Rajeswaran, K. Lee, A. Grover, M. Laskin, P. Abbeel, A. Srinivas, and I. Mordatch, “Decision transformer: Reinforcement learning via sequence modeling,” in NeurIPS , 2021, pp. 15 084–15 097
2021
Earlier work this paper cites.
M. Janner, Q. Li, and S. Levine, “Offline reinforcement learning as one big sequence modeling problem,” in NeurIPS , 2021, pp. 1273–1286
2021
Earlier work this paper cites.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever, “Learning transferable visual models from natural language supervision,” in ICML , vol. 139. PMLR, 2021, pp. 8748–8763
2021
Earlier work this paper cites.
M. Shridhar, L. Manuelli, and D. Fox, “Cliport: What and where pathways for robotic manipulation,” in CoRL , vol. 164. PMLR, 2021, pp. 894–906
2021
Earlier work this paper cites.
D. Hafner, T. P. Lillicrap, M. Norouzi, and J. Ba, “Mastering atari with discrete world models,” in ICLR , 2021
2021
Earlier work this paper cites.
E. Jang, A. Irpan, M. Khansari, D. Kappler, F. Ebert, C. Lynch, S. Levine, and C. Finn, “BC-Z: zero-shot task generalization with robotic imitation learning,” in CoRL , vol. 164. PMLR, 2021, pp. 991–1002
2021
Earlier work this paper cites.
C. Lynch and P. Sermanet, “Language conditioned imitation learning over unstructured data,” in RSS , 2021
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
S. Srivastava, C. Li, M. Lingelbach, R. Martín-Martín, F. Xia, K. E. Vainio, Z. Lian, C. Gokmen, S. Buch, C. K. Liu, S. Savarese, H. Gweon, J. Wu, and L. Fei-Fei, “BEHAVIOR: benchmark for everyday household activities in virtual, interactive, and ecological environments,” in CoRL , vol. 164. PMLR, 2021, pp. 477–490
2021
Earlier work this paper cites.
C. Li, F. Xia, R. Martín-Martín, M. Lingelbach, S. Srivastava, B. Shen, K. E. Vainio, C. Gokmen, G. Dharan, T. Jain, A. Kurenkov, C. K. Liu, H. Gweon, J. Wu, L. Fei-Fei, and S. Savarese, “igibson 2.0: Object-centric simulation for robot learning of everyday household tasks,” in CoRL , vol. 164. PMLR, 2021, pp. 455–465
2021
Earlier work this paper cites.
B. Shen, F. Xia, C. Li, R. Martín-Martín, L. Fan, G. Wang, C. Pérez-D’Arpino, S. Buch, S. Srivastava, L. Tchapmi, M. Tchapmi, K. Vainio, J. Wong, L. Fei-Fei, and S. Savarese, “igibson 1.0: A simulation environment for interactive tasks in large realistic scenes,” in IROS . IEEE, 2021, pp. 7520–7527
2021
Earlier work this paper cites.
K. Ehsani, W. Han, A. Herrasti, E. VanderBilt, L. Weihs, E. Kolve, A. Kembhavi, and R. Mottaghi, “Manipulathor: A framework for visual object manipulation,” in CVPR . Computer Vision Foundation / IEEE, 2021, pp. 4497–4506
2021
Earlier work this paper cites.
L. Weihs, M. Deitke, A. Kembhavi, and R. Mottaghi, “Visual room rearrangement,” in CVPR . Computer Vision Foundation / IEEE, 2021, pp. 5922–5931
2021
Earlier work this paper cites.
C. Gan, J. Schwartz, S. Alter, D. Mrowca, M. Schrimpf, J. Traer, J. D. Freitas, J. Kubilius, A. Bhandwaldar, N. Haber, M. Sano, K. Kim, E. Wang, M. Lingelbach, A. Curtis, K. T. Feigelis, D. Bear, D. Gutfreund, D. D. Cox, A. Torralba, J. J. DiCarlo, J. Tenenbaum, J. H. McDermott, and D. Yamins, “Threedworld: A platform for interactive multi-modal physical simulation,” in NeurIPS Datasets and Benchmarks , 2021
2021
Earlier work this paper cites.
A. Szot, A. Clegg, E. Undersander, E. Wijmans, Y. Zhao, J. M. Turner, N. Maestre, M. Mukadam, D. S. Chaplot, O. Maksymets, A. Gokaslan, V. Vondrus, S. Dharur, F. Meier, W. Galuba, A. X. Chang, Z. Kira, V. Koltun, J. Malik, M. Savva, and D. Batra, “Habitat 2.0: Training home assistants to rearrange their habitat,” in NeurIPS , 2021, pp. 251–266
2021
Earlier work this paper cites.
A. Saran, R. Zhang, E. S. Short, and S. Niekum, “Efficiently guiding imitation learning agents with human gaze,” in AAMAS . ACM, 2021, pp. 1109–1117
2021
Earlier work this paper cites.
S. Guo, R. Zhang, B. Liu, Y. Zhu, D. H. Ballard, M. M. Hayhoe, and P. Stone, “Machine versus human attention in deep reinforcement learning tasks,” in NeurIPS , 2021, pp. 25 370–25 385
2021
Earlier work this paper cites.
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An image is worth 16x16 words: Transformers for image recognition at scale,” in ICLR , 2021
2021
Earlier work this paper cites.
R. Strudel, R. G. Pinel, I. Laptev, and C. Schmid, “Segmenter: Transformer for semantic segmentation,” in ICCV . IEEE, 2021, pp. 7242–7252
2021
Earlier work this paper cites.
Y. Guo, H. Wang, Q. Hu, H. Liu, L. Liu, and M. Bennamoun, “Deep learning for 3d point clouds: A survey,” IEEE Trans. Pattern Anal. Mach. Intell. , vol. 43, no. 12, pp. 4338–4364, 2021
2021
Earlier work this paper cites.
D. W. Otter, J. R. Medina, and J. K. Kalita, “A survey of the usages of deep learning for natural language processing,” IEEE Trans. Neural Networks Learn. Syst. , vol. 32, no. 2, pp. 604–624, 2021
2021
Earlier work this paper cites.
Z. Wu, S. Pan, F. Chen, G. Long, C. Zhang, and P. S. Yu, “A comprehensive survey on graph neural networks,” IEEE Trans. Neural Networks Learn. Syst. , vol. 32, no. 1, pp. 4–24, 2021
2021
Earlier work this paper cites.
V. G. Satorras, E. Hoogeboom, and M. Welling, “E(n) equivariant graph neural networks,” in ICML , vol. 139. PMLR, 2021, pp. 9323–9332
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
W. Kim, B. Son, and I. Kim, “Vilt: Vision-and-language transformer without convolution or region supervision,” in ICML , vol. 139. PMLR, 2021, pp. 5583–5594
2021
Earlier work this paper cites.
Z. Dai, H. Liu, Q. V. Le, and M. Tan, “Coatnet: Marrying convolution and attention for all data sizes,” in NeurIPS , 2021, pp. 3965–3977
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
C. Jia, Y. Yang, Y. Xia, Y. Chen, Z. Parekh, H. Pham, Q. V. Le, Y. Sung, Z. Li, and T. Duerig, “Scaling up visual and vision-language representation learning with noisy text supervision,” in ICML , vol. 139. PMLR, 2021, pp. 4904–4916
2021
Earlier work this paper cites.
J. Li, R. R. Selvaraju, A. Gotmare, S. R. Joty, C. Xiong, and S. C. Hoi, “Align before fuse: Vision and language representation learning with momentum distillation,” in NeurIPS , 2021, pp. 9694–9705
2021
Earlier work this paper cites.
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” in ICCV . IEEE, 2021, pp. 9992–10 002
2021
Earlier work this paper cites.
H. Wu, B. Xiao, N. Codella, M. Liu, X. Dai, L. Yuan, and L. Zhang, “Cvt: Introducing convolutions to vision transformers,” in ICCV . IEEE, 2021, pp. 22–31
2021
Earlier work this paper cites.
A. Brock, S. De, S. L. Smith, and K. Simonyan, “High-performance large-scale image recognition without normalization,” in ICML , vol. 139. PMLR, 2021, pp. 1059–1071
2021
Earlier work this paper cites.
L. Xue, N. Constant, A. Roberts, M. Kale, R. Al-Rfou, A. Siddhant, A. Barua, and C. Raffel, “mt5: A massively multilingual pre-trained text-to-text transformer,” in NAACL-HLT . Association for Computational Linguistics, 2021, pp. 483–498
2021
Earlier work this paper cites.
M. Tsimpoukelli, J. Menick, S. Cabi, S. M. A. Eslami, O. Vinyals, and F. Hill, “Multimodal few-shot learning with frozen language models,” in NeurIPS , 2021, pp. 200–212
2021
Earlier work this paper cites.
P. Esser, R. Rombach, and B. Ommer, “Taming transformers for high-resolution image synthesis,” in CVPR . Computer Vision Foundation / IEEE, 2021, pp. 12 873–12 883
2021
Earlier work this paper cites.
A. Kamath, M. Singh, Y. LeCun, G. Synnaeve, I. Misra, and N. Carion, “MDETR - modulated detection for end-to-end multi-modal understanding,” in ICCV . IEEE, 2021, pp. 1760–1770
2021
Earlier work this paper cites.
V. Blukis, C. Paxton, D. Fox, A. Garg, and Y. Artzi, “A persistent spatial semantic representation for high-level natural language instruction execution,” in CoRL , vol. 164. PMLR, 2021, pp. 706–717
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
M. Bhardwaj, B. Sundaralingam, A. Mousavian, N. D. Ratliff, D. Fox, F. Ramos, and B. Boots, “STORM: an integrated framework for fast joint-space model-predictive control for reactive manipulation,” in CoRL , vol. 164. PMLR, 2021, pp. 750–759
2021
Earlier work this paper cites.
Y. Song, J. Sohl-Dickstein, D. P. Kingma, A. Kumar, S. Ermon, and B. Poole, “Score-based generative modeling through stochastic differential equations,” in ICLR , 2021
2021
Earlier work this paper cites.
J. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds, R. Ring, E. Rutherford, S. Cabi, T. Han, Z. Gong, S. Samangooei, M. Monteiro, J. L. Menick, S. Borgeaud, A. Brock, A. Nematzadeh, S. Sharifzadeh, M. Binkowski, R. Barreira, O. Vinyals, A. Zisserman, and K. Simonyan, “Flamingo: a visual language model for few-shot learning,” in NeurIPS , 2022
2022
Earlier work this paper cites.
W. Huang, F. Xia, T. Xiao, H. Chan, J. Liang, P. Florence, A. Zeng, J. Tompson, I. Mordatch, Y. Chebotar, P. Sermanet, T. Jackson, N. Brown, L. Luu, S. Levine, K. Hausman, and B. Ichter, “Inner monologue: Embodied reasoning through planning with language models,” in CoRL , vol. 205. PMLR, 2022, pp. 1769–1782
2022
Earlier work this paper cites.
B. Ichter, A. Brohan, Y. Chebotar, C. Finn, K. Hausman, A. Herzog, D. Ho, J. Ibarz, A. Irpan, E. Jang, R. Julian, D. Kalashnikov, S. Levine, Y. Lu, C. Parada, K. Rao, P. Sermanet, A. Toshev, V. Vanhoucke, F. Xia, T. Xiao, P. Xu, M. Yan, N. Brown, M. Ahn, O. Cortes, N. Sievers, C. Tan, S. Xu, D. Reyes, J. Rettinghouse, J. Quiambao, P. Pastor, L. Luu, K. Lee, Y. Kuang, S. Jesmonth, N. J. Joshi, K. Jeffrey, R. J. Ruano, J. Hsu, K. Gopalakrishnan, B. David, A. Zeng, and C. K. Fu, “Do as I can, not as I say: Grounding language in robotic affordances,” in CoRL , vol. 205. PMLR, 2022, pp. 287–318
2022
Earlier work this paper cites.
S. E. Reed, K. Zolna, E. Parisotto, S. G. Colmenarejo, A. Novikov, G. Barth-Maron, M. Gimenez, Y. Sulsky, J. Kay, J. T. Springenberg, T. Eccles, J. Bruce, A. Razavi, A. Edwards, N. Heess, Y. Chen, R. Hadsell, O. Vinyals, M. Bordbar, and N. de Freitas, “A generalist agent,” Trans. Mach. Learn. Res. , vol. 2022, 2022
2022
Earlier work this paper cites.
S. Nair, A. Rajeswaran, V. Kumar, C. Finn, and A. Gupta, “R3M: A universal visual representation for robot manipulation,” in CoRL , vol. 205. PMLR, 2022, pp. 892–909
2022
Earlier work this paper cites.
I. Radosavovic, T. Xiao, S. James, P. Abbeel, J. Malik, and T. Darrell, “Real-world robot learning with masked visual pre-training,” in CoRL , vol. 205. PMLR, 2022, pp. 416–426
2022
Earlier work this paper cites.
A. Khandelwal, L. Weihs, R. Mottaghi, and A. Kembhavi, “Simple but effective: CLIP embeddings for embodied AI,” in CVPR . IEEE, 2022, pp. 14 809–14 818
2022
Earlier work this paper cites.
S. Parisi, A. Rajeswaran, S. Purushwalkam, and A. Gupta, “The unsurprising effectiveness of pre-trained vision models for control,” in ICML , vol. 162. PMLR, 2022, pp. 17 359–17 371
2022
Earlier work this paper cites.
Y. LeCun, “A path towards autonomous machine intelligence,” 2022. [Online]. Available: https://openreview.net/pdf?id=BZ5a1r-kVsf
2022
Earlier work this paper cites.
F. Liu, H. Liu, A. Grover, and P. Abbeel, “Masked autoencoding for scalable and generalizable decision making,” in NeurIPS , 2022
2022
Earlier work this paper cites.
B. Baker, I. Akkaya, P. Zhokhov, J. Huizinga, J. Tang, A. Ecoffet, B. Houghton, R. Sampedro, and J. Clune, “Video pretraining (VPT): learning to act by watching unlabeled online videos,” in NeurIPS , 2022
2022
Earlier work this paper cites.
P. Wu, A. Escontrela, D. Hafner, P. Abbeel, and K. Goldberg, “Daydreamer: World models for physical robot learning,” in CoRL , vol. 205. PMLR, 2022, pp. 2226–2240
2022
Earlier work this paper cites.
T. Kojima, S. S. Gu, M. Reid, Y. Matsuo, and Y. Iwasawa, “Large language models are zero-shot reasoners,” in NeurIPS , 2022
2022
Earlier work this paper cites.
J. Wei, X. Wang, D. Schuurmans, M. Bosma, B. Ichter, F. Xia, E. H. Chi, Q. V. Le, and D. Zhou, “Chain-of-thought prompting elicits reasoning in large language models,” in NeurIPS , 2022
2022
Earlier work this paper cites.
O. Mees, L. Hermann, and W. Burgard, “What matters in language conditioned robotic imitation learning over unstructured data,” IEEE Robotics Autom. Lett. , vol. 7, no. 4, pp. 11 205–11 212, 2022
2022
Earlier work this paper cites.
P. Sharma, B. Sundaralingam, V. Blukis, C. Paxton, T. Hermans, A. Torralba, J. Andreas, and D. Fox, “Correcting robot plans with natural language feedback,” in RSS , 2022
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
P. Guhur, S. Chen, R. G. Pinel, M. Tapaswi, I. Laptev, and C. Schmid, “Instruction-driven history-aware policies for robotic manipulations,” in CoRL , vol. 205. PMLR, 2022, pp. 175–187
2022
Earlier work this paper cites.
M. Shridhar, L. Manuelli, and D. Fox, “Perceiver-actor: A multi-task transformer for robotic manipulation,” in CoRL , vol. 205. PMLR, 2022, pp. 785–799
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
W. Huang, P. Abbeel, D. Pathak, and I. Mordatch, “Language models as zero-shot planners: Extracting actionable knowledge for embodied agents,” in ICML , vol. 162. PMLR, 2022, pp. 9118–9147
2022
Earlier work this paper cites.
P. Sharma, A. Torralba, and J. Andreas, “Skill induction and planning with latent language,” in ACL (1) , 2022, pp. 1713–1726
2022
Earlier work this paper cites.
S. Li, X. Puig, C. Paxton, Y. Du, C. Wang, L. Fan, T. Chen, D. Huang, E. Akyürek, A. Anandkumar, J. Andreas, I. Mordatch, A. Torralba, and Y. Zhu, “Pre-trained language models for interactive decision-making,” in NeurIPS , 2022
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
F. Ebert, Y. Yang, K. Schmeckpeper, B. Bucher, G. Georgakis, K. Daniilidis, C. Finn, and S. Levine, “Bridge data: Boosting generalization of robotic skills with cross-domain datasets,” in RSS , 2022
2022
Earlier work this paper cites.
O. Mees, L. Hermann, E. Rosete-Beas, and W. Burgard, “CALVIN: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks,” IEEE Robotics Autom. Lett. , vol. 7, no. 3, pp. 7327–7334, 2022
2022
Earlier work this paper cites.
B. Jia, T. Lei, S. Zhu, and S. Huang, “Egotaskqa: Understanding human tasks in egocentric videos,” in NeurIPS , 2022
2022
Earlier work this paper cites.
M. S. Laursen, J. S. Pedersen, S. A. Just, T. R. Savarimuthu, B. Blomholt, J. K. H. Andersen, and P. J. Vinholt, “Factors facilitating the acceptance of diagnostic robots in healthcare: A survey,” in ICHI . IEEE, 2022, pp. 442–448
2022
Earlier work this paper cites.
S. H. Khan, M. Naseer, M. Hayat, S. W. Zamir, F. S. Khan, and M. Shah, “Transformers in vision: A survey,” ACM Comput. Surv. , vol. 54, no. 10s, pp. 200:1–200:41, 2022
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. L. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, J. Schulman, J. Hilton, F. Kelton, L. Miller, M. Simens, A. Askell, P. Welinder, P. F. Christiano, J. Leike, and R. Lowe, “Training language models to follow instructions with human feedback,” in NeurIPS , 2022
2022
Earlier work this paper cites.
J. Wei, M. Bosma, V. Y. Zhao, K. Guu, A. W. Yu, B. Lester, N. Du, A. M. Dai, and Q. V. Le, “Finetuned language models are zero-shot learners,” in ICLR , 2022
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
W. Du, H. Zhang, Y. Du, Q. Meng, W. Chen, N. Zheng, B. Shao, and T. Liu, “SE(3) equivariant graph neural networks with complete local frames,” in ICML , vol. 162. PMLR, 2022, pp. 5583–5608
2022
Earlier work this paper cites.
Z. Wang, J. Yu, A. W. Yu, Z. Dai, Y. Tsvetkov, and Y. Cao, “Simvlm: Simple visual language model pretraining with weak supervision,” in ICLR , 2022
2022
Earlier work this paper cites.
J. Wang, Z. Yang, X. Hu, L. Li, K. Lin, Z. Gan, Z. Liu, C. Liu, and L. Wang, “GIT: A generative image-to-text transformer for vision and language,” Trans. Mach. Learn. Res. , vol. 2022, 2022
2022
Earlier work this paper cites.
L. Yao, R. Huang, L. Hou, G. Lu, M. Niu, H. Xu, X. Liang, Z. Li, X. Jiang, and C. Xu, “FILIP: fine-grained interactive language-image pre-training,” in ICLR , 2022
2022
Earlier work this paper cites.
A. Singh, R. Hu, V. Goswami, G. Couairon, W. Galuba, M. Rohrbach, and D. Kiela, “FLAVA: A foundational language and vision alignment model,” in CVPR . IEEE, 2022, pp. 15 617–15 629
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
X. Zhai, A. Kolesnikov, N. Houlsby, and L. Beyer, “Scaling vision transformers,” in CVPR . IEEE, 2022, pp. 1204–1213
2022
Earlier work this paper cites.
Z. Du, Y. Qian, X. Liu, M. Ding, J. Qiu, Z. Yang, and J. Tang, “GLM: general language model pretraining with autoregressive blank infilling,” in ACL (1) . Association for Computational Linguistics, 2022, pp. 320–335
2022
Earlier work this paper cites.
H. Bao, W. Wang, L. Dong, Q. Liu, O. K. Mohammed, K. Aggarwal, S. Som, S. Piao, and F. Wei, “Vlmo: Unified vision-language pre-training with mixture-of-modality-experts,” in NeurIPS , 2022
2022
Earlier work this paper cites.
X. Zhai, X. Wang, B. Mustafa, A. Steiner, D. Keysers, A. Kolesnikov, and L. Beyer, “Lit: Zero-shot transfer with locked-image text tuning,” in CVPR . IEEE, 2022, pp. 18 102–18 112
2022
Earlier work this paper cites.
J. Yu, Z. Wang, V. Vasudevan, L. Yeung, M. Seyedhosseini, and Y. Wu, “Coca: Contrastive captioners are image-text foundation models,” Trans. Mach. Learn. Res. , vol. 2022, 2022
2022
Earlier work this paper cites.
P. Wang, A. Yang, R. Men, J. Lin, S. Bai, Z. Li, J. Ma, C. Zhou, J. Zhou, and H. Yang, “OFA: unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework,” in ICML , vol. 162. PMLR, 2022, pp. 23 318–23 340
2022
Earlier work this paper cites.
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen, “Lora: Low-rank adaptation of large language models,” in ICLR , 2022
2022
Earlier work this paper cites.
C. Zheng, T. Vuong, J. Cai, and D. Phung, “Movq: Modulating quantized vectors for high-fidelity image generation,” in NeurIPS , 2022
2022
Earlier work this paper cites.
X. Gu, T. Lin, W. Kuo, and Y. Cui, “Open-vocabulary object detection via vision and language knowledge distillation,” in ICLR , 2022
2022
Earlier work this paper cites.
K. He, X. Chen, S. Xie, Y. Li, P. Dollár, and R. B. Girshick, “Masked autoencoders are scalable vision learners,” in CVPR . IEEE, 2022, pp. 15 979–15 988
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
Z. Liu, H. Mao, C. Wu, C. Feichtenhofer, T. Darrell, and S. Xie, “A convnet for the 2020s,” in CVPR . IEEE, 2022, pp. 11 966–11 976
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
M. S. M. Sajjadi, D. Duckworth, A. Mahendran, S. van Steenkiste, F. Pavetic, M. Lucic, L. J. Guibas, K. Greff, and T. Kipf, “Object scene representation transformer,” in NeurIPS , 2022
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
A. Jaegle, S. Borgeaud, J. Alayrac, C. Doersch, C. Ionescu, D. Ding, S. Koppula, D. Zoran, A. Brock, E. Shelhamer, O. J. Hénaff, M. M. Botvinick, A. Zisserman, O. Vinyals, and J. Carreira, “Perceiver IO: A general architecture for structured inputs & outputs,” in ICLR , 2022
2022
Earlier work this paper cites.
Y. Seo, D. Hafner, H. Liu, F. Liu, S. James, K. Lee, and P. Abbeel, “Masked world models for visual control,” in CoRL , vol. 205. PMLR, 2022, pp. 1332–1344
2022
Earlier work this paper cites.
M. Pan, X. Zhu, Y. Wang, and X. Yang, “Iso-dream: Isolating and leveraging noncontrollable visual dynamics in world models,” in NeurIPS , 2022
2022
Earlier work this paper cites.
N. M. Shafiullah, Z. J. Cui, A. Altanzaya, and L. Pinto, “Behavior transformers: Cloning $k$ modes with one stone,” in NeurIPS , 2022
2022
Earlier work this paper cites.
2023
Earlier work this paper cites.
J. Li, D. Li, S. Savarese, and S. C. H. Hoi, “BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models,” in ICML , vol. 202. PMLR, 2023, pp. 19 730–19 742
2023
Earlier work this paper cites.
H. Liu, C. Li, Q. Wu, and Y. J. Lee, “Visual instruction tuning,” CoRR , vol. abs/2304.08485, 2023
2023
Earlier work this paper cites.
D. Driess, F. Xia, M. S. M. Sajjadi, C. Lynch, A. Chowdhery, B. Ichter, A. Wahid, J. Tompson, Q. Vuong, T. Yu, W. Huang, Y. Chebotar, P. Sermanet, D. Duckworth, S. Levine, V. Vanhoucke, K. Hausman, M. Toussaint, K. Greff, A. Zeng, I. Mordatch, and P. Florence, “Palm-e: An embodied multimodal language model,” in ICML , vol. 202. PMLR, 2023, pp. 8469–8488
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
A. Hiranaka, M. Hwang, S. Lee, C. Wang, L. Fei-Fei, J. Wu, and R. Zhang, “Primitive skill-based robot learning from human evaluative feedback,” in IROS , 2023, pp. 7817–7824
2023
Earlier work this paper cites.
N. Shinn, F. Cassano, A. Gopinath, K. Narasimhan, and S. Yao, “Reflexion: language agents with verbal reinforcement learning,” in NeurIPS , 2023
2023
Earlier work this paper cites.
Y. J. Ma, S. Sodhani, D. Jayaraman, O. Bastani, V. Kumar, and A. Zhang, “VIP: towards universal visual reward and representation via value-implicit pre-training,” in ICLR , 2023
2023
Earlier work this paper cites.
I. Radosavovic, B. Shi, L. Fu, K. Goldberg, T. Darrell, and J. Malik, “Robot learning with sensorimotor pre-training,” in CoRL , vol. 229. PMLR, 2023, pp. 683–693
2023
Earlier work this paper cites.
S. Y. Gadre, M. Wortsman, G. Ilharco, L. Schmidt, and S. Song, “Cows on pasture: Baselines and benchmarks for language-driven zero-shot object navigation,” in CVPR . IEEE, 2023, pp. 23 171–23 181
2023
Earlier work this paper cites.
A. Majumdar, K. Yadav, S. Arnaud, Y. J. Ma, C. Chen, S. Silwal, A. Jain, V. Berges, T. Wu, J. Vakil, P. Abbeel, J. Malik, D. Batra, Y. Lin, O. Maksymets, A. Rajeswaran, and F. Meier, “Where are we in the search for an artificial visual cortex for embodied intelligence?” in NeurIPS , 2023
2023
Earlier work this paper cites.
S. Karamcheti, S. Nair, A. S. Chen, T. Kollar, C. Finn, D. Sadigh, and P. Liang, “Language-driven representation learning for robotics,” in RSS , 2023
2023
Earlier work this paper cites.
M. Assran, Q. Duval, I. Misra, P. Bojanowski, P. Vincent, M. G. Rabbat, Y. LeCun, and N. Ballas, “Self-supervised learning from images with a joint-embedding predictive architecture,” in CVPR . IEEE, 2023, pp. 15 619–15 629
2023
Earlier work this paper cites.
W. Shen, G. Yang, A. Yu, J. Wong, L. P. Kaelbling, and P. Isola, “Distilled feature fields enable few-shot language-guided manipulation,” in CoRL , vol. 229. PMLR, 2023, pp. 405–424
2023
Earlier work this paper cites.
Y. Hong, H. Zhen, P. Chen, S. Zheng, Y. Du, Z. Chen, and C. Gan, “3d-llm: Injecting the 3d world into large language models,” in NeurIPS , 2023
2023
Earlier work this paper cites.
B. Kerbl, G. Kopanas, T. Leimkühler, and G. Drettakis, “3d gaussian splatting for real-time radiance field rendering,” ACM Trans. Graph. , vol. 42, no. 4, pp. 139:1–139:14, 2023
2023
Earlier work this paper cites.
A. Thankaraj and L. Pinto, “That sounds right: Auditory self-supervision for dynamic robot manipulation,” in CoRL , vol. 229. PMLR, 2023, pp. 1036–1049
2023
Earlier work this paper cites.
Y. Jing, X. Zhu, X. Liu, Q. Sima, T. Yang, Y. Feng, and T. Kong, “Exploring visual pre-training for robot manipulation: Datasets, models and methods,” in IROS , 2023, pp. 11 390–11 395
2023
Cited alongside, same era.
Y. Sun, S. Ma, R. Madaan, R. Bonatti, F. Huang, and A. Kapoor, “SMART: self-supervised multi-task pretraining with control transformers,” in ICLR , 2023
2023
Cited alongside, same era.
R. Bonatti, S. Vemprala, S. Ma, F. Frujeri, S. Chen, and A. Kapoor, “PACT: perception-action causal transformer for autoregressive robotics pre-training,” in IROS , 2023, pp. 3621–3627
2023
Cited alongside, same era.
2023
Cited alongside, same era.
V. Micheli, E. Alonso, and F. Fleuret, “Transformers are sample-efficient world models,” in ICLR , 2023
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
C. W. Yang Zhang and, O. Lu, Y. Zhao, Y. Ge, Z. S. 0001, X. L. 0001, C. Z. 0012, C. Bai, and X. L. 0001, “Align-then-steer: Adapting the vision-language action models through unified latent guidance,” 2025. [Online]. Available: https://dblp.org/rec/journals/corr/abs-2509-02055
2025
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2023
Cited alongside, same era.
J. Robine, M. Höftmann, T. Uelwer, and S. Harmeling, “Transformer-based world models are happy with 100k interactions,” in ICLR , 2023
2023
Cited alongside, same era.
K. Nottingham, P. Ammanabrolu, A. Suhr, Y. Choi, H. Hajishirzi, S. Singh, and R. Fox, “Do embodied agents dream of pixelated sheep: Embodied decision making using language guided world modelling,” in ICML , vol. 202. PMLR, 2023, pp. 26 311–26 325
2023
Cited alongside, same era.
Z. Song, Y. Zhang, and I. King, “No change, no gain: Empowering graph neural networks with expected model change maximization for active learning,” in NeurIPS , 2023
2023
Cited alongside, same era.
Y. Ma, Z. Song, X. Hu, J. Li, Y. Zhang, and I. King, “Graph component contrastive learning for concept relatedness estimation,” in AAAI . AAAI Press, 2023, pp. 13 362–13 370
2023
Cited alongside, same era.
L. Guan, K. Valmeekam, S. Sreedharan, and S. Kambhampati, “Leveraging pre-trained large language models to construct and utilize world models for model-based task planning,” in NeurIPS , 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
S. Hao, Y. Gu, H. Ma, J. J. Hong, Z. Wang, D. Z. Wang, and Z. Hu, “Reasoning with language model is planning with world model,” in EMNLP , 2023, pp. 8154–8173
2023
Cited alongside, same era.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
J. L. Lei Xiao and, J. Gao, F. Ye, Y. Jin, J. Qian, J. Zhang, Y. Wu, and X. Yu, “Ava-vla: Improving vision-language-action models with active visual attention,” 2025. [Online]. Available: https://dblp.org/rec/journals/corr/abs-2511-18960
2025
Closest in time.
L. L. Harris Song and, “Avi: Action from volumetric inference,” CoRR , vol. abs/2510.21746, 2025
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
A. Z. R. Asher J. Hancock and and A. Majumdar, “Run-time observation interventions make vision-language-action models, more visually robust,” in ICRA . IEEE, 2025, pp. 9499–9506
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
S. Wu, Y. Ji, Q. Li, Z. Zhang, Q. He, W. Xie, G. Zhang, B. Bayramli, Y. Ding, and H. Lu, “Dejavu: Towards experience feedback learning for embodied intelligence,” arXiv preprint , 2025
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
Z. X. Cheng Chi and, S. Feng, E. Cousineau, Y. Du, B. Burchfiel, R. Tedrake, and S. Song, “Diffusion policy: Visuomotor policy learning via action diffusion,” Int. J. Robotics Res. , vol. 44, no. 10-11, pp. 1684–1704, 2025
2025
Closest in time.
2025
Closest in time.
R. Liang, Y. Zheng, K. Zheng, T. Tan, J. Li, L. Mao, Z. Wang, G. Chen, H. Ye, J. Liu, J. Wang, and X. Zhan, “Dichotomous diffusion policy optimization,” arXiv preprint , 2025
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
P. H. Qi Sun and, P. T. Deep, V. Toh, U. Tan, D. Ghosal, and S. Poria, “Emma-x: An embodied multimodal action model with grounded chain of, thought and look-ahead spatial reasoning,” in ACL (1) . Association for Computational Linguistics, 2025, pp. 14 199–14 214
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
H. Chen, J. Liu, C. Gu, Z. Liu, R. Zhang, X. Li, and X. He, “Fast-in-slow: A dual-system vla model unifying fast manipulation within slow reasoning,” arXiv preprint , 2025
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
Z. Zhong, H. Yan, J. Li, X. Liu, X. Gong, T. Zhang, W. Song, J. Chen, X. Zheng, H. Wang, and H. Li, “Flowvla: Visual chain of thought-based motion reasoning for vision-language-action models,” arXiv preprint , 2025
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
H. N. Arjun Vaithilingam Sudhakar and, M. Reymond, M. Liu, J. Rajendran, and S. Chandar, “A generalist hanabi agent,” in ICLR . OpenReview.net, 2025
2025
Closest in time.
G. A. Team, “Gen-0: Embodied foundation models that scale with physical interaction,” arXiv preprint , 2025
2025
Closest in time.
2025
Closest in time.
P. D. Hongyin Zhang and, S. Lyu, Y. Peng, and D. Wang, “GEVRM: goal-expressive video generation model for robust visual, manipulation,” in ICLR . OpenReview.net, 2025
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
J. Jabbour, D.-K. Kim, M. Smith, J. Patrikar, R. Ghosal, Y. Wang, A. Agha, V. J. Reddi, and S. Omidshafiei, “Dont run with scissors,” arXiv preprint , 2025
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
Y. D. Yi Li and, J. Zhang, J. Jang, M. Memmel, C. R. Garrett, F. Ramos, D. Fox, A. Li, A. Gupta, and A. Goyal, “HAMSTER: hierarchical action models for open-world robot manipulation,” in ICLR . OpenReview.net, 2025
2025
Closest in time.
FIGURE, “Helix: A vision-language-action model for generalist humanoid control,” arXiv preprint , 2025
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
B. I. Lucy Xiaoyang Shi and, M. R. Equi, L. Ke, K. Pertsch, Q. Vuong, J. Tanner, A. Walling, H. Wang, N. Fusai, A. Li-Bell, D. Driess, L. Groom, S. Levine, and C. Finn, “Hi robot: Open-ended instruction following with hierarchical vision-language-action, models,” in ICML . OpenReview.net, 2025
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
J. Z. Yanjiang Guo and, X. Chen, X. Ji, Y. Wang, Y. Hu, and J. Chen, “Improving vision-language-action model with online reinforcement learning,” in ICRA . IEEE, 2025, pp. 15 665–15 672
2025
Closest in time.
2025
Closest in time.
W. Xu, L. Zhuang, and L. Shan, “Kv-efficient vla: A method to speed up vision language models with rnn-gated chunked kv cache,” arXiv preprint , 2025
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
Y. Z. Zhenyu Wu and, X. Xu, Z. Wang, and H. Yan, “Momanipvla: Transferring vision-language-action models for general, mobile manipulation,” in CVPR . Computer Vision Foundation / IEEE, 2025, pp. 1714–1723
2025
Closest in time.
2025
Closest in time.
W. S. Han Zhao and, D. Wang, X. Tong, P. Ding, X. Cheng, and Z. Ge, “More: Unlocking scalability in reinforcement learning for quadruped, vision-language-action models,” in ICRA . IEEE, 2025, pp. 11 212–11 218
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
F. L. Huang Huang and, L. Fu, T. Wu, M. Mukadam, J. Malik, K. Goldberg, and P. Abbeel, “OTTER: A vision-language-action model with text-aware visual feature, extraction,” in ICML . OpenReview.net, 2025
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
P. D. Xinyang Tong and, Y. Fan, D. Wang, W. Zhang, C. Cui, M. Sun, H. Zhao, H. Zhang, Y. Dang, S. Huang, and S. Lyu, “Quart-online: Latency-free multimodal large language model for quadruped, robot learning,” in ICRA . IEEE, 2025, pp. 9533–9539
2025
Closest in time.
2025
Closest in time.
L. W. Songming Liu and, B. Li, H. Tan, H. Chen, Z. Wang, K. Xu, H. Su, and J. Zhu, “RDT-1B: a diffusion foundation model for bimanual manipulation,” in ICLR . OpenReview.net, 2025
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
S. D. Kaustubh Sridhar and, D. Jayaraman, and I. Lee, “REGENT: A retrieval-augmented generalist agent that can act in-context, in new environments,” in ICLR . OpenReview.net, 2025
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
J. Z. Sombit Dey and, N. Nikolov, L. V. Gool, and D. P. Paudel, “Revla: Reverting visual domain limitation of robotic foundation models,” in ICRA . IEEE, 2025, pp. 8679–8686
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
J. W. Shunlei Li and, R. Dai, W. Ma, W. Y. Ng, Y. Hu, and Z. Li, “Robonurse-vla: Robotic scrub nurse system based on vision-language-action, model,” in IROS . IEEE, 2025, pp. 3986–3993
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
W. B. Tobias Jülg and and F. Walter, “Refined policy distillation: From VLA generalists to RL experts,” in IROS . IEEE, 2025, pp. 11 677–11 684
2025
Closest in time.
2025
Closest in time.
S. K. Soroush Nasiriany and, T. Ding, L. Smith, Y. Zhu, D. Driess, D. Sadigh, and T. Xiao, “Rt-affordance: Affordances are versatile intermediate representations, for robot manipulation,” in ICRA . IEEE, 2025, pp. 8249–8257
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
B. Zhang, Y. Zhang, J. Ji, Y. Lei, J. Dai, Y. Chen, and Y. Yang, “Safevla: Towards safety alignment of vision-language-action model via constrained learning,” arXiv preprint , 2025
2025
Closest in time.
Y. Z. Minjie Zhu and, J. Li, J. Wen, Z. Xu, N. Liu, R. Cheng, C. Shen, Y. Peng, F. Feng, and J. Tang, “Scaling diffusion policy in transformer to 1 billion parameters for, robotic manipulation,” in ICRA . IEEE, 2025, pp. 10 838–10 845
2025
Closest in time.
2025
Closest in time.
J. Z. Beichen Wang and, S. Dong, I. Fang, and C. Feng, “VLM see, robot do: Human demo video to robot action plan via vision, language model,” in IROS . IEEE, 2025, pp. 17 215–17 222
2025
Closest in time.
2025
Closest in time.
L. L. Kevin Qinghong Lin and, D. Gao, Z. Yang, S. Wu, Z. Bai, S. W. Lei, L. Wang, and M. Z. Shou, “Showui: One vision-language-action model for GUI visual agent,” in CVPR . Computer Vision Foundation / IEEE, 2025, pp. 19 498–19 508
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
W. X. Jianping Jiang and, Z. Lin, H. Zhang, T. Ren, Y. Gao, Z. Lin, Z. Cai, L. Yang, and Z. Liu, “SOLAMI: social vision-language-action modeling for immersive interaction, with 3d autonomous characters,” in CVPR . Computer Vision Foundation / IEEE, 2025, pp. 26 887–26 898
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
Z. W. Guanxing Lu and, C. Liu, J. Lu, and Y. Tang, “Thinkbot: Embodied instruction following with thought chain reasoning,” in ICLR . OpenReview.net, 2025
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
Y. Bai, Z. Wang, Y. Liu, K. Luo, Y. Wen, M. Dai, W. Chen, Z. Chen, L. Liu, G. Li, and L. Lin, “Learning to see and act: Task-aware virtual view exploration for robotic manipulation,” arXiv preprint , 2025
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
Y. G. Jianke Zhang and, Y. Hu, X. Chen, X. Zhu, and J. Chen, “UP-VLA: A unified understanding and prediction model for embodied, agent,” in ICML . OpenReview.net, 2025
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
S. M. Hyunki Seong and, H. Ahn, J. Kang, and D. H. Shim, “Vla-r: Vision-language action retrieval toward open-world end-to-end autonomous driving,” 2025. [Online]. Available: https://dblp.org/rec/journals/corr/abs-2511-12405
2025
Closest in time.
Z. Z. Angen Ye and, B. Wang, X. Wang, D. Zhang, and Z. Zhu, “Vla-r1: Enhancing reasoning in vision-language-action models,” 2025. [Online]. Available: https://dblp.org/rec/journals/corr/abs-2510-01623
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
P. D. Wei Zhao and, M. Zhang, Z. Gong, S. Bai, H. Zhao, and D. Wang, “VLAS: vision-language-action model with speech instructions for, customized robot manipulation,” in ICLR . OpenReview.net, 2025
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
M. W. Mingjie Xu and, Y. Zhao, J. C. L. Li, and W. Ou, “Llava-spacesgg: Visual instruct tuning for open-vocabulary scene graph, generation with enhanced spatial relations,” in WACV . IEEE, 2025, pp. 6362–6372
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
L. Magne, A. Awadalla, G. Wang, Y. Xu, J. Belofsky, F. Hu, J. Kim, L. Schmidt, G. Gkioxari, J. Kautz, Y. Yue, Y. Choi, Y. Zhu, and L. J. Fan, “Nitrogen: An open foundation model for generalist gaming agents,” arXiv preprint , 2026
2026
Closest in time.
S. Fujimoto, D. Meger, and D. Precup, “Off-policy deep reinforcement learning without exploration,” in ICML , vol. 97. PMLR, 2019, pp. 2052–2062
2062
Closest in time.
H. van Hasselt, A. Guez, and D. Silver, “Deep reinforcement learning with double q-learning,” in AAAI . AAAI Press, 2016, pp. 2094–2100
2094
Closest in time.