Fetching the paper…
Reading the bibliography…
Benefiting from language flexibility and compositionality, humans naturally intend to use language to command an embodied agent for complex tasks such as navigation and object manipulation.
1912
Earlier work this paper cites.
Şucan, I.A., Moll, M., Kavraki, L.E.: The Open Motion Planning Library. IEEE Robotics & Automation Magazine 19
2012
Earlier work this paper cites.
Rohmer, E., Singh, S.P.N., Freese, M.: V-rep: A versatile and scalable robot simulation framework pp. 1321–1326 (2013). https://doi.org/10.1109/IROS.2013.6696520
2013
Earlier work this paper cites.
Antol, S., Agrawal, A., Lu, J., Mitchell, M., Batra, D., Zitnick, C.L., Parikh, D.: Vqa: Visual question answering pp. 2425–2433 (2015)
2015
Earlier work this paper cites.
Gao, L., Guo, Z., Zhang, H., Xu, X., Shen, H.T.: Video captioning with attention-based lstm and semantic consistency. IEEE Transactions on Multimedia 19
2017
Earlier work this paper cites.
ten Pas, A., Gualtieri, M., Saenko, K., Platt, R.: Grasp pose detection in point clouds. The International Journal of Robotics Research 36
2017
Earlier work this paper cites.
Anderson, P., Wu, Q., Teney, D., Bruce, J., Johnson, M., Sünderhauf, N., Reid, I., Gould, S., Van Den Hengel, A.: Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments pp. 3674–3683 (2018)
2018
Earlier work this paper cites.
Das, A., Datta, S., Gkioxari, G., Lee, S., Parikh, D., Batra, D.: Embodied question answering pp. 1–10 (2018)
2018
Earlier work this paper cites.
Wang, B., Ma, L., Zhang, W., Liu, W.: Reconstruction network for video captioning pp. 7622–7631 (2018)
2018
Earlier work this paper cites.
2019
Earlier work this paper cites.
Liu, R., Liu, C., Bai, Y., Yuille, A.L.: Clevr-ref+: Diagnosing visual reasoning with referring expressions pp. 4185–4194 (2019)
2019
Earlier work this paper cites.
Lu, J., Batra, D., Parikh, D., Lee, S.: Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. Advances in neural information processing systems 32
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2020
Cited alongside, same era.
Chen, H., Ding, G., Liu, X., Lin, Z., Liu, J., Han, J.: Imram: Iterative matching with recurrent attention memory for cross-modal image-text retrieval pp. 12655–12663 (2020)
2020
Cited alongside, same era.
James, S., Ma, Z., Arrojo, D.R., Davison, A.J.: Rlbench: The robot learning benchmark & learning environment. IEEE Robotics and Automation Letters 5
2020
Cited alongside, same era.
2020
Cited alongside, same era.
2021
Later among the works it cites.
2021
Later among the works it cites.
Pashevich, A., Schmid, C., Sun, C.: Episodic transformer for vision-and-language navigation pp. 15942–15952 (2021)
2021
Later among the works it cites.
Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al.: Learning transferable visual models from natural language supervision pp. 8748–8763 (2021)
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2020
Cited alongside, same era.
Stepputtis, S., Campbell, J., Phielipp, M., Lee, S., Baral, C., Amor, H.B.: Language-conditioned imitation learning for robot manipulation tasks (2020)
2020
Cited alongside, same era.
Yu, T., Quillen, D., He, Z., Julian, R., Hausman, K., Finn, C., Levine, S.: Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning pp. 1094–1100 (2020)
2020
Cited alongside, same era.
Zeng, A., Florence, P., Tompson, J., Welker, S., Chien, J., Attarian, M., Armstrong, T., Krasin, I., Duong, D., Sindhwani, V., Lee, J.: Transporter networks: Rearranging the visual world for robotic manipulation. Conference on Robot Learning (CoRL) (2020)
2020
Cited alongside, same era.
Zhang, Q., Lei, Z., Zhang, Z., Li, S.Z.: Context-aware attention network for image-text retrieval pp. 3536–3545 (2020)
2020
Cited alongside, same era.
2020
Cited alongside, same era.
Ehsani, K., Han, W., Herrasti, A., VanderBilt, E., Weihs, L., Kolve, E., Kembhavi, A., Mottaghi, R.: Manipulathor: A framework for visual object manipulation pp. 4497–4506 (2021)
2021
Cited alongside, same era.
2021
Cited alongside, same era.
Shao, L., Migimatsu, T., Zhang, Q., Yang, K., Bohg, J.: Concept2robot: Learning manipulation concepts from instructions and human demonstrations. The International Journal of Robotics Research 40
2021
Later among the works it cites.
Szot, A., Clegg, A., Undersander, E., Wijmans, E., Zhao, Y., Turner, J., Maestre, N., Mukadam, M., Chaplot, D.S., Maksymets, O., et al.: Habitat 2.0: Training home assistants to rearrange their habitat. Advances in Neural Information Processing Systems 34
2021
Later among the works it cites.
2021
Later among the works it cites.
2022
Closest in time.
Mees, O., Hermann, L., Rosete-Beas, E., Burgard, W.: Calvin: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks. IEEE Robotics and Automation Letters (RA-L) 7
2022
Closest in time.
Shridhar, M., Manuelli, L., Fox, D.: Cliport: What and where pathways for robotic manipulation pp. 894–906 (2022)
2022
Closest in time.
Srivastava, S., Li, C., Lingelbach, M., Martín-Martín, R., Xia, F., Vainio, K.E., Lian, Z., Gokmen, C., Buch, S., Liu, K., et al.: Behavior: Benchmark for everyday household activities in virtual, interactive, and ecological environments pp. 477–490 (2022)
2022
Closest in time.
Yuan, W., Paxton, C., Desingh, K., Fox, D.: Sornet: Spatial object-centric representations for sequential manipulation pp. 148–157 (2022)
2022
Closest in time.