Fetching the paper…
Reading the bibliography…
People always desire an embodied agent that can perform a task by understanding language instruction.
Lewis, M., Liu, Y., Goyal, N., Ghazvininejad, M., Mohamed, A., Levy, O., Stoyanov, V., and Zettlemoyer, L. (2019) · 1910
Earlier work this paper cites.
Huggingface’s transformers: State-of-the-art natural language processing
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., et al. (2019) · 1910
Earlier work this paper cites.
A density-based algorithm for discovering clusters in large spatial databases with noise
Ester, M., Kriegel, H.-P., Sander, J., Xu, X., et al. (1996) · 1996
Earlier work this paper cites.
A fast marching level set method for monotonically advancing fronts
Sethian, J. A. (1996) · 1996
Earlier work this paper cites.
Scikit-learn: Machine learning in python
Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., et al. (2011) · 2011
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Sutskever, I., Vinyals, O., and Le, Q. V. (2014) · 2014
Earlier work this paper cites.
Virtual embodiment: A scalable long-term strategy for artificial intelligence research
Kiela, D., Bulat, L., Vero, A. L., and Clark, S. (2016) · 2016
Earlier work this paper cites.
Ai2-thor: An interactive 3d environment for visual ai
Kolve, E., Mottaghi, R., Han, W., VanderBilt, E., Weihs, L., Herrasti, A., Gordon, D., Zhu, Y., Gupta, A., and Farhadi, A. (2017) · 2017
Earlier work this paper cites.
Embodied Question Answering
Das, A., Datta, S., Gkioxari, G., Lee, S., Parikh, D., and Batra, D. (2018) · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. (2018) · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G. (2018) · 2018
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., et al. (2019) · 2019
Cited alongside, same era.
Why build an assistant in minecraft?
Szlam, A., Gray, J., Srinet, K., Jernite, Y., Joulin, A., Synnaeve, G., Kiela, D., Yu, H., Chen, Z., Goyal, S., Guo, D., Rothermel, D., Zitnick, C. L., and Weston, J. (2019) · 2019
Cited alongside, same era.
Alfred: A benchmark for interpreting grounded instructions for everyday tasks
Shridhar, M., Thomason, J., Gordon, D., Bisk, Y., Han, W., Mottaghi, R., Zettlemoyer, L., and Fox, D. (2020) · 2020
Cited alongside, same era.
Film: Following instructions in language with modular methods
Min, S. Y., Chaplot, D. S., Ravikumar, P., Bisk, Y., and Salakhutdinov, R. (2021) · 2021
Later among the works it cites.
Look wide and interpret twice: Improving performance on interactive instruction-following tasks
Nguyen, V.-Q., Suganuma, M., and Okatani, T. (2021) · 2021
Later among the works it cites.
Modular framework for visuomotor language grounding
Nottingham, K., Liang, L., Shin, D., Fowlkes, C. C., Fox, R., and Singh, S. (2021) · 2021
Later among the works it cites.
Episodic transformer for vision-and-language navigation
Pashevich, A., Schmid, C., and Sun, C. (2021) · 2021
Later among the works it cites.
igibson 1.0: a simulation environment for interactive tasks in large realistic scenes
Shen, B., Xia, F., Li, C., Martín-Martín, R., Fan, L., Wang, G., Pérez-D’Arpino, C., Buch, S., Srivastava, S., Tchapmi, L. P., Tchapmi, M. E., Vainio, K., Wong, J., Fei-Fei, L., and Savarese, S. (2021) · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Moca: A modular object-centric approach for interactive instruction following
Singh, K. P., Bhambri, S., Kim, B., Mottaghi, R., and Choi, J. (2020) · 2020
Cited alongside, same era.
Interactive gibson benchmark: A benchmark for interactive navigation in cluttered environments
Xia, F., Shen, W. B., Li, C., Kasimbeg, P., Tchapmi, M. E., Toshev, A., Martín-Martín, R., and Savarese, S. (2020) · 2020
Cited alongside, same era.
A persistent spatial semantic representation for high-level natural language instruction execution
Blukis, V., Paxton, C., Fox, D., Garg, A., and Artzi, Y. (2021) · 2021
Cited alongside, same era.
Agent with the big picture: Perceiving surroundings for interactive instruction following
Kim, B., Bhambri, S., Singh, K. P., Mottaghi, R., and Choi, J. (2021) · 2021
Cited alongside, same era.
Later among the works it cites.
Embodied : A transformer model for embodied, language-guided visual task completion
Suglia, A., Gao, Q., Thomason, J., Thattai, G., and Sukhatme, G. (2021) · 2021
Later among the works it cites.
Visual room rearrangement
Weihs, L., Deitke, M., Kembhavi, A., and Mottaghi, R. (2021) · 2021
Later among the works it cites.
Hierarchical task learning from language instructions with unified transformers and self-monitoring
Zhang, Y. and Chai, J. (2021) · 2021
Later among the works it cites.