Fetching the paper…
Reading the bibliography…
Instruction-following agents must ground language into their observation and action spaces.
Word & Object
Quine, W. V. O · 1960
Earlier work this paper cites.
Understanding natural language
Winograd, T · 1972
Earlier work this paper cites.
Ad hoc categories
Barsalou, L. W · 1983
Earlier work this paper cites.
The symbol grounding problem
Harnad, S · 1990
Earlier work this paper cites.
Learning to connect language and perception
Mooney, R. J · 2008
Earlier work this paper cites.
Toward understanding natural language directions
Kollar, T., Tellex, S., Roy, D., and Roy, N · 2010
Earlier work this paper cites.
Learning to interpret natural language navigation instructions from observations
Chen, D. and Mooney, R · 2011
Earlier work this paper cites.
Understanding natural language commands for robotic navigation and mobile manipulation
Tellex, S., Kollar, T., Dickerson, S., Walter, M., Banerjee, A., Teller, S., and Roy, N · 2011
Earlier work this paper cites.
Learning to win by reading manuals in a Monte-Carlo framework
Branavan, S., Silver, D., and Barzilay, R · 2012
Earlier work this paper cites.
Weakly supervised learning of semantic parsers for mapping instructions to actions
Artzi, Y. and Zettlemoyer, L · 2013
Earlier work this paper cites.
Active learning for teaching a robot grounded relational symbols
Kulick, J., Toussaint, M., Lang, T., and Lopes, M · 2013
Earlier work this paper cites.
Linguistic regularities in continuous space word representations
Mikolov, T., Yih, W.-t., and Zweig, G · 2013
Earlier work this paper cites.
Learning goal-oriented hierarchical tasks from situated interactive instruction
Mohan, S. and Laird, J · 2014
Earlier work this paper cites.
Teaching robots new actions through natural language instructions
She, L., Cheng, Y., Chai, J. Y., Jia, Y., Yang, S., and Xi, N · 2014
Earlier work this paper cites.
Listen, attend, and walk: Neural mapping of navigational instructions to action sequences
Mei, H., Bansal, M., and Walter, M. R · 2016
Earlier work this paper cites.
Supersizing self-supervision: Learning to grasp from 50k tries and 700 robot hours
Pinto, L. and Gupta, A · 2016
Earlier work this paper cites.
Hindsight experience replay
Andrychowicz, M., Wolski, F., Ray, A., Schneider, J., Fong, R., Welinder, P., McGrew, B., Tobin, J., Pieter Abbeel, O., and Zaremba, W · 2017
Earlier work this paper cites.
Mapping instructions and visual observations to actions with reinforcement learning
Misra, D., Langford, J., and Artzi, Y · 2017
Earlier work this paper cites.
Opportunistic active learning for grounding natural language descriptions
Thomason, J., Padmakumar, A., Sinapov, J., Hart, J., Stone, P., and Mooney, R · 2017
Earlier work this paper cites.
Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments
Anderson, P., Wu, Q., Teney, D., Bruce, J., Johnson, M., Sünderhauf, N., Reid, I., Gould, S., and van den Hengel, A · 2018
Earlier work this paper cites.
Learning to understand goal specifications by modelling reward
Bahdanau, D., Hill, F., Leike, J., Hughes, E., Hosseini, A., Kohli, P., and Grefenstette, E · 2018
Earlier work this paper cites.
Language to action: Towards interactive task learning with physical agents
Chai, J. Y., Gao, Q., She, L., Yang, S., Saba-Sadiya, S., and Xu, G · 2018
Earlier work this paper cites.
Gated-attention architectures for task-oriented language grounding
Chaplot, D. S., Sathyendra, K. M., Pasumarthi, R. K., Rajagopal, D., and Salakhutdinov, R · 2018
Earlier work this paper cites.
Guiding policies with language via meta-learning
Co-Reyes, J. D., Gupta, A., Sanjeev, S., Altieri, N., Andreas, J., DeNero, J., Abbeel, P., and Levine, S · 2018
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. N · 2018
Earlier work this paper cites.
Speaker-follower models for vision-and-language navigation
Fried, D., Hu, R., Cirik, V., Rohrbach, A., Andreas, J., Morency, L.-P., Berg-Kirkpatrick, T., Saenko, K., Klein, D., and Darrell, T · 2018
Earlier work this paper cites.
Grounding language for transfer in deep reinforcement learning
Narasimhan, K., Barzilay, R., and Jaakkola, T · 2018
Cited alongside, same era.
Interactive grounded language acquisition and generalization in a 2D world
Yu, H., Zhang, H., and Xu, W · 2018
Cited alongside, same era.
Learning to map natural language instructions to physical quadcopter control using simulated flight
Blukis, V., Terme, Y., Niklasson, E., Knepper, R. A., and Artzi, Y · 2019
Cited alongside, same era.
Actrce: Augmenting experience via teacher’s advice for multi-goal reinforcement learning
Chan, H., Wu, Y., Kiros, J., Fidler, S., and Ba, J · 2019
Cited alongside, same era.
Robust navigation with language pretraining and stochastic sampling
Li, X., Li, C., Xia, Q., Bisk, Y., Celikyilmaz, A., Gao, J., Smith, N. A., and Choi, Y · 2019
Cited alongside, same era.
SILG: The multi-domain symbolic interactive language grounding benchmark
Zhong, V., Hanjie, A. W., Wang, S., Narasimhan, K., and Zettlemoyer, L · 2021
Later among the works it cites.
Do as I can, not as I say: Grounding language in robotic affordances
Ahn, M., Brohan, A., Brown, N., Chebotar, Y., Cortes, O., David, B., Finn, C., Gopalakrishnan, K., Hausman, K., Herzog, A., et al · 2022
Later among the works it cites.
Flamingo: a visual language model for few-shot learning
Alayrac, J.-B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., Ring, R., Rutherford, E., Cabi, S., Han, T., Gong, Z., Samangooei, S., Monteiro, M., Menick, J., Borgeaud, S., Brock, A., Nematzadeh, A., Sharifzadeh, S., Binkowski, M., Barreira, R., Vinyals, O., Zisserman, A., and Simonyan, K · 2022
Later among the works it cites.
LaTTe: Language trajectory transformEr
Bucker, A. F. C., Figueredo, L. F. C., Haddadin, S., Kapoor, A., Ma, S., Vemprala, S., and Bonatti, R · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Habitat: A platform for embodied ai research
Savva, M., Kadian, A., Maksymets, O., Zhao, Y., Wijmans, E., Jain, B., Straub, J., Liu, J., Koltun, V., Malik, J., et al · 2019
Cited alongside, same era.
Imitating interactive intelligence
Abramson, J., Ahuja, A., Barr, I., Brussee, A., Carnevale, F., Cassin, M., Chhaparia, R., Clark, S., Damoc, B., Dudzik, A., et al · 2020
Cited alongside, same era.
Few-shot object grounding and mapping for natural language robot instruction following
Blukis, V., Knepper, R. A., and Artzi, Y · 2020
Cited alongside, same era.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T. J., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Cited alongside, same era.
Ask your humans: Using human instructions to improve generalization in reinforcement learning
Chen, V., Gupta, A. K., and Marino, K · 2020
Cited alongside, same era.
Higher: Improving instruction following with hindsight generation for experience replay
Cideron, G., Seurin, M., Strub, F., and Pietquin, O · 2020
Cited alongside, same era.
Towards learning a generic agent for vision-and-language navigation via pre-training
Hao, W., Li, C., Li, X., Carin, L., and Gao, J · 2020
Cited alongside, same era.
Chen, S., Guhur, P.-L., Tapaswi, M., Schmid, C., and Laptev, I · 2022
Later among the works it cites.
Collaborating with language models for embodied reasoning
Dasgupta, I., Kaeser-Chen, C., Marino, K., Ahuja, A., Babayan, S., Hill, F., and Fergus, R · 2022
Later among the works it cites.
MineDojo: Building open-ended embodied agents with internet-scale knowledge
Fan, L. J., Wang, G., Jiang, Y., Mandlekar, A., Yang, Y., Zhu, H., Tang, A., Huang, D.-A., Zhu, Y., and Anandkumar, A · 2022
Later among the works it cites.
What makes certain pre-trained visual representations better for robotic learning?
Hsu, K., Lum, T. G. W., Gao, R., Gu, S. S., Wu, J., and Finn, C · 2022
Later among the works it cites.
F-vlm: Open-vocabulary object detection upon frozen vision and language models
Kuo, W., Cui, Y., Gu, X., Piergiovanni, A., and Angelova, A · 2022
Later among the works it cites.
Inferring rewards from language in context
Lin, J., Fried, D., Klein, D., and Dragan, A · 2022
Later among the works it cites.
ZSON: Zero-shot object-goal navigation using multimodal goal embeddings
Majumdar, A., Aggarwal, G., Devnani, B. S., Hoffman, J., and Batra, D · 2022
Later among the works it cites.
Simple open-vocabulary object detection
Minderer, M., Gritsenko, A., Stone, A., Neumann, M., Weissenborn, D., Dosovitskiy, A., Mahendran, A., Arnab, A., Dehghani, M., Shen, Z., Wang, X., Zhai, X., Kipf, T., and Houlsby, N · 2022
Later among the works it cites.
R3M: A universal visual representation for robot manipulation
Nair, S., Rajeswaran, A., Kumar, V., Finn, C., and Gupta, A · 2022
Later among the works it cites.
Hierarchical text-conditional image generation with clip latents
Ramesh, A., Dhariwal, P., Nichol, A., Chu, C., and Chen, M · 2022
Later among the works it cites.
LM-Nav: Robotic navigation with large pre-trained models of language, vision, and action
Shah, D., Osiński, B., Levine, S., et al · 2022
Later among the works it cites.
Skill induction and planning with latent language
Sharma, P., Torralba, A., and Andreas, J · 2022
Later among the works it cites.
Progprompt: Generating situated robot task plans using large language models
Singh, I., Blukis, V., Mousavian, A., Goyal, A., Xu, D., Tremblay, J., Fox, D., Thomason, J., and Garg, A · 2022
Later among the works it cites.
Semantic exploration from language abstractions and pretrained representations
Tam, A., Rabinowitz, N. C., Lampinen, A. K., Roy, N. A., Chan, S. C., Strouse, D., Wang, J. X., Banino, A., and Hill, F · 2022
Later among the works it cites.
Robotic skill acquisition via instruction augmentation with vision-language models
Xiao, T., Chan, H., Sermanet, P., Wahid, A., Brohan, A., Hausman, K., Levine, S., and Tompson, J · 2022
Later among the works it cites.
Intra-agent speech permits zero-shot task acquisition
Yan, C., Carnevale, F., Georgiev, P., Santoro, A., Guy, A., Muldal, A., Hung, C.-C., Abramson, J. S., Lillicrap, T. P., and Wayne, G · 2022
Later among the works it cites.
Socratic models: Composing zero-shot multimodal reasoning with language
Zeng, A., Wong, A., Welker, S., Choromanski, K., Tombari, F., Purohit, A., Ryoo, M., Sindhwani, V., Lee, J., Vanhoucke, V., et al · 2022
Later among the works it cites.
Improving policy learning via language dynamics distillation
Zhong, V., Mu, J., Zettlemoyer, L., Grefenstette, E., and Rocktäschel, T · 2022
Later among the works it cites.
Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L., Xiao, T., Whitehead, S., Berg, A. C., Lo, W.-Y., et al · 2023
Closest in time.
Prismer: A vision-language model with an ensemble of experts
Liu, S., Fan, L., Johns, E., Yu, Z., Xiao, C., and Anandkumar, A · 2023
Closest in time.
Gpt-4 technical report, 2023
OpenAI · 2023
Closest in time.