Fetching the paper…
Reading the bibliography…
To interact with humans and act in the world, agents need to understand the range of language that people use and relate it to the visual world.
Sentence-bert: Sentence embeddings using siamese bert-networks
Reimers, N. and Gurevych, I · 1908
Earlier work this paper cites.
The origin of speech
Hockett, C. F. and Hockett, C. D · 1960
Earlier work this paper cites.
Understanding natural language
Winograd, T · 1972
Earlier work this paper cites.
Dyna, an integrated architecture for learning, planning, and reacting
Sutton, R. S · 1991
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J · 1992
Earlier work this paper cites.
Language conditioned imitation learning over unstructured data
Lynch, C. and Sermanet, P · 2005
Earlier work this paper cites.
Reading to learn: Constructing features from semantic abstracts
Eisenstein, J., Clarke, J., Goldwasser, D., and Roth, D · 2009
Earlier work this paper cites.
Reading between the lines: Learning to map high-level instructions to commands
Branavan, S., Zettlemoyer, L., and Barzilay, R · 2010
Earlier work this paper cites.
Alfworld: Aligning text and embodied environments for interactive learning
Shridhar, M., Yuan, X., Côté, M.-A., Bisk, Y., Trischler, A., and Hausknecht, M · 2010
Earlier work this paper cites.
Estimating or propagating gradients through stochastic neurons for conditional computation
Bengio, Y., Léonard, N., and Courville, A · 2013
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. P. and Welling, M · 2013
Earlier work this paper cites.
Learning phrase representations using rnn encoder-decoder for statistical machine translation
Cho, K., Van Merriënboer, B., Gulcehre, C., Bahdanau, D., Bougares, F., Schwenk, H., and Bengio, Y · 2014
Earlier work this paper cites.
Stochastic backpropagation and approximate inference in deep generative models
Rezende, D. J., Mohamed, S., and Wierstra, D · 2014
Earlier work this paper cites.
Alignment-based compositional semantics for instruction following
Andreas, J. and Klein, D · 2015
Earlier work this paper cites.
Improved variational inference with inverse autoregressive flow
Kingma, D. P., Salimans, T., Jozefowicz, R., Chen, X., Sutskever, I., and Welling, M · 2016
Earlier work this paper cites.
Hierarchical question-image co-attention for visual question answering
Lu, J., Yang, J., Batra, D., and Parikh, D · 2016
Earlier work this paper cites.
Modular multitask reinforcement learning with policy sketches
Andreas, J., Klein, D., and Levine, S · 2017
Earlier work this paper cites.
Matterport3d: Learning from rgb-d data in indoor environments
Chang, A., Dai, A., Funkhouser, T., Halber, M., Niessner, M., Savva, M., Song, S., Zeng, A., and Zhang, Y · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
Playing text-adventure games with graph-based deep reinforcement learning
Ammanabrolu, P. and Riedl, M. O · 2018
Earlier work this paper cites.
Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments
Anderson, P., Wu, Q., Teney, D., Bruce, J., Johnson, M., Sünderhauf, N., Reid, I. D., Gould, S., and van den Hengel, A · 2018
Earlier work this paper cites.
Embodied Question Answering
Das, A., Datta, S., Gkioxari, G., Lee, S., Parikh, D., and Batra, D · 2018
Earlier work this paper cites.
Impala: Scalable distributed deep-rl with importance weighted actor-learner architectures
Espeholt, L., Soyer, H., Munos, R., Simonyan, K., Mnih, V., Ward, T., Doron, Y., Firoiu, V., Harley, T., Dunning, I., et al · 2018
Earlier work this paper cites.
Recurrent world models facilitate policy evolution
Ha, D. and Schmidhuber, J · 2018
Earlier work this paper cites.
Learning latent dynamics for planning from pixels
Hafner, D., Lillicrap, T., Fischer, I., Villegas, R., Ha, D., Lee, H., and Davidson, J · 2018
Earlier work this paper cites.
Grounding language for transfer in deep reinforcement learning
Narasimhan, K., Barzilay, R., and Jaakkola, T · 2018
Earlier work this paper cites.
Reinforcement learning: An introduction
Sutton, R. S. and Barto, A. G · 2018
Cited alongside, same era.
Seed rl: Scalable and efficient deep-rl with accelerated central inference
Espeholt, L., Marinier, R., Stanczyk, P., Wang, K., and Michalski, M · 2019
Cited alongside, same era.
Language as an abstraction for hierarchical deep reinforcement learning
Jiang, Y., Gu, S. S., Murphy, K. P., and Finn, C · 2019
Cited alongside, same era.
Recurrent experience replay in distributed reinforcement learning
Kapturowski, S., Ostrovski, G., Quan, J., Munos, R., and Dabney, W · 2019
Cited alongside, same era.
A survey of reinforcement learning informed by natural language
Luketina, J., Nardelli, N., Farquhar, G., Foerster, J. N., Andreas, J., Grefenstette, E., Whiteson, S., and Rocktäschel, T · 2019
Cited alongside, same era.
Vision-and-dialog navigation
Visualgpt: Data-efficient adaptation of pretrained language models for image captioning
Chen, J., Guo, H., Yi, K., Li, B., and Elhoseiny, M · 2022
Later among the works it cites.
Vima: General robot manipulation with multimodal prompts
Jiang, Y., Gupta, A., Zhang, Z., Wang, G., Dou, Y., Chen, Y., Fei-Fei, L., Anandkumar, A., Zhu, Y., and Fan, L · 2022
Later among the works it cites.
Pre-trained language models for interactive decision-making
Li, S., Puig, X., Du, Y., Wang, C., Akyurek, E., Torralba, A., Andreas, J., and Mordatch, I · 2022
Later among the works it cites.
Unified-io: A unified model for vision, language, and multi-modal tasks
Lu, J., Clark, C., Zellers, R., Mottaghi, R., and Kembhavi, A · 2022
Later among the works it cites.
Improving intrinsic exploration with language abstractions
Mu, J., Zhong, V., Raileanu, R., Jiang, M., Goodman, N. D., Rocktäschel, T., and Grefenstette, E · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Thomason, J., Murray, M., Cakmak, M., and Zettlemoyer, L · 2019
Cited alongside, same era.
Fine-tuning language models from human preferences
Ziegler, D. M., Stiennon, N., Wu, J., Brown, T. B., Radford, A., Amodei, D., Christiano, P., and Irving, G · 2019
Cited alongside, same era.
Imitating interactive intelligence
Abramson, J., Ahuja, A., Barr, I., Brussee, A., Carnevale, F., Cassin, M., Chhaparia, R., Clark, S., Damoc, B., Dudzik, A., et al · 2020
Cited alongside, same era.
Experience grounds language
Bisk, Y., Holtzman, A., Thomason, J., Andreas, J., Bengio, Y., Chai, J., Lapata, M., Lazaridou, A., May, J., Nisnevich, A., Pinto, N., and Turian, J · 2020
Cited alongside, same era.
Mastering atari with discrete world models
Hafner, D., Lillicrap, T., Norouzi, M., and Ba, J · 2020
Cited alongside, same era.
Beyond the nav-graph: Vision-and-language navigation in continuous environments
Krantz, J., Wijmans, E., Majumdar, A., Batra, D., and Lee, S · 2020
Cited alongside, same era.
Room-across-room: Multilingual vision-and-language navigation with dense spatiotemporal grounding
Ku, A., Anderson, P., Patel, R., Ie, E., and Baldridge, J · 2020
Cited alongside, same era.
Later among the works it cites.
TEACh: Task-driven Embodied Agents that Chat
Padmakumar, A., Thomason, J., Shrivastava, A., Lange, P., Narayan-Chen, A., Gella, S., Piramuthu, R., and Gokhan Tur and, D. H.-T · 2022
Later among the works it cites.
Meaning without reference in large language models
Piantadosi, S. T. and Hill, F · 2022
Later among the works it cites.
Reed, S., Zolna, K., Parisotto, E., Colmenarejo, S. G., Novikov, A., Barth-Maron, G., Gimenez, M., Sulsky, Y., Kay, J., Springenberg, J. T., et al · 2022
Later among the works it cites.
Cliport: What and where pathways for robotic manipulation
Shridhar, M., Manuelli, L., and Fox, D · 2022
Later among the works it cites.
Semantic exploration from language abstractions and pretrained representations
Tam, A. C., Rabinowitz, N. C., Lampinen, A. K., Roy, N. A., Chan, S. C. Y., Strouse, D., Wang, J., Banino, A., and Hill, F · 2022
Later among the works it cites.
Improving policy learning via language dynamics distillation
Zhong, V., Mu, J., Zettlemoyer, L., Grefenstette, E., and Rocktaschel, T · 2022
Later among the works it cites.
Etpnav: Evolving topological planning for vision-language navigation in continuous environments
An, D., Wang, H., Wang, W., Wang, Z., Huang, Y., He, K., and Wang, L · 2023
Closest in time.
Rt-2: Vision-language-action models transfer web knowledge to robotic control
Brohan, A., Brown, N., Carbajal, J., Chebotar, Y., Chen, X., Choromanski, K., Ding, T., Driess, D., Dubey, A., Finn, C., et al · 2023
Closest in time.
Grounding large language models in interactive environments with online reinforcement learning
Carta, T., Romac, C., Wolf, T., Lamprier, S., Sigaud, O., and Oudeyer, P.-Y · 2023
Closest in time.
Learning the effects of physical actions in a multi-modal environment
Dagan, G., Keller, F., and Lascarides, A · 2023
Closest in time.
Collaborating with language models for embodied reasoning
Dasgupta, I., Kaeser-Chen, C., Marino, K., Ahuja, A., Babayan, S., Hill, F., and Fergus, R · 2023
Closest in time.
Palm-e: An embodied multimodal language model
Driess, D., Xia, F., Sajjadi, M. S., Lynch, C., Chowdhery, A., Ichter, B., Wahid, A., Tompson, J., Vuong, Q., Yu, T., et al · 2023
Closest in time.
Tinystories: How small can language models be and still speak coherent english?, 2023
Eldan, R. and Li, Y · 2023
Closest in time.
From images to textual prompts: Zero-shot visual question answering with frozen large language models
Guo, J., Li, J., Li, D., Tiong, A. M. H., Li, B., Tao, D., and Hoi, S · 2023
Closest in time.
Mastering diverse domains through world models
Hafner, D., Pasukonis, J., Ba, J., and Lillicrap, T · 2023
Closest in time.
Visual instruction tuning
Liu, H., Li, C., Wu, Q., and Lee, Y. J · 2023
Closest in time.
Transformer-based world models are happy with 100k interactions
Robine, J., Höftmann, M., Uelwer, T., and Harmeling, S · 2023
Closest in time.
Voyager: An open-ended embodied agent with large language models
Wang, G., Xie, Y., Jiang, Y., Mandlekar, A., Xiao, C., Zhu, Y., Fan, L., and Anandkumar, A · 2023
Closest in time.
SPRING: Studying papers and reasoning to play games
Wu, Y., Min, S. Y., Prabhumoye, S., Bisk, Y., Salakhutdinov, R., Azaria, A., Mitchell, T., and Li, Y · 2023
Closest in time.
Learning interactive real-world simulators
Yang, M., Du, Y., Ghasemipour, K., Tompson, J., Schuurmans, D., and Abbeel, P · 2023
Closest in time.