Fetching the paper…
Reading the bibliography…
One of the most challenging topics in Natural Language Processing (NLP) is visually-grounded language understanding and reasoning.
Wanrong Zhu, Zhiting Hu, and Eric P. Xing. 2019 · 1901
Earlier work this paper cites.
Visualbert: A simple and performant baseline for vision and language
Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang. 2019 · 1908
Earlier work this paper cites.
A learning algorithm for continually running fully recurrent neural networks
Ronald J. Williams and David Zipser. 1989 · 1989
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Harsh Mehta, Yoav Artzi, Jason Baldridge, Eugene Ie, and Piotr Mirowski. 2020 · 2001
Earlier work this paper cites.
Univilm: A unified video and language pre-training model for multimodal understanding and generation
Huaishao Luo, Lei Ji, Botian Shi, Haoyang Huang, Nan Duan, Tianrui Li, Xilin Chen, and Ming Zhou. 2020 · 2002
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Pixel-bert: Aligning image pixels with text by deep multi-modal transformers
Zhicheng Huang, Zhaoyang Zeng, Bei Liu, Dongmei Fu, and Jianlong Fu. 2020b · 2004
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Natural Language Processing with Python
Steven Bird, Ewan Klein, and Edward Loper. 2009 · 2009
Earlier work this paper cites.
Hogwild: A lock-free approach to parallelizing stochastic gradient descent
Benjamin Recht, Christopher Ré, Stephen J. Wright, and Feng Niu. 2011 · 2011
Earlier work this paper cites.
Image description using visual dependency representations
Desmond Elliott and Frank Keller. 2013 · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba. 2015 · 2015
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael S. Bernstein, Alexander C. Berg, and Fei-Fei Li. 2015 · 2015
Earlier work this paper cites.
Cider: Consensus-based image description evaluation
Ramakrishna Vedantam, C. Lawrence Zitnick, and Devi Parikh. 2015 · 2015
Earlier work this paper cites.
SPICE: semantic propositional image caption evaluation
Peter Anderson, Basura Fernando, Mark Johnson, and Stephen Gould. 2016 · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
Extracting optimal performance from dynamic time warping
Abdullah Mueen and Eamonn J. Keogh. 2016 · 2016
Earlier work this paper cites.
Toward controlled generation of text
Zhiting Hu, Zichao Yang, Xiaodan Liang, Ruslan Salakhutdinov, and Eric P. Xing. 2017 · 2017
Earlier work this paper cites.
Style transfer from non-parallel text by cross-alignment
Tianxiao Shen, Tao Lei, Regina Barzilay, and Tommi S. Jaakkola. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments
Peter Anderson, Qi Wu, Damien Teney, Jake Bruce, Mark Johnson, Niko Sünderhauf, Ian D. Reid, Stephen Gould, and Anton van den Hengel. 2018 · 2018
Cited alongside, same era.
Unsupervised neural machine translation
Mikel Artetxe, Gorka Labaka, Eneko Agirre, and Kyunghyun Cho. 2018 · 2018
Cited alongside, same era.
Gated-attention architectures for task-oriented language grounding
Devendra Singh Chaplot, Kanthashree Mysore Sathyendra, Rama Kumar Pasumarthi, Dheeraj Rajagopal, and Ruslan Salakhutdinov. 2018 · 2018
Cited alongside, same era.
The Book of Why: The New Science of Cause and Effect, Judea Pearl, Dana Mackenzie. Basic Books (2018) , volume 284
Norman E. Fenton, Martin Neil, and Anthony C. Constantinou. 2020 · 2018
Cited alongside, same era.
TIGS: an inference algorithm for text infilling with gradient search
Dayiheng Liu, Jie Fu, Pengfei Liu, and Jiancheng Lv. 2019 · 2019
Later among the works it cites.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. 2019 · 2019
Later among the works it cites.
Self-monitoring navigation agent via auxiliary progress estimation
Chih-Yao Ma, Jiasen Lu, Zuxuan Wu, Ghassan AlRegib, Zsolt Kira, Richard Socher, and Caiming Xiong. 2019a · 2019
Later among the works it cites.
The regretful agent: Heuristic-aided navigation through progress estimation
Chih-Yao Ma, Zuxuan Wu, Ghassan AlRegib, Caiming Xiong, and Zsolt Kira. 2019b · 2019
Later among the works it cites.
Videobert: A joint model for video and language representation learning
Chen Sun, Austin Myers, Carl Vondrick, Kevin Murphy, and Cordelia Schmid. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Speaker-follower models for vision-and-language navigation
Daniel Fried, Ronghang Hu, Volkan Cirik, Anna Rohrbach, Jacob Andreas, Louis-Philippe Morency, Taylor Berg-Kirkpatrick, Kate Saenko, Dan Klein, and Trevor Darrell. 2018 · 2018
Cited alongside, same era.
Style transfer in text: Exploration and evaluation
Zhenxin Fu, Xiaoye Tan, Nanyun Peng, Dongyan Zhao, and Rui Yan. 2018 · 2018
Cited alongside, same era.
Unsupervised machine translation using monolingual corpora only
Guillaume Lample, Alexis Conneau, Ludovic Denoyer, and Marc’Aurelio Ranzato. 2018a · 2018
Cited alongside, same era.
Phrase-based & neural unsupervised machine translation
Guillaume Lample, Myle Ott, Alexis Conneau, Ludovic Denoyer, and Marc’Aurelio Ranzato. 2018b · 2018
Cited alongside, same era.
Learning to navigate in cities without a map
Piotr Mirowski, Matthew Koichi Grimes, Mateusz Malinowski, Karl Moritz Hermann, Keith Anderson, Denis Teplyashin, Karen Simonyan, Koray Kavukcuoglu, Andrew Zisserman, and Raia Hadsell. 2018 · 2018
Cited alongside, same era.
Look before you leap: Bridging model-free and model-based reinforcement learning for planned-ahead vision-and-language navigation
Xin Wang, Wenhan Xiong, Hongmin Wang, and William Yang Wang. 2018 · 2018
Cited alongside, same era.
Unsupervised text style transfer using language models as discriminators
Zichao Yang, Zhiting Hu, Chris Dyer, Eric P. Xing, and Taylor Berg-Kirkpatrick. 2018 · 2018
Cited alongside, same era.
LXMERT: learning cross-modality encoder representations from transformers
Hao Tan and Mohit Bansal. 2019 · 2019
Later among the works it cites.
Learning to navigate unseen environments: Back translation with environmental dropout
Hao Tan, Licheng Yu, and Mohit Bansal. 2019 · 2019
Later among the works it cites.
Multimodal transformer for unaligned multimodal language sequences
Yao-Hung Hubert Tsai, Shaojie Bai, Paul Pu Liang, J. Zico Kolter, Louis-Philippe Morency, and Ruslan Salakhutdinov. 2019 · 2019
Later among the works it cites.
Reinforced cross-modal matching and self-supervised imitation learning for vision-language navigation
Xin Wang, Qiuyuan Huang, Asli Çelikyilmaz, Jianfeng Gao, Dinghan Shen, Yuan-Fang Wang, William Yang Wang, and Lei Zhang. 2019 · 2019
Later among the works it cites.
UNITER: universal image-text representation learning
Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu. 2020 · 2020
Closest in time.
Enabling language models to fill in the blanks
Chris Donahue, Mina Lee, and Percy Liang. 2020 · 2020
Closest in time.
Towards learning a generic agent for vision-and-language navigation via pre-training
Weituo Hao, Chunyuan Li, Xiujun Li, Lawrence Carin, and Jianfeng Gao. 2020 · 2020
Closest in time.
INSET: sentence infilling with inter-sentential transformer
Yichen Huang, Yizhe Zhang, Oussama Elachqar, and Yu Cheng. 2020a · 2020
Closest in time.
Unicoder-vl: A universal encoder for vision and language by cross-modal pre-training
Gen Li, Nan Duan, Yuejian Fang, Ming Gong, and Daxin Jiang. 2020 · 2020
Closest in time.
Improving vision-and-language navigation with image-text pairs from the web
Arjun Majumdar, Ayush Shrivastava, Stefan Lee, Peter Anderson, Devi Parikh, and Dhruv Batra. 2020 · 2020
Closest in time.
Multi-modality cross attention network for image and sentence matching
Xi Wei, Tianzhu Zhang, Yan Li, Yongdong Zhang, and Feng Wu. 2020 · 2020
Closest in time.
Learning to stop: A simple yet effective approach to urban vision-language navigation
Jiannan Xiang, Xin Wang, and William Yang Wang. 2020 · 2020
Closest in time.
Diagnosing the environment bias in vision-and-language navigation
Yubo Zhang, Hao Tan, and Mohit Bansal. 2020 · 2020
Closest in time.
Cross-modality relevance for reasoning on language and vision
Chen Zheng, Quan Guo, and Parisa Kordjamshidi. 2020 · 2020
Closest in time.