Fetching the paper…
Reading the bibliography…
Cross-domain alignment between image objects and text sequences is key to many visual-language tasks, and it poses a fundamental challenge to both computer vision and natural language processing.
Unicoder-VL: A universal encoder for vision and language by cross-modal pre-training
Gen Li, Nan Duan, Yuejian Fang, Daxin Jiang, and Ming Zhou · 1908
Earlier work this paper cites.
Uniter: Learning universal image-text representations
Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu · 1909
Earlier work this paper cites.
The hungarian method for the assignment problem
Harold W Kuhn · 1955
Earlier work this paper cites.
Combinatorial matrix theory , volume 39
Richard A Brualdi and Herbert J Ryser · 1991
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter et al · 1997
Earlier work this paper cites.
Bidirectional recurrent neural networks
Mike Schuster and Kuldip K Paliwal · 1997
Earlier work this paper cites.
Object recognition with gradient-based learning
Yann LeCun et al · 1999
Earlier work this paper cites.
The earth mover’s distance as a metric for image retrieval
Yossi Rubner, Carlo Tomasi, and Leonidas J Guibas · 2000
Earlier work this paper cites.
Efficient non-maximum suppression
Alexander Neubeck and Luc Van Gool · 2006
Earlier work this paper cites.
Optimal transport: old and new , volume 338
Cédric Villani · 2008
Earlier work this paper cites.
An optimal transport approach to robust reconstruction and simplification of 2d shapes
Fernando De Goes et al · 2011
Earlier work this paper cites.
Sinkhorn distances: Lightspeed computation of optimal transport
Marco Cuturi · 2013
Earlier work this paper cites.
Unifying visual-semantic embeddings with multimodal neural language models
Ryan Kiros, Ruslan Salakhutdinov, and Richard S Zemel · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin et al · 2014
Earlier work this paper cites.
A multi-world approach to question answering about real-world scenes based on uncertain input
Mateusz Malinowski and Mario Fritz · 2014
Earlier work this paper cites.
Vqa: Visual question answering
Stanislaw Antol et al · 2015
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2015
Earlier work this paper cites.
From captions to visual concepts and back
Hao Fang, Saurabh Gupta, et al · 2015
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions
Andrej Karpathy and Li Fei-Fei · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
From word embeddings to document distances
Matt Kusner et al · 2015
Earlier work this paper cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
Bryan A Plummer et al · 2015
Cited alongside, same era.
Faster r-cnn: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun · 2015
Cited alongside, same era.
Show and tell: A neural image caption generator
Oriol Vinyals et al · 2015
Cited alongside, same era.
Show, attend and tell: Neural image caption generation with visual attention
Kelvin Xu et al · 2015
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Densecap: Fully convolutional localization networks for dense captioning
Justin Johnson, Andrej Karpathy, and Li Fei-Fei · 2016
Bottom-up and top-down attention for image captioning and visual question answering
Peter Anderson et al · 2018
Later among the works it cites.
Finding beans in burgers: Deep semantic-visual embedding with localization
Martin Engilberge, Louis Chevallier, Patrick Pérez, and Matthieu Cord · 2018
Later among the works it cites.
Vse++: Improved visual-semantic embeddings
Fartash Faghri, David J Fleet, Jamie Ryan Kiros, and Sanja Fidler · 2018
Later among the works it cites.
Look, imagine and match: Improving textual-visual cross-modal retrieval with generative models
Jiuxiang Gu, Jianfei Cai, Shafiq R Joty, Li Niu, and Gang Wang · 2018
Later among the works it cites.
Jointly discovering visual objects and spoken words from raw sensory input
David Harwath, Adria Recasens, Dídac Surís, Galen Chuang, Antonio Torralba, and James Glass · 2018
Later among the works it cites.
Learning semantic concepts and order for image and sentence matching
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Generative adversarial text to image synthesis
Scott Reed et al · 2016
Cited alongside, same era.
Order-embeddings of images and language
Ivan Vendrov, Ryan Kiros, Sanja Fidler, and Raquel Urtasun · 2016
Cited alongside, same era.
Learning deep structure-preserving image-text embeddings
Liwei Wang et al · 2016
Cited alongside, same era.
Wasserstein generative adversarial networks
Martin Arjovsky et al · 2017
Cited alongside, same era.
Computational optimal transport
M Cuturi and G Peyré · 2017
Cited alongside, same era.
Linking image and text with 2-way nets
Aviv Eisenschtat and Lior Wolf · 2017
Cited alongside, same era.
Yan Huang, Qi Wu, Chunfeng Song, and Liang Wang · 2018
Later among the works it cites.
Stacked cross attention for image-text matching
Kuang-Huei Lee et al · 2018
Later among the works it cites.
Scene classification using hierarchical wasserstein cnn
Yishu Liu et al · 2018
Later among the works it cites.
A fast proximal point method for computing exact wasserstein distance
Yujia Xie, Xiangfeng Wang, Ruijia Wang, and Hongyuan Zha · 2018
Later among the works it cites.
Weakly supervised phrase localization with multi-scale anchored transformer network
Fang Zhao, Jianshu Li, Jian Zhao, and Jiashi Feng · 2018
Later among the works it cites.
Align2ground: Weakly supervised phrase grounding guided by image-caption alignment
Samyak Datta, Karan Sikka, Anirban Roy, Karuna Ahuja, Devi Parikh, and Ajay Divakaran · 2019
Later among the works it cites.
Focus your attention: A bidirectional focal attention network for image-text matching
Chunxiao Liu, Zhendong Mao, An-An Liu, Tianzhu Zhang, Bin Wang, and Yongdong Zhang · 2019
Later among the works it cites.
VilBERT: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee · 2019
Later among the works it cites.
Mirrorgan: Learning text-to-image generation by redescription
Tingting Qiao et al · 2019
Later among the works it cites.
VL-BERT: Pre-training of generic visual-linguistic representations
Weijie Su, Xizhou Zhu, Yue Cao, Bin Li, Lewei Lu, Furu Wei, and Jifeng Dai · 2019
Later among the works it cites.
VideoBERT: A joint model for video and language representation learning
Chen Sun, Austin Myers, Carl Vondrick, Kevin Murphy, and Cordelia Schmid · 2019
Later among the works it cites.
LXMERT: Learning cross-modality encoder representations from transformers
Hao Tan and Mohit Bansal · 2019
Later among the works it cites.
Position focused attention network for image-text matching
Yaxiong Wang, Hao Yang, Xueming Qian, Lin Ma, Jing Lu, Biao Li, and Xin Fan · 2019
Later among the works it cites.
Towards learning a generic agent for vision-and-language navigation via pre-training
Weituo Hao, Chunyuan Li, Xiujun Li, Lawrence Carin, and Jianfeng Gao · 2020
Closest in time.
Oscar: Object-semantics aligned pre-training for vision-language tasks
Xiujun Li, Xi Yin, Chunyuan Li, Xiaowei Hu, Pengchuan Zhang, Lei Zhang, Lijuan Wang, Houdong Hu, Li Dong, Furu Wei, et al · 2020
Closest in time.