Fetching the paper…
Reading the bibliography…
Probabilistic embeddings have proven useful for capturing polysemous word meanings, as well as ambiguity in image matching.
Distinctive image features from scale-invariant keypoints
David G. Lowe · 2004
Earlier work this paper cites.
Object retrieval with large vocabularies and fast spatial matching
James Philbin, Ondrej Chum, Michael Isard, Josef Sivic, and Andrew Zisserman · 2007
Earlier work this paper cites.
Large scale online learning of image similarity through ranking
G. Chechik, V. Sharma, U. Shalit, and S. Bengio · 2010
Earlier work this paper cites.
Aggregating local descriptors into a compact image representation
Hervé Jégou, Matthijs Douze, Cordelia Schmid, and Patrick Pérez · 2010
Earlier work this paper cites.
Aggregating local image descriptors into compact codes
Hervé Jégou, Florent Perronnin, Matthijs Douze, Jorge Sánchez, Patrick Pérez, and Cordelia Schmid · 2012
Earlier work this paper cites.
Devise: A deep visual-semantic embedding model
Andrea Frome, Gregory S. Corrado, Jonathon Shlens, Samy Bengio, Jeffrey Dean, Marc’Aurelio Ranzato, and Tomás Mikolov · 2013
Earlier work this paper cites.
Learning phrase representations using RNN encoder-decoder for statistical machine translation
Kyunghyun Cho, Bart van Merrienboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio · 2014
Earlier work this paper cites.
Microsoft COCO: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick · 2014
Earlier work this paper cites.
Word representations via Gaussian embedding
Luke Vilnis and Andrew McCallum · 2014
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Peter Young, Alice Lai, Micah Hodosh, and Julia Hockenmaier · 2014
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions
A. Karpathy and L. Fei-Fei · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
FaceNet: A unified embedding for face recognition and clustering
Florian Schroff, Dmitry Kalenichenko, and James Philbin · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Focused evaluation for image description with binary forced-choice tasks
Micah Hodosh and Julia Hockenmaier · 2016
Earlier work this paper cites.
Visual Genome: Connecting language and vision using crowdsourced dense image annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, Michael Bernstein, and Li Fei-Fei · 2016
Earlier work this paper cites.
Gaussian visual-linguistic embedding for zero-shot recognition
Tanmoy Mukherjee and Timothy Hospedales · 2016
Cited alongside, same era.
Order-embeddings of images and language
Ivan Vendrov, Ryan Kiros, Sanja Fidler, and Raquel Urtasun · 2016
Cited alongside, same era.
Multimodal word distributions
Ben Athiwaratkun and Andrew Gordon Wilson · 2017
Cited alongside, same era.
Instance-aware image and sentence matching with selective multimodal LSTM
Yan Huang, Wei Wang, and Liang Wang · 2017
Cited alongside, same era.
What uncertainties do we need in Bayesian deep learning for computer vision?
Alex Kendall and Yarin Gal · 2017
Cited alongside, same era.
Learning from uncertain curves: The 2-Wasserstein metric for Gaussian processes
Anton Mallasto and Aasa Feragen · 2017
Cited alongside, same era.
Visual semantic reasoning for image-text matching
Kunpeng Li, Yulun Zhang, Kai Li, Yuanyuan Li, and Yun Fu · 2019
Later among the works it cites.
ViLBERT: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee · 2019
Later among the works it cites.
Modeling uncertainty with hedged instance embedding
Seong Joon Oh, Kevin Murphy, Jiyan Pan, Joseph Roth, Florian Schroff, and Andrew Gallagher · 2019
Later among the works it cites.
PyTorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Köpf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala · 2019
Later among the works it cites.
Polysemous visual-semantic embedding for cross-modal retrieval
Yale Song and Mohammad Soleymani · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dual attention networks for multimodal reasoning and matching
Hyeonseob Nam, Jung-Woo Ha, and Jeonghee Kim · 2017
Cited alongside, same era.
Foil it! Find one mismatch between image and language caption
Ravi Shekhar, Sandro Pezzelle, Yauhen Klimovich, Aurelie Herbelot, Moin Nabi, Enver Sangineto, and Raffaella Bernardi · 2017
Cited alongside, same era.
Bottom-up and top-down attention for image captioning and visual question answering
Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang · 2018
Cited alongside, same era.
NetVLAD: CNN architecture for weakly supervised place recognition
Relja Arandjelovic, Petr Gronát, Akihiko Torii, Tomás Pajdla, and Josef Sivic · 2018
Cited alongside, same era.
Hierarchical density order embeddings
Ben Athiwaratkun and Andrew Gordon Wilson · 2018
Cited alongside, same era.
Probabilistic FastText for multi-sense word embeddings
Ben Athiwaratkun, Andrew Gordon Wilson, and Anima Anandkumar · 2018
Cited alongside, same era.
CAMP: Cross-modal adaptive message passing for text-image retrieval
Zihao Wang, Xihui Liu, Hongsheng Li, Lu Sheng, Junjie Yan, Xiaogang Wang, and Jing Shao · 2019
Later among the works it cites.
UniVSE: Robust visual semantic embeddings via structured semantic representations
Hao Wu, Jiayuan Mao, Yufeng Zhang, Yuning Jiang, Lei Li, Weiwei Sun, and Wei-Ying Ma · 2019
Later among the works it cites.
12-in-1: Multi-task vision and language representation learning
Jiasen Lu, Vedanuj Goswami, Marcus Rohrbach, Devi Parikh, and Stefan Lee · 2020
Later among the works it cites.
A metric learning reality check
Kevin Musgrave, Serge Belongie, and Ser-Nam Lim · 2020
Later among the works it cites.
Multi-modality cross attention network for image and sentence matching
Xi Wei, Tianzhu Zhang, Yan Li, Yongdong Zhang, and Feng Wu · 2020
Later among the works it cites.
Learning to represent image and text with denotation graph
Bowen Zhang, Hexiang Hu, Vihan Jain, Eugene Ie, and Fei Sha · 2020
Later among the works it cites.
Context-aware attention network for image-text retrieval
Qi Zhang, Zhen Lei, Zhaoxiang Zhang, and Stan Z. Li · 2020
Later among the works it cites.
Dual-path convolutional image-text embeddings with instance loss
Zhedong Zheng, Liang Zheng, Michael Garrett, Yi Yang, Mingliang Xu, and Yi-Dong Shen · 2020
Later among the works it cites.
Integrating information theory and adversarial learning for cross-modal retrieval
Wei Chen, Yu Liu, Erwin M. Bakker, and Michael S. Lew · 2021
Later among the works it cites.
Probabilistic embeddings for cross-modal retrieval
Sanghyuk Chun, Seong Joon Oh, Rafael Sampaio de Rezende, Yannis Kalantidis, and Diane Larlus · 2021
Later among the works it cites.
Crisscrossed captions: Extended intramodal and intermodal semantic similarity judgments for MS-COCO
Zarana Parekh, Jason Baldridge, Daniel Cer, Austin Waters, and Yinfei Yang · 2021
Later among the works it cites.