Fetching the paper…
Reading the bibliography…
Despite the potential of multi-modal pre-training to learn highly discriminative feature representations from complementary data modalities, current progress is being slowed by the lack of large-scale modality-diverse datasets.
Combining evidence from residual phase and mfcc features for speaker recognition
Ksr Murty and B. Yegnanarayana · 2005
Earlier work this paper cites.
Spectral hashing
Yair Weiss, Antonio Torralba, and Robert Fergus · 2008
Earlier work this paper cites.
Document clustering using locality preserving indexing and support vector machines
Chengfu Yang and Zhang Yi · 2008
Earlier work this paper cites.
NUS-WIDE: a real-world web image database from national university of singapore
Tat-Seng Chua, Jinhui Tang, Richang Hong, Haojie Li, Zhiping Luo, and Yantao Zheng · 2009
Earlier work this paper cites.
Whose vote should count more: Optimal integration of labels from labelers of unknown expertise
Jacob Whitehill, Paul Ruvolo, Tingfan Wu, Jacob Bergsma, and Javier R. Movellan · 2009
Earlier work this paper cites.
Improving web image search results using query-relative classifiers
Josip Krapac, Moray Allan, Jakob J. Verbeek, and Frédéric Jurie · 2010
Earlier work this paper cites.
Im2text: Describing images using 1 million captioned photographs
Vicente Ordonez, Girish Kulkarni, and Tamara L. Berg · 2011
Earlier work this paper cites.
Iterative quantization: A procrustean approach to learning binary codes for large-scale image retrieval
Yunchao Gong, Svetlana Lazebnik, Albert Gordo, and Florent Perronnin · 2013
Earlier work this paper cites.
Microsoft COCO: common objects in context
Tsung-Yi Lin, Michael Maire, Serge J. Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick · 2014
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Peter Young, Alice Lai, Micah Hodosh, and Julia Hockenmaier · 2014
Earlier work this paper cites.
VQA: Visual Question Answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Movieqa: Understanding stories in movies through question-answering
Makarand Tapaswi, Yukun Zhu, Rainer Stiefelhagen, Antonio Torralba, Raquel Urtasun, and Sanja Fidler · 2016
Earlier work this paper cites.
MSR-VTT: A large video description dataset for bridging video and language
Jun Xu, Tao Mei, Ting Yao, and Yong Rui · 2016
Earlier work this paper cites.
Leveraging weakly annotated data for fashion image retrieval and label prediction
Charles Corbiere, Hedi Ben-Younes, Alexandre Ramé, and Charles Ollion · 2017
Earlier work this paper cites.
The lj speech dataset
Keith Ito and Linda Johnson · 2017
Earlier work this paper cites.
TGIF-QA: toward spatio-temporal reasoning in visual question answering
Yunseok Jang, Yale Song, Youngjae Yu, Youngjin Kim, and Gunhee Kim · 2017
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A. Shamma, Michael S. Bernstein, and Li Fei-Fei · 2017
Earlier work this paper cites.
Latent semantic minimal hashing for image retrieval
Xiaoqiang Lu, Xiangtao Zheng, and Xuelong Li · 2017
Earlier work this paper cites.
Automatic differentiation in pytorch
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer · 2017
Cited alongside, same era.
Faster R-CNN: towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross B. Girshick, and Jian Sun · 2017
Cited alongside, same era.
https://storage.googleapis.com/openimages/web/index.html/ , 2018
Open images dataset · 2018
Cited alongside, same era.
Bottom-up and top-down attention for image captioning and visual question answering
Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Cited alongside, same era.
Rpc: A large-scale retail product checkout dataset
Xiu-Shen Wei, Quan Cui, Lei Yang, Peng Wang, and Lingqiao Liu · 2019
Later among the works it cites.
Deep multi-modal latent representation learning for automated dementia diagnosis
Tao Zhou, Mingxia Liu, Huazhu Fu, Jun Wang, Jianbing Shen, Ling Shao, and Dinggang Shen · 2019
Later among the works it cites.
UNITER: universal image-text representation learning
Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu · 2020
Later among the works it cites.
Modality-agnostic attention fusion for visual search with text feedback
Eric Dodds, Jack Culpepper, Simao Herdade, Yang Zhang, and Kofi Boakye · 2020
Later among the works it cites.
Multi-modal transformer for video retrieval
Valentin Gabeur, Chen Sun, Karteek Alahari, and Cordelia Schmid · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Twitter100k: A real-world dataset for weakly supervised cross-media retrieval
Yuting Hu, Liang Zheng, Yi Yang, and Yongfeng Huang · 2018
Cited alongside, same era.
TVQA: localized, compositional video question answering
Jie Lei, Licheng Yu, Mohit Bansal, and Tamara L. Berg · 2018
Cited alongside, same era.
Spoken squad: A study of mitigating the impact of speech recognition errors on listening comprehension
Chia-Hsuan Li, Szu-Lin Wu, Chi-Liang Liu, and Hung-yi Lee · 2018
Cited alongside, same era.
An overview of cross-media retrieval: Concepts, methodologies, benchmarks, and challenges
Yuxin Peng, Xin Huang, and Yunzhen Zhao · 2018
Cited alongside, same era.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut · 2018
Cited alongside, same era.
Multimodal language analysis in the wild: CMU-MOSEI dataset and interpretable dynamic fusion graph
Amir Zadeh, Paul Pu Liang, Soujanya Poria, Erik Cambria, and Louis-Philippe Morency · 2018
Cited alongside, same era.
Towards automatic learning of procedures from web instructional videos
Luowei Zhou, Chenliang Xu, and Jason J. Corso · 2018
Cited alongside, same era.
Fashionbert: Text and image matching with adaptive loss for cross-modal retrieval
Dehong Gao, Linbo Jin, Ben Chen, Minghui Qiu, Peng Li, Yi Wei, Yi Hu, and Hao Wang · 2020
Later among the works it cites.
Does my multimodal model learn cross-modal interactions? it’s harder to tell than you might think!
Jack Hessel and Lillian Lee · 2020
Later among the works it cites.
HERO: hierarchical encoder for video+language omni-representation pre-training
Linjie Li, Yen-Chun Chen, Yu Cheng, Zhe Gan, Licheng Yu, and Jingjing Liu · 2020
Later among the works it cites.
Imagebert: Cross-modal pre-training with large-scale weak-supervised image-text data
Di Qi, Lin Su, Jia Song, Edward Cui, Taroon Bharti, and Arun Sacheti · 2020
Later among the works it cites.
VL-BERT: pre-training of generic visual-linguistic representations
Weijie Su, Xizhou Zhu, Yue Cao, Bin Li, Lewei Lu, Furu Wei, and Jifeng Dai · 2020
Later among the works it cites.
Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts
Soravit Changpinyo, Piyush Sharma, Nan Ding, and Radu Soricut · 2021
Closest in time.
Mep-3m: A large-scale multi-modal e-commerce products dataset
Delong Chen, Fan Liu, Xiaoyu Du, Ruizhuo Gao, and Feng Xu · 2021
Closest in time.
What makes multimodal learning better than single (provably)
Yu Huang, Chenzhuang Du, Zihui Xue, Xuanyao Chen, Hang Zhao, and Longbo Huang · 2021
Closest in time.
M6: A chinese multimodal pretrainer
Junyang Lin, Rui Men, An Yang, Chang Zhou, Ming Ding, Yichang Zhang, Peng Wang, Ang Wang, Le Jiang, Xianyan Jia, Jie Zhang, Jianwei Zhang, Xu Zou, Zhikang Li, Xiaodong Deng, Jie Liu, Jinbao Xue, Huiling Zhou, Jianxin Ma, Jin Yu, Yong Li, Wei Lin, Jingren Zhou, Jie Tang, and Hongxia Yang · 2021
Closest in time.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever · 2021
Closest in time.
Zero-shot text-to-image generation
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever · 2021
Closest in time.
Product1m: Towards weakly supervised instance-level product retrieval via cross-modal pretraining
Xunlin Zhan, Yangxin Wu, Xiao Dong, Yunchao Wei, Minlong Lu, Yichi Zhang, Hang Xu, and Xiaodan Liang · 2021
Closest in time.
Product1m: Towards weakly supervised instance-level product retrieval via cross-modal pretraining
Xunlin Zhan, Yangxin Wu, Xiao Dong, Yunchao Wei, Minlong Lu, Yichi Zhang, Hang Xu, and Xiaodan Liang · 2021
Closest in time.
Kaleido-bert: Vision-language pre-training on fashion domain
Mingchen Zhuge, Dehong Gao, Deng-Ping Fan, Linbo Jin, Ben Chen, Haoming Zhou, Minghui Qiu, and Ling Shao · 2021
Closest in time.