Fetching the paper…
Reading the bibliography…
Vision-language (V+L) pretraining models have achieved great success in supporting multimedia applications by understanding the alignments between images and text.
A relationship between arbitrary positive matrices and doubly stochastic matrices
Richard Sinkhorn · 1964
Earlier work this paper cites.
Image retrieval: Ideas, influences, and trends of the new age
Ritendra Datta, Dhiraj Joshi, Jia Li, and James Z Wang · 2008
Earlier work this paper cites.
Modeling mutual context of object and human pose in human-object interaction activities
Bangpeng Yao and Li Fei-Fei · 2010
Earlier work this paper cites.
Open language learning for information extraction
Michael Schmitz, Stephen Soderland, Robert Bart, Oren Etzioni, et al · 2012
Earlier work this paper cites.
Unstructured human activity detection from rgbd images
Jaeyong Sung, Colin Ponce, Bart Selman, and Ashutosh Saxena · 2012
Earlier work this paper cites.
Sinkhorn distances: Lightspeed computation of optimal transport
Marco Cuturi · 2013
Earlier work this paper cites.
A thousand frames in just a few words: Lingual description of videos through latent topics and sparse object stitching
Pradipto Das, Chenliang Xu, Richard F Doell, and Jason J Corso · 2013
Earlier work this paper cites.
The stanford corenlp natural language processing toolkit
Christopher D Manning, Mihai Surdeanu, John Bauer, Jenny Rose Finkel, Steven Bethard, and David McClosky · 2014
Earlier work this paper cites.
Leveraging linguistic structure for open domain information extraction
Gabor Angeli, Melvin Jose Johnson Premkumar, and Christopher D Manning · 2015
Earlier work this paper cites.
Hico: A benchmark for recognizing human-object interactions in images
Yu-Wei Chao, Zhan Wang, Yugeng He, Jiaxuan Wang, and Jia Deng · 2015
Earlier work this paper cites.
Saurabh Gupta and Jitendra Malik · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun · 2015
Earlier work this paper cites.
Visual relationship detection with language priors
Cewu Lu, Ranjay Krishna, Michael Bernstein, and Li Fei-Fei · 2016
Earlier work this paper cites.
Situation recognition: Visual semantic role labeling for image understanding
Mark Yatskar, Luke Zettlemoyer, and Ali Farhadi · 2016
Earlier work this paper cites.
Foil it! find one mismatch between image and language caption
Ravi Shekhar, Sandro Pezzelle, Yauhen Klimovich, Aurélie Herbelot, Moin Nabi, Enver Sangineto, and Raffaella Bernardi · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Compositional learning for human object interaction
Keizo Kato, Yin Li, and Abhinav Gupta · 2018
Cited alongside, same era.
Fixing weight decay regularization in adam
Ilya Loshchilov and Frank Hutter · 2018
Cited alongside, same era.
Transferable interactiveness knowledge for human-object interaction detection
Yong-Lu Li, Siyuan Zhou, Xijie Huang, Liang Xu, Ze Ma, Hao-Shu Fang, Yanfeng Wang, and Cewu Lu · 2019
Cited alongside, same era.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee · 2019
Cited alongside, same era.
Lxmert: Learning cross-modality encoder representations from transformers
A joint neural model for information extraction with global features
Ying Lin, Heng Ji, Fei Huang, and Lingfei Wu · 2020
Later among the works it cites.
Visualcomet: Reasoning about the dynamic context of a still image
Jae Sung Park, Chandra Bhagavatula, Roozbeh Mottaghi, Ali Farhadi, and Yejin Choi · 2020
Later among the works it cites.
Grounded situation recognition
Sarah Pratt, Mark Yatskar, Luca Weihs, Ali Farhadi, and Aniruddha Kembhavi · 2020
Later among the works it cites.
Learning human-object interaction detection using interaction points
Tiancai Wang, Tong Yang, Martin Danelljan, Fahad Shahbaz Khan, Xiangyu Zhang, and Jian Sun · 2020
Later among the works it cites.
Ernie-vil: Knowledge enhanced vision-language representations through scene graph
Fei Yu, Jiji Tang, Weichong Yin, Yu Sun, Hao Tian, Hua Wu, and Haifeng Wang · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hao Tan and Mohit Bansal · 2019
Cited alongside, same era.
Scalable gromov-wasserstein learning for graph partitioning and matching
Hongteng Xu, Dixin Luo, and Lawrence Carin · 2019
Cited alongside, same era.
From recognition to cognition: Visual commonsense reasoning
Rowan Zellers, Yonatan Bisk, Ali Farhadi, and Yejin Choi · 2019
Cited alongside, same era.
Graph optimal transport for cross-domain alignment
Liqun Chen, Zhe Gan, Yu Cheng, Linjie Li, Lawrence Carin, and Jingjing Liu · 2020
Cited alongside, same era.
Uniter: Universal image-text representation learning
Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu · 2020
Cited alongside, same era.
Gpt-3: Its nature, scope, limits, and consequences
Luciano Floridi and Massimo Chiriatti · 2020
Cited alongside, same era.
The open images dataset v4
Alina Kuznetsova, Hassan Rom, Neil Alldrin, Jasper Uijlings, Ivan Krasin, Jordi Pont-Tuset, Shahab Kamali, Stefan Popov, Matteo Malloci, Alexander Kolesnikov, et al · 2020
Cited alongside, same era.
Alireza Zareian, Svebor Karaman, and Shih-Fu Chang · 2020
Later among the works it cites.
Unified vision-language pre-training for image captioning and vqa
Luowei Zhou, Hamid Palangi, Lei Zhang, Houdong Hu, Jason Corso, and Jianfeng Gao · 2020
Later among the works it cites.
Cascaded human-object interaction recognition
Tianfei Zhou, Wenguan Wang, Siyuan Qi, Haibin Ling, and Jianbing Shen · 2020
Later among the works it cites.
Probing image-language transformers for verb understanding
Lisa Anne Hendricks and Aida Nematzadeh · 2021
Later among the works it cites.
Seeing out of the box: End-to-end pre-training for vision-language representation learning
Zhicheng Huang, Zhaoyang Zeng, Yupan Huang, Bei Liu, Dongmei Fu, and Jianlong Fu · 2021
Later among the works it cites.
Scaling up visual and vision-language representation learning with noisy text supervision
Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc V Le, Yunhsuan Sung, Zhen Li, and Tom Duerig · 2021
Later among the works it cites.
Vilt: Vision-and-language transformer without convolution or region supervision
Wonjae Kim, Bokyung Son, and Ildoo Kim · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al · 2021
Later among the works it cites.
Simvlm: Simple visual language model pretraining with weak supervision
Zirui Wang, Jiahui Yu, Adams Wei Yu, Zihang Dai, Yulia Tsvetkov, and Yuan Cao · 2021
Later among the works it cites.
Vinvl: Revisiting visual representations in vision-language models
Pengchuan Zhang, Xiujun Li, Xiaowei Hu, Jianwei Yang, Lei Zhang, Lijuan Wang, Yejin Choi, and Jianfeng Gao · 2021
Later among the works it cites.