Fetching the paper…
Reading the bibliography…
With the urgent demand for generalized deep models, many pre-trained big models are proposed, such as BERT, ViT, GPT, etc.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. Lecun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
The pascal recognising textual entailment challenge
Ido Dagan, Oren Glickman, and Bernardo Magnini · 2005
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Im2text: Describing images using 1 million captioned photographs
Vicente Ordonez, Girish Kulkarni, and Tamara Berg · 2011
Earlier work this paper cites.
A three-way model for collective learning on multi-relational data
Maximilian Nickel, Volker Tresp, and Hans-Peter Kriegel · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Framing image description as a ranking task: Data, models and evaluation metrics
Micah Hodosh, Peter Young, and Julia Hockenmaier · 2013
Earlier work this paper cites.
Translating embeddings for modeling multi-relational data
Antoine Bordes, Nicolas Usunier, Alberto Garcia-Duran, Jason Weston, and Oksana Yakhnenko · 2013
Earlier work this paper cites.
Reasoning with neural tensor networks for knowledge base completion
Richard Socher, Danqi Chen, Christopher D Manning, and Andrew Ng · 2013
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D Manning · 2014
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Peter Young, Alice Lai, Micah Hodosh, and Julia Hockenmaier · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
Knowledge graph embedding by translating on hyperplanes
Zhen Wang, Jianwen Zhang, Jianlin Feng, and Zheng Chen · 2014
Earlier work this paper cites.
Embedding entities and relations for learning and inference in knowledge bases, 2014
Bishan Yang, Wen-tau Yih, Xiaodong He, Jianfeng Gao, and Li Deng · 2014
Earlier work this paper cites.
A semantic matching energy function for learning with multi-relational data
Antoine Bordes, Xavier Glorot, Jason Weston, and Yoshua Bengio · 2014
Earlier work this paper cites.
Spectral networks and locally connected networks on graphs
Joan Bruna, Wojciech Zaremba, Arthur Szlam, and Yann Lecun · 2014
Earlier work this paper cites.
Large-scale object classification using label relation graphs
Jia Deng, Nan Ding, Yangqing Jia, Andrea Frome, Kevin Murphy, Samy Bengio, Yuan Li, Hartmut Neven, and Hartwig Adam · 2014
Earlier work this paper cites.
Robust entity linking via random walks
Zhaochen Guo and Denilson Barbosa · 2014
Earlier work this paper cites.
Skip-thought vectors
Ryan Kiros, Yukun Zhu, Russ R Salakhutdinov, Richard Zemel, Raquel Urtasun, Antonio Torralba, and Sanja Fidler · 2015
Earlier work this paper cites.
Microsoft coco captions: Data collection and evaluation server
Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Dollár, and C Lawrence Zitnick · 2015
Earlier work this paper cites.
Knowledge graph embedding via dynamic mapping matrix
Guoliang Ji, Shizhu He, Liheng Xu, Kang Liu, and Jun Zhao · 2015
Earlier work this paper cites.
Learning entity and relation embeddings for knowledge graph completion
Yankai Lin, Zhiyuan Liu, Maosong Sun, Yang Liu, and Xuan Zhu · 2015
Earlier work this paper cites.
Vqa: Visual question answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
Robust multi-modal medical image fusion via anisotropic heat diffusion guided low-rank structural analysis
Qingzheng Wang, Shuai Li, Hong Qin, and Aimin Hao · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Yfcc100m: The new data in multimedia research
Bart Thomee, David A Shamma, Gerald Friedland, Benjamin Elizalde, Karl Ni, Douglas Poland, Damian Borth, and Li-Jia Li · 2016
Earlier work this paper cites.
Knowledge graph completion with adaptive sparse transfer matrix
Guoliang Ji, Kang Liu, Shizhu He, and Jun Zhao · 2016
Earlier work this paper cites.
Holographic embeddings of knowledge graphs
Maximilian Nickel, Lorenzo Rosasco, and Tomaso Poggio · 2016
Earlier work this paper cites.
Variational graph auto-encoders, 2016
Thomas N. Kipf and Max Welling · 2016
Earlier work this paper cites.
Inception-v4, inception-resnet and the impact of residual connections on learning
Christian Szegedy, Sergey Ioffe, Vincent Vanhoucke, and Alexander A Alemi · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Densely connected convolutional networks
Gao Huang, Zhuang Liu, Laurens van der Maaten, and Kilian Q. Weinberger · 2017
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, et al · 2017
Earlier work this paper cites.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh · 2017
Earlier work this paper cites.
Revisiting unreasonable effectiveness of data in deep learning era
Chen Sun, Abhinav Shrivastava, Saurabh Singh, and Abhinav Gupta · 2017
Earlier work this paper cites.
Scene parsing through ade20k dataset
Bolei Zhou, Hang Zhao, Xavier Puig, Sanja Fidler, Adela Barriuso, and Antonio Torralba · 2017
Earlier work this paper cites.
Semi-supervised classification with graph convolutional networks
Thomas N. Kipf and Max Welling · 2017
Earlier work this paper cites.
Inductive representation learning on large graphs
Will Hamilton, Zhitao Ying, and Jure Leskovec · 2017
Earlier work this paper cites.
Graph attention networks, 2017
Petar Veličković, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Liò, and Yoshua Bengio · 2017
Earlier work this paper cites.
Explicit knowledge-based reasoning for visual question answering
Peng Wang, Qi Wu, Chunhua Shen, Anthony Dick, and Anton van den Hengel · 2017
Earlier work this paper cites.
Visual dialog
Abhishek Das, Satwik Kottur, Khushi Gupta, Avi Singh, Deshraj Yadav, José MF Moura, Devi Parikh, and Dhruv Batra · 2017
Earlier work this paper cites.
A corpus of natural language for visual reasoning
Alane Suhr, Mike Lewis, James Yeh, and Yoav Artzi · 2017
Earlier work this paper cites.
Person search with natural language description
Shuang Li, Tong Xiao, Hongsheng Li, Bolei Zhou, Dayu Yue, and Xiaogang Wang · 2017
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever · 2018
Earlier work this paper cites.
Fashion-gen: The generative fashion dataset and challenge
Negar Rostamzadeh, Seyedarian Hosseini, Thomas Boquet, Wojciech Stokowiec, Ying Zhang, Christian Jauvin, and Chris Pal · 2018
Earlier work this paper cites.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut · 2018
Earlier work this paper cites.
Tvqa: Localized, compositional video question answering
Jie Lei, Licheng Yu, Mohit Bansal, and Tamara Berg · 2018
Earlier work this paper cites.
Exploring the limits of weakly supervised pretraining
Dhruv Mahajan, Ross Girshick, Vignesh Ramanathan, Kaiming He, Manohar Paluri, Yixuan Li, Ashwin Bharambe, and Laurens Van Der Maaten · 2018
Earlier work this paper cites.
Deep contextualized word representations
Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer · 2018
Earlier work this paper cites.
Modeling relational data with graph convolutional networks
Michael Schlichtkrull, Thomas N. Kipf, Peter Bloem, Rianne van den Berg, Ivan Titov, and Max Welling · 2018
Earlier work this paper cites.
Convolutional 2d knowledge graph embeddings
Tim Dettmers, Pasquale Minervini, Pontus Stenetorp, and Sebastian Riedel · 2018
Earlier work this paper cites.
Fvqa: Fact-based visual question answering
Peng Wang, Qi Wu, Chunhua Shen, Anthony Dick, and Anton van den Hengel · 2018
Earlier work this paper cites.
HotpotQA: A dataset for diverse, explainable multi-hop question answering
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William Cohen, Ruslan Salakhutdinov, and Christopher D. Manning · 2018
Earlier work this paper cites.
FEVER: a large-scale dataset for fact extraction and VERification
James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal · 2018
Earlier work this paper cites.
Contextual inter-modal attention for multi-modal sentiment analysis
Deepanway Ghosal, Md Shad Akhtar, Dushyant Chauhan, Soujanya Poria, Asif Ekbal, and Pushpak Bhattacharyya · 2018
Earlier work this paper cites.
Grounding referring expressions in images by variational context
Hanwang Zhang, Yulei Niu, and Shih-Fu Chang · 2018
Earlier work this paper cites.
Xiao Wang, Chenglong Li, Rui Yang, Tianzhu Zhang, Jin Tang, and Bin Luo · 2018
Earlier work this paper cites.
Stacked cross attention for image-text matching
Kuang-Huei Lee, Xi Chen, Gang Hua, Houdong Hu, and Xiaodong He · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin Ming-Wei Chang Kenton and Lee Kristina Toutanova · 2019
Earlier work this paper cites.
Xlnet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R Salakhutdinov, and Quoc V Le · 2019
Earlier work this paper cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
Nezha: Neural contextualized representation for chinese language understanding
Junqiu Wei, Xiaozhe Ren, Xiaoguang Li, Wenyong Huang, Yi Liao, Yasheng Wang, Jiashu Lin, Xin Jiang, Xiao Chen, and Qun Liu · 2019
Earlier work this paper cites.
wav2vec: Unsupervised pre-training for speech recognition
Steffen Schneider, Alexei Baevski, Ronan Collobert, and Michael Auli · 2019
Earlier work this paper cites.
Effectiveness of self-supervised pre-training for speech recognition
Alexei Baevski, Michael Auli, and Abdelrahman Mohamed · 2019
Earlier work this paper cites.
Gqa: A new dataset for real-world visual reasoning and compositional question answering
Drew A Hudson and Christopher D Manning · 2019
Earlier work this paper cites.
Howto100m: Learning a text-video embedding by watching hundred million narrated video clips
Antoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi, Ivan Laptev, and Josef Sivic · 2019
Earlier work this paper cites.
Lxmert: Learning cross-modality encoder representations from transformers
Hao Tan and Mohit Bansal · 2019
Earlier work this paper cites.
Unified language model pre-training for natural language understanding and generation
Li Dong, Nan Yang, Wenhui Wang, Furu Wei, Xiaodong Liu, Yu Wang, Jianfeng Gao, Ming Zhou, and Hsiao-Wuen Hon · 2019
Earlier work this paper cites.
Computational optimal transport: With applications to data science
Gabriel Peyré, Marco Cuturi, et al · 2019
Earlier work this paper cites.
Visualbert: A simple and performant baseline for vision and language
Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang · 2019
Earlier work this paper cites.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee · 2019
Earlier work this paper cites.
Fusion of detected objects in text for visual question answering
Chris Alberti, Jeffrey Ling, Michael Collins, and David Reitter · 2019
Earlier work this paper cites.
Vl-bert: Pre-training of generic visual-linguistic representations
Weijie Su, Xizhou Zhu, Yue Cao, Bin Li, Lewei Lu, Furu Wei, and Jifeng Dai · 2019
Earlier work this paper cites.
Videobert: A joint model for video and language representation learning
Chen Sun, Austin Myers, Carl Vondrick, Kevin Murphy, and Cordelia Schmid · 2019
Earlier work this paper cites.
Learning video representations using contrastive bidirectional transformer
Chen Sun, Fabien Baradel, Kevin Murphy, and Cordelia Schmid · 2019
Earlier work this paper cites.
End-to-end structure-aware convolutional networks for knowledge base completion
Chao Shang, Yun Tang, Jing Huang, Jinbo Bi, Xiaodong He, and Bowen Zhou · 2019
Earlier work this paper cites.
Learning attention-based embeddings for relation prediction in knowledge graphs
Deepak Nathani, Jatin Chauhan, Charu Sharma, and Manohar Kaul · 2019
Earlier work this paper cites.
ERNIE: Enhanced language representation with informative entities
Zhengyan Zhang, Xu Han, Zhiyuan Liu, Xin Jiang, Maosong Sun, and Qun Liu · 2019
Earlier work this paper cites.
Knowledge enhanced contextual word representations
Matthew E. Peters, Mark Neumann, Robert Logan, Roy Schwartz, Vidur Joshi, Sameer Singh, and Noah A. Smith · 2019
Earlier work this paper cites.
Natural questions: A benchmark for question answering research
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov · 2019
Earlier work this paper cites.
BoolQ: Exploring the surprising difficulty of natural yes/no questions
Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova · 2019
Earlier work this paper cites.
CommonsenseQA: A question answering challenge targeting commonsense knowledge
Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant · 2019
Earlier work this paper cites.
Social IQa: Commonsense reasoning about social interactions
Maarten Sap, Hannah Rashkin, Derek Chen, Ronan Le Bras, and Yejin Choi · 2019
Earlier work this paper cites.
“going on a vacation” takes longer than “going for a walk”: A study of temporal commonsense understanding
Ben Zhou, Daniel Khashabi, Qiang Ning, and Dan Roth · 2019
Earlier work this paper cites.
Nocaps: Novel object captioning at scale
Harsh Agrawal, Karan Desai, Yufei Wang, Xinlei Chen, Rishabh Jain, Mark Johnson, Dhruv Batra, Devi Parikh, Stefan Lee, and Peter Anderson · 2019
Earlier work this paper cites.
Visual entailment: A novel task for fine-grained image understanding
Ning Xie, Farley Lai, Derek Doran, and Asim Kadav · 2019
Earlier work this paper cites.
From recognition to cognition: Visual commonsense reasoning
Rowan Zellers, Yonatan Bisk, Ali Farhadi, and Yejin Choi · 2019
Cited alongside, same era.
Cross-modal relationship inference for grounding referring expressions
Sibei Yang, Guanbin Li, and Yizhou Yu · 2019
Cited alongside, same era.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Cited alongside, same era.
Align before fuse: Vision and language representation learning with momentum distillation
Junnan Li, Ramprasaath Selvaraju, Akhilesh Gotmare, Shafiq Joty, Caiming Xiong, and Steven Chu Hong Hoi · 2021
Later among the works it cites.
Proposal-free one-stage referring expression via grid-word cross-attention
Wei Suo, Mengyang Sun, Peng Wang, and Qi Wu · 2021
Later among the works it cites.
Actionclip: A new paradigm for video action recognition
Mengmeng Wang, Jiazheng Xing, and Yong Liu · 2021
Later among the works it cites.
How much can clip benefit vision-and-language tasks?
Sheng Shen, Liunian Harold Li, Hao Tan, Mohit Bansal, Anna Rohrbach, Kai-Wei Chang, Zhewei Yao, and Kurt Keutzer · 2021
Later among the works it cites.
Kvl-bert: Knowledge enhanced visual-and-linguistic bert for visual commonsense reasoning
Dandan Song, Siyi Ma, Zhanchen Sun, Sicheng Yang, and Lejian Liao · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Oscar: Object-semantics aligned pre-training for vision-language tasks
Xiujun Li, Xi Yin, Chunyuan Li, Pengchuan Zhang, Xiaowei Hu, Lei Zhang, Lijuan Wang, Houdong Hu, Li Dong, Furu Wei, et al · 2020
Cited alongside, same era.
Uniter: Universal image-text representation learning
Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu · 2020
Cited alongside, same era.
Pixel-bert: Aligning image pixels with text by deep multi-modal transformers
Zhicheng Huang, Zhaoyang Zeng, Bei Liu, Dongmei Fu, and Jianlong Fu · 2020
Cited alongside, same era.
A short survey of pre-trained language models for conversational ai-a new age in nlp
Munazza Zaib, Quan Z Sheng, and Wei Emma Zhang · 2020
Cited alongside, same era.
A survey on contextual embeddings
Qi Liu, Matt J Kusner, and Phil Blunsom · 2020
Cited alongside, same era.
Pre-trained models for natural language processing: A survey
Xipeng Qiu, Tianxiang Sun, Yige Xu, Yunfan Shao, Ning Dai, and Xuanjing Huang · 2020
Cited alongside, same era.
A multi-layer bidirectional transformer encoder for pre-trained word embedding: A survey of bert
Rohit Kumar Kaliyar · 2020
Cited alongside, same era.
Later among the works it cites.
Unifying vision-and-language tasks via text generation
Jaemin Cho, Jie Lei, Hao Tan, and Mohit Bansal · 2021
Later among the works it cites.
Vilt: Vision-and-language transformer without convolution or region supervision
Wonjae Kim, Bokyung Son, and Ildoo Kim · 2021
Later among the works it cites.
Mdetr-modulated detection for end-to-end multi-modal understanding
Aishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve, Ishan Misra, and Nicolas Carion · 2021
Later among the works it cites.
Seeing out of the box: End-to-end pre-training for vision-language representation learning
Zhicheng Huang, Zhaoyang Zeng, Yupan Huang, Bei Liu, Dongmei Fu, and Jianlong Fu · 2021
Later among the works it cites.
Probing inter-modality: Visual parsing with self-attention for vision-and-language pre-training
Hongwei Xue, Yupan Huang, Bei Liu, Houwen Peng, Jianlong Fu, Houqiang Li, and Jiebo Luo · 2021
Later among the works it cites.
Mural: multimodal, multitask retrieval across languages
Aashi Jain, Mandy Guo, Krishna Srinivasan, Ting Chen, Sneha Kudugunta, Chao Jia, Yinfei Yang, and Jason Baldridge · 2021
Later among the works it cites.
Vlmo: Unified vision-language pre-training with mixture-of-modality-experts
Wenhui Wang, Hangbo Bao, Li Dong, and Furu Wei · 2021
Later among the works it cites.
An empirical study of training end-to-end vision-and-language transformers
Zi-Yi Dou, Yichong Xu, Zhe Gan, Jianfeng Wang, Shuohang Wang, Lijuan Wang, Chenguang Zhu, Zicheng Liu, Michael Zeng, et al · 2021
Later among the works it cites.
Video-text pre-training with learned regions
Rui Yan, Mike Zheng Shou, Yixiao Ge, Alex Jinpeng Wang, Xudong Lin, Guanyu Cai, and Jinhui Tang · 2021
Later among the works it cites.
Zero-shot text-to-image generation
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever · 2021
Later among the works it cites.
Cogview: Mastering text-to-image generation via transformers
Ming Ding, Zhuoyi Yang, Wenyi Hong, Wendi Zheng, Chang Zhou, Da Yin, Junyang Lin, Xu Zou, Zhou Shao, Hongxia Yang, et al · 2021
Later among the works it cites.
Vatt: Transformers for multimodal self-supervised learning from raw video, audio and text
Hassan Akbari, Liangzhe Yuan, Rui Qian, Wei-Hong Chuang, Shih-Fu Chang, Yin Cui, and Boqing Gong · 2021
Later among the works it cites.
Florence: A new foundation model for computer vision
Lu Yuan, Dongdong Chen, Yi-Ling Chen, Noel Codella, Xiyang Dai, Jianfeng Gao, Houdong Hu, Xuedong Huang, Boxin Li, Chunyuan Li, et al · 2021
Later among the works it cites.
Gilbert: Generative vision-language pre-training for image-text retrieval
Weixiang Hong, Kaixiang Ji, Jiajia Liu, Jian Wang, Jingdong Chen, and Wei Chu · 2021
Later among the works it cites.
Unsupervised vision-and-language pre-training without parallel images and captions
Liunian Harold Li, Haoxuan You, Zhecan Wang, Alireza Zareian, Shih-Fu Chang, and Kai-Wei Chang · 2021
Later among the works it cites.
M3p: Learning universal representations via multitask multilingual multimodal pre-training
Minheng Ni, Haoyang Huang, Lin Su, Edward Cui, Taroon Bharti, Lijuan Wang, Dongdong Zhang, and Nan Duan · 2021
Later among the works it cites.
Regionclip: Region-based language-image pretraining
Yiwu Zhong, Jianwei Yang, Pengchuan Zhang, Chunyuan Li, Noel Codella, Liunian Harold Li, Luowei Zhou, Xiyang Dai, Lu Yuan, Yin Li, et al · 2021
Later among the works it cites.
Slip: Self-supervision meets language-image pre-training
Norman Mu, Alexander Kirillov, David Wagner, and Saining Xie · 2021
Later among the works it cites.
Filip: Fine-grained interactive language-image pre-training
Lewei Yao, Runhui Huang, Lu Hou, Guansong Lu, Minzhe Niu, Hang Xu, Xiaodan Liang, Zhenguo Li, Xin Jiang, and Chunjing Xu · 2021
Later among the works it cites.
Semvlp: Vision-language pre-training by aligning semantics at multiple levels
Chenliang Li, Ming Yan, Haiyang Xu, Fuli Luo, Wei Wang, Bin Bi, and Songfang Huang · 2021
Later among the works it cites.
Do syntax trees help pre-trained transformers extract information?
Devendra Sachan, Yuhao Zhang, Peng Qi, and William L. Hamilton · 2021
Later among the works it cites.
Temporal reasoning on implicit events from distant supervision
Ben Zhou, Kyle Richardson, Qiang Ning, Tushar Khot, Ashish Sabharwal, and Dan Roth · 2021
Later among the works it cites.
Deep image retrieval: A survey
Wei Chen, Yang Liu, Weiping Wang, Erwin M Bakker, TK Georgiou, Paul Fieguth, Li Liu, and MSK Lew · 2021
Later among the works it cites.
Human-centric spatio-temporal video grounding with visual transformers
Zongheng Tang, Yue Liao, Si Liu, Guanbin Li, Xiaojie Jin, Hongxu Jiang, Qian Yu, and Dong Xu · 2021
Later among the works it cites.
Towards more flexible and accurate object tracking with natural language: Algorithms and benchmark
Xiao Wang, Xiujun Shu, Zhipeng Zhang, Bo Jiang, Yaowei Wang, Yonghong Tian, and Feng Wu · 2021
Later among the works it cites.
Siamese natural language tracker: Tracking by natural language descriptions with siamese trackers
Qi Feng, Vitaly Ablavsky, Qinxun Bai, and Stan Sclaroff · 2021
Later among the works it cites.
Cpt: Colorful prompt tuning for pre-trained vision-language models
Yuan Yao, Ao Zhang, Zhengyan Zhang, Zhiyuan Liu, Tat-Seng Chua, and Maosong Sun · 2021
Later among the works it cites.
Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig · 2021
Later among the works it cites.
Hybrid dynamic contrast and probability distillation for unsupervised person re-id
De Cheng, Jingyu Zhou, Nannan Wang, and Xinbo Gao · 2022
Later among the works it cites.
Vlp: A survey on vision-language pre-training
Feilong Chen, Duzhen Zhang, Minglun Han, Xiuyi Chen, Jing Shi, Shuang Xu, and Bo Xu · 2022
Later among the works it cites.
A survey of vision-language pre-trained models
Yifan Du, Zikang Liu, Junyi Li, and Wayne Xin Zhao · 2022
Later among the works it cites.
A survey of controllable text generation using transformer-based pre-trained language models
Hanqing Zhang, Haolin Song, Shaoyu Li, Ming Zhou, and Dawei Song · 2022
Later among the works it cites.
A survey of knowledge-intensive nlp with pre-trained language models
Da Yin, Li Dong, Hao Cheng, Xiaodong Liu, Kai-Wei Chang, Furu Wei, and Jianfeng Gao · 2022
Later among the works it cites.
Commonsense knowledge reasoning and generation with pre-trained language models: A survey
Prajjwal Bhargava and Vincent Ng · 2022
Later among the works it cites.
A survey of vision-language pre-trained models
Yifan Du, Zikang Liu, Junyi Li, and Wayne Xin Zhao · 2022
Later among the works it cites.
Survey: Transformer based video-language pre-training
Ludan Ruan and Qin Jin · 2022
Later among the works it cites.
Vision-language intelligence: Tasks, representation learning, and large models
Feng Li, Hao Zhang, Yi-Fan Zhang, Shilong Liu, Jian Guo, Lionel M Ni, PengChuan Zhang, and Lei Zhang · 2022
Later among the works it cites.
A survey on vision transformer
Kai Han, Yunhe Wang, Hanting Chen, Xinghao Chen, Jianyuan Guo, Zhenhua Liu, Yehui Tang, An Xiao, Chunjing Xu, Yixing Xu, et al · 2022
Later among the works it cites.
Javier Selva, Anders S Johansen, Sergio Escalera, Kamal Nasrollahi, Thomas B Moeslund, and Albert Clapés · 2022
Later among the works it cites.
Threats to pre-trained language models: Survey and taxonomy
Shangwei Guo, Chunlong Xie, Jiwei Li, Lingjuan Lyu, and Tianwei Zhang · 2022
Later among the works it cites.
Sha Yuan, Hanyu Zhao, Shuai Zhao, Jiahong Leng, Yangxiao Liang, Xiaozhi Wang, Jifan Yu, Xin Lv, Zhou Shao, Jiaao He, et al · 2022
Later among the works it cites.
Vision-and-language pretrained models: A survey
Soyeon Caren Han Siqu Long, Feiqi Cao and Haiqing Yang · 2022
Later among the works it cites.
Multimodal learning with transformers: A survey
Xu Peng, Zhu Xiatian, and A. Clifton David · 2022
Later among the works it cites.
Prompt-based learning for unpaired image captioning
Peipei Zhu, Xiao Wang, Lin Zhu, Zhenglong Sun, Weishi Zheng, Yaowei Wang, and Changwen Chen · 2022
Later among the works it cites.
Class-aware visual prompt tuning for vision-language pre-trained model
Yinghui Xing, Qirui Wu, De Cheng, Shizhou Zhang, Guoqiang Liang, and Yanning Zhang · 2022
Later among the works it cites.
Wukong: 100 million large-scale chinese cross-modal pre-training dataset and a foundation framework, 2022
Jiaxi Gu, Xiaojun Meng, Guansong Lu, Lu Hou, Minzhe Niu, Hang Xu, Xiaodan Liang, Wei Zhang, Xin Jiang, and Chunjing Xu · 2022
Later among the works it cites.
Wudaomm: A large-scale multi-modal dataset for pre-training models
Leng Jiahong Xue Zhao Zhao Hanyu Sha Yuan, Zhao Shuai and Tang Jie · 2022
Later among the works it cites.
Vision-language pre-training for multimodal aspect-based sentiment analysis
Yan Ling, Rui Xia, et al · 2022
Later among the works it cites.
Attention mechanisms in computer vision: A survey
Meng-Hao Guo, Tian-Xing Xu, Jiang-Jiang Liu, Zheng-Ning Liu, Peng-Tao Jiang, Tai-Jiang Mu, Song-Hai Zhang, Ralph R Martin, Ming-Ming Cheng, and Shi-Min Hu · 2022
Later among the works it cites.
i-code: An integrative and composable multimodal learning framework
Ziyi Yang, Yuwei Fang, Chenguang Zhu, Reid Pryzant, Dongdong Chen, Yu Shi, Yichong Xu, Yao Qian, Mei Gao, Yi-Ling Chen, et al · 2022
Later among the works it cites.
Clip-event: Connecting text and images with event structures
Manling Li, Ruochen Xu, Shuohang Wang, Luowei Zhou, Xudong Lin, Chenguang Zhu, Michael Zeng, Heng Ji, and Shih-Fu Chang · 2022
Later among the works it cites.
Yufeng Cui, Lichen Zhao, Feng Liang, Yangguang Li, and Jing Shao · 2022
Later among the works it cites.
Prototypical contrastive language image pretraining
Chen Delong, Wu Zhao, Liu Fan, Yang Zaiquan, Huang Yixiang, Bao Yiping, and Zhou Erjin · 2022
Later among the works it cites.
Pyramidclip: Hierarchical feature alignment for vision-language model pretraining
Gao Yuting, Liu Jinfeng, Xu Zihan, Zhang Jun, Li Ke, and Shen Chunhua · 2022
Later among the works it cites.
Training vision-language transformers from captions alone
Alex Hauptmann Yonatan Bisk Jianfeng Gao Liangke Gui, Qiuyuan Huang · 2022
Later among the works it cites.
Hivlp: Hierarchical vision-language pre-training for fast image-text retrieval
Mickael Coustaty Marçal Rusiñol Oriol Ramos Terrades Souhail Bakkali, Zuheng Ming · 2022
Later among the works it cites.
Mvp: Multimodality-guided visual pre-training
Longhui Wei, Lingxi Xie, Wengang Zhou, Houqiang Li, and Qi Tian · 2022
Later among the works it cites.
Cots: Collaborative two-stream vision-language pre-training model for cross-modal retrieval
Yuqi Huo Yizhao Gao Zhiwu Lu Ji-Rong Wen Haoyu Lu, Nanyi Fei · 2022
Later among the works it cites.
Flamingo: a visual language model for few-shot learning
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katie Millican, Malcolm Reynolds, et al · 2022
Later among the works it cites.
Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation, 2022
Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi · 2022
Later among the works it cites.
Vision-language pre-training with triple contrastive learning
Jinyu Yang, Jiali Duan, Son Tran, Yi Xu, Sampath Chanda, Liqun Chen, Belinda Zeng, Trishul Chilimbi, and Junzhou Huang · 2022
Later among the works it cites.
Clinical-bert: Vision-language pre-training for radiograph diagnosis and reports generation
Bin Yan and Mingtao Pei · 2022
Later among the works it cites.
Visual-language navigation pretraining via prompt-based environmental self-exploration
Xiwen Liang, Fengda Zhu, Lingling Li, Hang Xu, and Xiaodan Liang · 2022
Later among the works it cites.
Grounded language-image pre-training
Liunian Harold Li*, Pengchuan Zhang*, Haotian Zhang*, Jianwei Yang, Chunyuan Li, Yiwu Zhong, Lijuan Wang, Lu Yuan, Lei Zhang, Jenq-Neng Hwang, Kai-Wei Chang, and Jianfeng Gao · 2022
Later among the works it cites.
Zero and r2d2: A large-scale chinese cross-modal benchmark and a vision-language framework
Xie Chunyu, Cai Heng, Song Jianfei, Li Jincheng, Kong Fanjing, Wu Xiaoyu, Morimitsu Henrique, Yao Lin, Wang Dexin, Leng Dawei, Ji Xiangyang, and Deng Yafeng · 2022
Later among the works it cites.
Coca: Contrastive captioners are image-text foundation models
Jiahui Yu, Zirui Wang, Vijay Vasudevan, Legg Yeung, Mojtaba Seyedhosseini, and Yonghui Wu · 2022
Later among the works it cites.
Hivlp: Hierarchical vision-language pre-training for fast image-text retrieval
Jiaxin Shi Duzhen Zhang Jianlong Chang Feilong Chen, Xiuyi Chen and Qi Tian · 2022
Later among the works it cites.
Audioclip: Extending clip to image, text and audio
Andrey Guzhov, Federico Raue, Jörn Hees, and Andreas Dengel · 2022
Later among the works it cites.
Vl-beit: Generative vision-language pretraining
Hangbo Bao, Wenhui Wang, Li Dong, and Furu Wei · 2022
Later among the works it cites.
End-to-end generative pretraining for multimodal video captioning
Paul Hongsuck Seo, Arsha Nagrani, Anurag Arnab, and Cordelia Schmid · 2022
Later among the works it cites.
A unified continuous learning framework for multi-modal knowledge discovery and pre-training
Fan Zhihao, Wei Zhongyu, Chen Jingjing, Wang Siyuan, Li Zejun, Xu Jiarong, and Huang Xuanjing · 2022
Later among the works it cites.
Glipv2: Unifying localization and vision-language understanding
Zhang Haotian, Zhang Pengchuan, Hu Xiaowei, Chen Yen-Chun, Harold Li Liunian, Dai Xiyang, Wang Lijuan, Yuan Lu, Hwang Jenq-Neng, and Gao Jianfeng · 2022
Later among the works it cites.
Multimodal contrastive learning with limoe: the language-image mixture of experts
Mustafa Basil, Riquelme Carlos, Puigcerver Joan, Jenatton Rodolphe, and Houlsby Neil · 2022
Later among the works it cites.
Vlmixer: Unpaired vision-language pre-training via cross-modal cutmix
Wang Teng, Jiang Wenhao, Lu Zhichao, Zheng Feng, Cheng Ran, Yin Chengguo, and Ping Luo · 2022
Later among the works it cites.
Pedestrian attribute recognition: A survey
Xiao Wang, Shaofei Zheng, Rui Yang, Aihua Zheng, Zhe Chen, Jin Tang, and Bin Luo · 2022
Later among the works it cites.
Vision-and-language navigation: A survey of tasks, methods, and future directions
Jing Gu, Eliana Stefani, Qi Wu, Jesse Thomason, and Xin Eric Wang · 2022
Later among the works it cites.
Visual language navigation: a survey and open challenges
Sang-Min Park and Young-Gab Kim · 2022
Later among the works it cites.
Exploring language hierarchy for video grounding
Xinpeng Ding, Nannan Wang, Shiwei Zhang, Ziyuan Huang, Xiaomeng Li, Mingqian Tang, Tongliang Liu, and Xinbo Gao · 2022
Later among the works it cites.
Cpl: Counterfactual prompt learning for vision and language models
Xuehai He, Diji Yang, Weixi Feng, Tsu-Jui Fu, Arjun Akula, Varun Jampani, Pradyumna Narayana, Sugato Basu, William Yang Wang, and Xin Eric Wang · 2022
Later among the works it cites.
Menglin Jia, Luming Tang, Bor-Chun Chen, Claire Cardie, Serge Belongie, Bharath Hariharan, and Ser-Nam Lim · 2022
Later among the works it cites.
Learning to prompt for vision-language models
Kaiyang Zhou, Jingkang Yang, Chen Change Loy, and Ziwei Liu · 2022
Later among the works it cites.
Conditional prompt learning for vision-language models
Kaiyang Zhou, Jingkang Yang, Chen Change Loy, and Ziwei Liu · 2022
Later among the works it cites.
Mfgnet: Dynamic modality-aware filter generation for rgb-t tracking
Xiao Wang, Xiujun Shu, Shilliang Zhang, Bo Jiang, Yaowei Wang, Yonghong Tian, and Feng Wu · 2022
Later among the works it cites.