Fetching the paper…
Reading the bibliography…
Multimodality Representation Learning, as a technique of learning to embed information from different modalities and their correlations, has achieved remarkable success on a variety of applications, such as Visual Question Answering (VQA), Natural Language for Visual Reasoning (NLVR), and Vision Language Retrieval (VLR).
Hearing lips and seeing voices
Harry McGurk and John MacDonald · 1976
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Comparison of automatic shot boundary detection algorithms
Rainer W Lienhart · 1998
Earlier work this paper cites.
Active appearance models
Timothy F. Cootes, Gareth J. Edwards, and Christopher J. Taylor · 2001
Earlier work this paper cites.
The ami meeting corpus: A pre-announcement
Jean Carletta, Simone Ashby, Sebastien Bourban, Mike Flynn, Mael Guillemot, Thomas Hain, Jaroslav Kadlec, Vasilis Karaiskos, Wessel Kraaij, Melissa Kronenthal, Guillaume Lathoud, Mike Lincoln, Masson Agnes Lisowska, Iain McCowan, Wilfried Post, Dennis Reidsma, and Pierre D. Wellner · 2006
Earlier work this paper cites.
Extracting and composing robust features with denoising autoencoders
Pascal Vincent, Hugo Larochelle, Yoshua Bengio, and Pierre-Antoine Manzagol · 2008
Earlier work this paper cites.
Automated flower classification over a large number of classes
Maria-Elena Nilsback and Andrew Zisserman · 2008
Earlier work this paper cites.
Multimodal fusion for multimedia analysis: a survey
Pradeep K Atrey, M Anwar Hossain, Abdulmotaleb El Saddik, and Mohan S Kankanhalli · 2010
Earlier work this paper cites.
The semaine corpus of emotionally coloured character interactions
Gary McKeown, Michel F Valstar, Roderick Cowie, and Maja Pantic · 2010
Earlier work this paper cites.
A new approach to cross-modal multimedia retrieval
Nikhil Rasiwasia, Jose Costa Pereira, Emanuele Coviello, Gabriel Doyle, Gert RG Lanckriet, Roger Levy, and Nuno Vasconcelos · 2010
Earlier work this paper cites.
Avec 2011–the first international audio/visual emotion challenge
Björn Schuller, Michel Valstar, Florian Eyben, Gary McKeown, Roddy Cowie, and Maja Pantic · 2011
Earlier work this paper cites.
Multimodal deep learning
Jiquan Ngiam, Aditya Khosla, Mingyu Kim, Juhan Nam, Honglak Lee, and Andrew Y Ng · 2011
Earlier work this paper cites.
The caltech-ucsd birds-200-2011 dataset
Catherine Wah, Steve Branson, Peter Welinder, Pietro Perona, and Serge Belongie · 2011
Earlier work this paper cites.
Indoor scene segmentation using a structured light sensor
N. Silberman and R. Fergus · 2011
Earlier work this paper cites.
Indoor segmentation and support inference from rgbd images
Nathan Silberman, Derek Hoiem, Pushmeet Kohli, and Rob Fergus · 2012
Earlier work this paper cites.
Multimodal saliency and fusion for movie summarization based on aural, visual, and textual attention
Georgios Evangelopoulos, Athanasia Zlatintsi, Alexandros Potamianos, Petros Maragos, Konstantinos Rapantzikos, Georgios Skoumas, and Yannis Avrithis · 2013
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean · 2013
Earlier work this paper cites.
Framing image description as a ranking task: Data, models and evaluation metrics
Micah Hodosh, Peter Young, and Julia Hockenmaier · 2013
Earlier work this paper cites.
The mit stata center dataset
Maurice Fallon, Hordur Johannsson, Michael Kaess, and John J Leonard · 2013
Earlier work this paper cites.
Avec 2014: 3d dimensional affect and depression recognition challenge
Michel Valstar, Björn Schuller, Kirsty Smith, Timur Almaev, Florian Eyben, Jarek Krajewski, Roddy Cowie, and Maja Pantic · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
Vqa: Visual question answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton · 2015
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2015
Earlier work this paper cites.
Vqa: Visual question answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models
Bryan A Plummer, Liwei Wang, Chris M Cervantes, Juan C Caicedo, Julia Hockenmaier, and Svetlana Lazebnik · 2015
Earlier work this paper cites.
Show and tell: A neural image caption generator
Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan · 2015
Earlier work this paper cites.
Tcd-timit: An audio-visual corpus of continuous speech
Naomi Harte and Eoin Gillen · 2015
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2015
Earlier work this paper cites.
Video2vec embeddings recognize events when examples are scarce
Amirhossein Habibian, Thomas Mensink, and Cees GM Snoek · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Revisiting visual question answering baselines
Allan Jabri, Armand Joulin, and Laurens van der Maaten · 2016
Earlier work this paper cites.
Gated graph sequence neural networks
Yujia Li, Richard Zemel, Marc Brockschmidt, and Daniel Tarlow · 2016
Earlier work this paper cites.
Yfcc100m: The new data in multimedia research
Bart Thomee, David A Shamma, Gerald Friedland, Benjamin Elizalde, Karl Ni, Douglas Poland, Damian Borth, and Li-Jia Li · 2016
Earlier work this paper cites.
Topic-based content and sentiment analysis of ebola virus on twitter and in the news
Erin Hea-Jin Kim, Yoo Kyung Jeong, Yuyoung Kim, Keun Young Kang, and Min Song · 2016
Earlier work this paper cites.
Attentive explanations: Justifying decisions and pointing to the evidence
Dong Huk Park, Lisa Anne Hendricks, Zeynep Akata, Bernt Schiele, Trevor Darrell, and Marcus Rohrbach · 2016
Earlier work this paper cites.
Generative adversarial text to image synthesis
Scott Reed, Zeynep Akata, Xinchen Yan, Lajanugen Logeswaran, Bernt Schiele, and Honglak Lee · 2016
Earlier work this paper cites.
Mimic-iii, a freely accessible critical care database
Alistair EW Johnson, Tom J Pollard, Lu Shen, Li-wei H Lehman, Mengling Feng, Mohammad Ghassemi, Benjamin Moody, Peter Szolovits, Leo Anthony Celi, and Roger G Mark · 2016
Earlier work this paper cites.
Msr-vtt: A large video description dataset for bridging video and language
Jun Xu, Tao Mei, Ting Yao, and Yong Rui · 2016
Earlier work this paper cites.
Mosi: multimodal corpus of sentiment intensity and subjectivity analysis in online opinion videos
Amir Zadeh, Rowan Zellers, Eli Pincus, and Louis-Philippe Morency · 2016
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2017
Earlier work this paper cites.
Deep multimodal learning: A survey on recent advances and trends
Dhanesh Ramachandram and Graham W Taylor · 2017
Earlier work this paper cites.
Visual question answering: A survey of methods and datasets
Qi Wu, Damien Teney, Peng Wang, Chunhua Shen, Anthony Dick, and Anton Van Den Hengel · 2017
Earlier work this paper cites.
Cascade recurrent neural network for image caption generation
Jie Wu and Haifeng Hu · 2017
Earlier work this paper cites.
Graph attention networks
Petar Velickovic, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Lio, and Yoshua Bengio · 2017
Earlier work this paper cites.
Amc: Attention guided multi-modal correlation learning for image search
Kan Chen, Trung Bui, Chen Fang, Zhaowen Wang, and Ram Nevatia · 2017
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A. Shamma, Michael S. Bernstein, and Li Fei-Fei · 2017
Earlier work this paper cites.
On the importance of super-gaussian speech priors for machine-learning based speech enhancement
Robert Rehr and Timo Gerkmann · 2017
Earlier work this paper cites.
Adaptive online event detection in news streams
Linmei Hu, Bin Zhang, Lei Hou, and Juanzi Li · 2017
Earlier work this paper cites.
Multimodal fusion with recurrent neural networks for rumor detection on microblogs
Zhiwei Jin, Juan Cao, Han Guo, Yongdong Zhang, and Jiebo Luo · 2017
Earlier work this paper cites.
Sca-cnn: Spatial and channel-wise attention in convolutional networks for image captioning
Long Chen, Hanwang Zhang, Jun Xiao, Liqiang Nie, Jian Shao, Wei Liu, and Tat-Seng Chua · 2017
Earlier work this paper cites.
Self-critical sequence training for image captioning
Steven J Rennie, Etienne Marcheret, Youssef Mroueh, Jerret Ross, and Vaibhava Goel · 2017
Earlier work this paper cites.
Vid2speech: speech reconstruction from silent video
Ariel Ephrat and Shmuel Peleg · 2017
Earlier work this paper cites.
Fashion 200K Benchmark
Xintong Han · 2017
Earlier work this paper cites.
Video question answering via gradually refined attention over appearance and motion
Dejing Xu, Zhou Zhao, Jun Xiao, Fei Wu, Hanwang Zhang, Xiangnan He, and Yueting Zhuang · 2017
Earlier work this paper cites.
Tgif-qa: Toward spatio-temporal reasoning in visual question answering
Yunseok Jang, Yale Song, Youngjae Yu, Youngjin Kim, and Gunhee Kim · 2017
Earlier work this paper cites.
An analysis of visual question answering algorithms
Kushal Kafle and Christopher Kanan · 2017
Earlier work this paper cites.
Multimodal machine learning: A survey and taxonomy
Tadas Baltrušaitis, Chaitanya Ahuja, and Louis-Philippe Morency · 2018
Earlier work this paper cites.
Improved fusion of visual and language representations by dense symmetric co-attention for visual question answering
Duy-Kien Nguyen and Takayuki Okatani · 2018
Earlier work this paper cites.
Video captioning via hierarchical reinforcement learning
Xin Wang, Wenhu Chen, Jiawei Wu, Yuan-Fang Wang, and William Yang Wang · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever · 2018
Earlier work this paper cites.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut · 2018
Earlier work this paper cites.
Allennlp: A deep semantic natural language processing platform
Matt Gardner, Joel Grus, Mark Neumann, Oyvind Tafjord, Pradeep Dasigi, Nelson F Liu, Matthew E Peters, Michael Schmitz, and Luke Zettlemoyer · 2018
Earlier work this paper cites.
How2: A large-scale dataset for multimodal language understanding
Ramon Sanabria, Ozan Caglayan, Shruti Palaskar, Desmond Elliott, Loïc Barrault, Lucia Specia, and Florian Metze · 2018
Earlier work this paper cites.
Lrs3-ted: a large-scale dataset for visual speech recognition
Triantafyllos Afouras, Joon Son Chung, and Andrew Zisserman · 2018
Earlier work this paper cites.
The sound of pixels
Hang Zhao, Chuang Gan, Andrew Rouditchenko, Carl Vondrick, Josh McDermott, and Antonio Torralba · 2018
Earlier work this paper cites.
On the role of text preprocessing in neural network architectures: An evaluation study on text categorization and sentiment analysis
Jose Camacho-Collados and Mohammad Taher Pilehvar · 2018
Earlier work this paper cites.
Crisismmd: Multimodal twitter datasets from natural disasters
Firoj Alam, Ferda Ofli, and Muhammad Imran · 2018
Earlier work this paper cites.
Eann: Event adversarial neural networks for multi-modal fake news detection
Yaqing Wang, Fenglong Ma, Zhiwei Jin, Ye Yuan, Guangxu Xun, Kishlay Jha, Lu Su, and Jing Gao · 2018
Earlier work this paper cites.
Don’t just assume; look and answer: Overcoming priors for visual question answering
Aishwarya Agrawal, Dhruv Batra, Devi Parikh, and Aniruddha Kembhavi · 2018
Earlier work this paper cites.
Attngan: Fine-grained text to image generation with attentional generative adversarial networks
Tao Xu, Pengchuan Zhang, Qiuyuan Huang, Han Zhang, Zhe Gan, Xiaolei Huang, and Xiaodong He · 2018
Earlier work this paper cites.
Deep audio-visual speech recognition
Triantafyllos Afouras, Joon Son Chung, Andrew Senior, Oriol Vinyals, and Andrew Zisserman · 2018
Cited alongside, same era.
Deep voice 3: Scaling text-to-speech with convolutional sequence learning
Wei Ping, Kainan Peng, Andrew Gibiansky, Sercan O Arik, Ajay Kannan, Sharan Narang, Jonathan Raiman, and John Miller · 2018
Cited alongside, same era.
Natural tts synthesis by conditioning wavenet on mel spectrogram predictions
Jonathan Shen, Ruoming Pang, Ron J Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Yang, Zhifeng Chen, Yu Zhang, Yuxuan Wang, Rj Skerrv-Ryan, Rif A. Saurous, Yannis Agiomyrgiannakis, and Yonghui Wu · 2018
Cited alongside, same era.
Lip2audspec: Speech reconstruction from silent lip movements video
Hassan Akbari, Himani Arora, Liangliang Cao, and Nima Mesgarani · 2018
Cited alongside, same era.
Multimodal language analysis in the wild: Cmu-mosei dataset and interpretable dynamic fusion graph
Amir Zadeh and Paul Pu · 2018
Cited alongside, same era.
Multi-gate attention network for image captioning
Weitao Jiang, Xiying Li, Haifeng Hu, Qiang Lu, and Bohong Liu · 2021
Later among the works it cites.
Gaussian process with graph convolutional kernel for relational learning
Jinyuan Fang, Shangsong Liang, Zaiqiao Meng, and Qiang Zhang · 2021
Later among the works it cites.
Vinvl: Revisiting visual representations in vision-language models
Pengchuan Zhang, Xiujun Li, Xiaowei Hu, Jianwei Yang, Lei Zhang, Lijuan Wang, Yejin Choi, and Jianfeng Gao · 2021
Later among the works it cites.
M6: Multi-modality-to-multi-modality multitask mega-transformer for unified pretraining
Junyang Lin, Rui Men, An Yang, Chang Zhou, Yichang Zhang, Peng Wang, Jingren Zhou, Jie Tang, and Hongxia Yang · 2021
Later among the works it cites.
Hatebert: Retraining bert for abusive language detection in english
Tommaso Caselli, Valerio Basile, Jelena Mitrović, and Michael Granitzer · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deep multimodal representation learning: A survey
Wenzhong Guo, Jianwen Wang, and Shiping Wang · 2019
Cited alongside, same era.
Visualbert: A simple and performant baseline for vision and language
Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang · 2019
Cited alongside, same era.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee · 2019
Cited alongside, same era.
Language-conditioned graph networks for relational reasoning
Ronghang Hu, Anna Rohrbach, Trevor Darrell, and Kate Saenko · 2019
Cited alongside, same era.
Dynamic fusion with intra-and inter-modality attention flow for visual question answering
Peng Gao, Zhengkai Jiang, Haoxuan You, Pan Lu, Steven CH Hoi, Xiaogang Wang, and Hongsheng Li · 2019
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin Ming-Wei Chang Kenton and Lee Kristina Toutanova · 2019
Cited alongside, same era.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Cited alongside, same era.
Zewen Chi, Li Dong, Furu Wei, Nan Yang, Saksham Singhal, Wenhui Wang, Xia Song, Xian-Ling Mao, He-Yan Huang, and Ming Zhou · 2021
Later among the works it cites.
Shuffled-token detection for refining pre-trained roberta
Subhadarshi Panda, Anjali Agrawal, Jeewon Ha, and Benjamin Bloch · 2021
Later among the works it cites.
Product1m: Towards weakly supervised instance-level product retrieval via cross-modal pretraining
Xunlin Zhan, Yangxin Wu, Xiao Dong, Yunchao Wei, Minlong Lu, Yichi Zhang, Hang Xu, and Xiaodan Liang · 2021
Later among the works it cites.
Uc2: Universal cross-lingual cross-modal vision-and-language pre-training
Mingyang Zhou, Luowei Zhou, Shuohang Wang, Yu Cheng, Linjie Li, Zhou Yu, and Jingjing Liu · 2021
Later among the works it cites.
Deformable detr: Deformable transformers for end-to-end object detection
Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, and Jifeng Dai · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever · 2021
Later among the works it cites.
Virtex: Learning visual representations from textual annotations
Karan Desai and Justin Johnson · 2021
Later among the works it cites.
Ernie-vil: Knowledge enhanced vision-language representations through scene graphs
Fei Yu, Jiji Tang, Weichong Yin, Yu Sun, Hao Tian, Hua Wu, and Haifeng Wang · 2021
Later among the works it cites.
Hubert: How much can a bad teacher benefit asr pre-training?
Wei-Ning Hsu, Yao-Hung Hubert Tsai, Benjamin Bolte, Ruslan Salakhutdinov, and Abdelrahman Mohamed · 2021
Later among the works it cites.
Towards a unified foundation model: Jointly pre-training transformers on unpaired images and text
Qing Li, Boqing Gong, Yin Cui, Dan Kondratyuk, Xianzhi Du, Ming-Hsuan Yang, and Matthew Brown · 2021
Later among the works it cites.
Unit: Multimodal multitask learning with a unified transformer
Ronghang Hu and Amanpreet Singh · 2021
Later among the works it cites.
Vatt: Transformers for multimodal self-supervised learning from raw video, audio and text
Hassan Akbari, Liangzhe Yuan, Rui Qian, Wei-Hong Chuang, Shih-Fu Chang, Yin Cui, and Boqing Gong · 2021
Later among the works it cites.
Self-supervised multimodal opinion summarization
Jinbae Im, Moonki Kim, Hoyeop Lee, Hyunsouk Cho, and Sehee Chung · 2021
Later among the works it cites.
Layoutlmv2: Multi-modal pre-training for visually-rich document understanding
Yang Xu, Yiheng Xu, Tengchao Lv, Lei Cui, Furu Wei, Guoxin Wang, Yijuan Lu, Dinei Florencio, Cha Zhang, Wanxiang Che, Min Zhang, and Lidong Zhou · 2021
Later among the works it cites.
Structext: Structured text understanding with multi-modal transformers
Yulin Li, Yuxi Qian, Yuechen Yu, Xiameng Qin, Chengquan Zhang, Yan Liu, Kun Yao, Junyu Han, Jingtuo Liu, and Errui Ding · 2021
Later among the works it cites.
Do we really need explicit position encodings for vision transformers
Xiangxiang Chu, Bo Zhang, Zhi Tian, Xiaolin Wei, and Huaxia Xia · 2021
Later among the works it cites.
Vision guided generative pre-trained language models for multimodal abstractive summarization
Tiezheng Yu, Wenliang Dai, Zihan Liu, and Pascale Ngan Fung · 2021
Later among the works it cites.
Leveraging category information for single-frame visual sound source separation
Lingyu Zhu and Esa Rahtu · 2021
Later among the works it cites.
Market strategies used by processed food manufacturers to increase and consolidate their power: a systematic review and document analysis
Benjamin Wood, Owain Williams, Vijaya Nagarajan, and Gary Sacks · 2021
Later among the works it cites.
Multi-modal generative adversarial networks for traffic event detection in smart cities
Qi Chen, Wei Wang, Kaizhu Huang, Suparna De, and Frans Coenen · 2021
Later among the works it cites.
Kvl-bert: Knowledge enhanced visual-and-linguistic bert for visual commonsense reasoning
Dandan Song, Siyi Ma, Zhanchen Sun, Sicheng Yang, and Lejian Liao · 2021
Later among the works it cites.
How to find a good image-text embedding for remote sensing visual question answering?
Christel Chappuis, Sylvain Lobry, Benjamin Alexander Kellenberger, Bertrand Le Saux, and Devis Tuia · 2021
Later among the works it cites.
An improved attention for visual question answering
Tanzila Rahman, Shih-Han Chou, Leonid Sigal, and Giuseppe Carenini · 2021
Later among the works it cites.
Learning robust patient representations from multi-modal electronic health records: a supervised deep learning approach
Xianli Zhang, Buyue Qian, Yang Li, Yang Liu, Xi Chen, Chong Guan, and Chen Li · 2021
Later among the works it cites.
Predicting the survival of cancer patients with multimodal graph neural network
Jianliang Gao, Tengfei Lyu, Fan Xiong, Jianxin Wang, Weimao Ke, and Zhao Li · 2021
Later among the works it cites.
Multi-modal neural machine translation with deep semantic interactions
Jinsong Su, Jinchang Chen, Hui Jiang, Chulun Zhou, Huan Lin, Yubin Ge, Qingqiang Wu, and Yongxuan Lai · 2021
Later among the works it cites.
M3p: Learning universal representations via multitask multilingual multimodal pre-training
Minheng Ni, Haoyang Huang, Lin Su, Edward Cui, Taroon Bharti, Lijuan Wang, Dongdong Zhang, and Nan Duan · 2021
Later among the works it cites.
Mural: multimodal, multitask retrieval across languages
Aashi Jain, Mandy Guo, Krishna Srinivasan, Ting Chen, Sneha Kudugunta, Chao Jia, Yinfei Yang, and Jason Baldridge · 2021
Later among the works it cites.
Align before fuse: Vision and language representation learning with momentum distillation
Junnan Li, Ramprasaath Selvaraju, Akhilesh Gotmare, Shafiq Joty, Caiming Xiong, and Steven Chu Hong Hoi · 2021
Later among the works it cites.
Multimodal few-shot learning with frozen language models
Maria Tsimpoukelli, Jacob L Menick, Serkan Cabi, SM Eslami, Oriol Vinyals, and Felix Hill · 2021
Later among the works it cites.
Kaleido-bert: Vision-language pre-training on fashion domain
Mingchen Zhuge, Dehong Gao, Deng-Ping Fan, Linbo Jin, Ben Chen, Haoming Zhou, Minghui Qiu, and Ling Shao · 2021
Later among the works it cites.
Compressing visual-linguistic model via knowledge distillation
Zhiyuan Fang, Jianfeng Wang, Xiaowei Hu, Lijuan Wang, Yezhou Yang, and Zicheng Liu · 2021
Later among the works it cites.
Scaling up visual and vision-language representation learning with noisy text supervision
Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, Quoc Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig · 2021
Later among the works it cites.
An overview of deep-learning-based audio-visual speech enhancement and separation
Daniel Michelsanti, Zheng-Hua Tan, Shi-Xiong Zhang, Yong Xu, Meng Yu, Dong Yu, and Jesper Jensen · 2021
Later among the works it cites.
Learning audio-visual speech representation by masked multimodal cluster prediction
Bowen Shi, Wei-Ning Hsu, Kushal Lakhotia, and Abdelrahman Mohamed · 2022
Later among the works it cites.
Revisiting parameter-efficient tuning: Are we really there yet?
Guanzheng Chen, Fangyu Liu, Zaiqiao Meng, and Shangsong Liang · 2022
Later among the works it cites.
Vision-language pre-training: Basics, recent advances, and future trends
Zhe Gan, Linjie Li, Chunyuan Li, Lijuan Wang, Zicheng Liu, and Jianfeng Gao · 2022
Later among the works it cites.
Vlmo: Unified vision-language pre-training with mixture-of-modality-experts
Hangbo Bao, Wenhui Wang, Li Dong, Qiang Liu, Owais Khan Mohammed, Kriti Aggarwal, Subhojit Som, Songhao Piao, and Furu Wei · 2022
Later among the works it cites.
Multimodal co-learning: challenges, applications with datasets, recent advances and future directions
Anil Rahate, Rahee Walambe, Sheela Ramanna, and Ketan Kotecha · 2022
Later among the works it cites.
A survey of vision-language pre-trained models
Yifan Du, Zikang Liu, Junyi Li, and Wayne Xin Zhao · 2022
Later among the works it cites.
Vision-and-language pretrained models: A survey
Siqu Long, Feiqi Cao, Soyeon Caren Han, and Haiqin Yang · 2022
Later among the works it cites.
Ammu: a survey of transformer-based biomedical pretrained language models
Katikapalli Subramanyam Kalyan, Ajit Rajasekharan, and Sivanesan Sangeetha · 2022
Later among the works it cites.
A survey of data representation for multi-modality event detection and evolution
Kejing Xiao, Zhaopeng Qian, and Biao Qin · 2022
Later among the works it cites.
Multi-modal knowledge graph construction and application: A survey
Xiangru Zhu, Zhixu Li, Xiaodan Wang, Xueyao Jiang, Penglei Sun, Xuwu Wang, Yanghua Xiao, and Nicholas Jing Yuan · 2022
Later among the works it cites.
Vqa-gnn: Reasoning with multimodal semantic graph for visual question answering
Yanan Wang, Michihiro Yasunaga, Hongyu Ren, Shinya Wada, and Jure Leskovec · 2022
Later among the works it cites.
Multi-relational graph representation learning with bayesian gaussian process network
Guanzheng Chen, Jinyuan Fang, Zaiqiao Meng, Qiang Zhang, and Shangsong Liang · 2022
Later among the works it cites.
Knowledge inheritance for pre-trained language models
Yujia Qin, Yankai Lin, Jing Yi, Jiajie Zhang, Xu Han, Zhengyan Zhang, Yusheng Su, Zhiyuan Liu, Peng Li, Maosong Sun, and Jie Zhou · 2022
Later among the works it cites.
Kd-vlp: Improving end-to-end vision-and-language pretraining with object knowledge distillation
Yongfei Liu, Chenfei Wu, Shao-Yen Tseng, Vasudev Lal, Xuming He, and Nan Duan · 2022
Later among the works it cites.
BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi · 2022
Later among the works it cites.
Ofa: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework
Peng Wang, An Yang, Rui Men, Junyang Lin, Shuai Bai, Zhikang Li, Jianxin Ma, Chang Zhou, Jingren Zhou, and Hongxia Yang · 2022
Later among the works it cites.
Scaling instruction-finetuned language models
Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, Albert Webson, Shixiang Shane Gu, Zhuyun Dai, Mirac Suzgun, Xinyun Chen, Aakanksha Chowdhery, Dasha Valter, Sharan Narang, Gaurav Mishra, Adams Wei Yu, Vincent Zhao, Yanping Huang, Andrew M. Dai, Hongkun Yu, Slav Petrov, Ed Huai-hsin Chi, Jeff Dean, Jacob Devlin, Adam Roberts, Denny Zhou, Quoc V. Le, and Jason Wei · 2022
Later among the works it cites.
Xylayoutlm: Towards layout-aware multimodal networks for visually-rich document understanding
Zhangxuan Gu, Changhua Meng, Ke Wang, Jun Lan, Weiqiang Wang, Ming Gu, and Liqing Zhang · 2022
Later among the works it cites.
Multi-source multimodal data and deep learning for disaster response: A systematic review
Nilani Algiriyage, Raj Prasanna, Kristin Stock, Emma EH Doyle, and David Johnston · 2022
Later among the works it cites.
Fmfn: Fine-grained multimodal fusion networks for fake news detection
Jingzi Wang, Hongyan Mao, and Hongwei Li · 2022
Later among the works it cites.
MS 2 \mathrm{MS}^{2} -gnn: Exploring gnn-based multimodal fusion network for depression detection
Tao Chen, Richang Hong, Yanrong Guo, Shijie Hao, and Bin Hu · 2022
Later among the works it cites.
Data2vec: A general framework for self-supervised learning in speech, vision and language
Alexei Baevski, Wei-Ning Hsu, Qiantong Xu, Arun Babu, Jiatao Gu, and Michael Auli · 2022
Later among the works it cites.
Flava: A foundational language and vision alignment model
Amanpreet Singh, Ronghang Hu, Vedanuj Goswami, Guillaume Couairon, Wojciech Galuba, Marcus Rohrbach, and Douwe Kiela · 2022
Later among the works it cites.
Cross-view language modeling: Towards unified cross-lingual cross-modal pre-training
Yan Zeng, Wangchunshu Zhou, Ao Luo, and Xinsong Zhang · 2022
Later among the works it cites.
Pali: A jointly-scaled multilingual language-image model
Xi Chen, Xiao Wang, Soravit Changpinyo, AJ Piergiovanni, Piotr Padlewski, Daniel Salz, Sebastian Goodman, Adam Grycner, Basil Mustafa, Lucas Beyer, Alexander Kolesnikov, Joan Puigcerver, Nan Ding, Keran Rong, Hassan Akbari, Gaurav Mishra, Linting Xue, Ashish V. Thapliyal, James Bradbury, Weicheng Kuo, Mojtaba Seyedhosseini, Chao Jia, Burcu Karagol Ayan, Carlos Riquelme, Andreas Steiner, Anelia Angelova, Xiaohua Zhai, Neil Houlsby, and Radu Soricut · 2022
Later among the works it cites.
Towards artificial general intelligence via a multimodal foundation model
Nanyi Fei, Zhiwu Lu, Yizhao Gao, Guoxing Yang, Yuqi Huo, Jingyuan Wen, Haoyu Lu, Ruihua Song, Xin Gao, Tao Xiang, Haoran Sun, and Jiling Wen · 2022
Later among the works it cites.
Vlp: A survey on vision-language pre-training
Fei-Long Chen, Du-Zhen Zhang, Ming-Lun Han, Xiu-Yi Chen, Jing Shi, Shuang Xu, and Bo Xu · 2023
Closest in time.
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi · 2023
Closest in time.
Instructblip: Towards general-purpose vision-language models with instruction tuning
Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale Fung, and Steven Hoi · 2023
Closest in time.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, March 2023
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing · 2023
Closest in time.
Multimodal learning with graphs
Yasha Ektefaie, George Dasoulas, Ayush Noori, Maha Farhat, and Marinka Zitnik · 2023
Closest in time.
Vlp: A survey on vision-language pre-training
Fei-Long Chen, Du-Zhen Zhang, Ming-Lun Han, Xiu-Yi Chen, Jing Shi, Shuang Xu, and Bo Xu · 2023
Closest in time.