Early vs late fusion in multimodal convolutional neural networks. In
Konrad Gadzicki, Razieh Khamsehashari, and Christoph Zetzsche. 2020 · 2020
Later among the works it cites.
RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models. In
Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A Smith. 2020 · 2020
Later among the works it cites.
Shortcut learning in deep neural networks
Robert Geirhos, Jörn-Henrik Jacobsen, Claudio Michaelis, Richard Zemel, Wieland Brendel, Matthias Bethge, and Felix A Wichmann. 2020 · 2020
Later among the works it cites.
Vqa-lol: Visual question answering under the lens of logic. In
Tejas Gokhale, Pratyay Banerjee, Chitta Baral, and Yezhou Yang. 2020 · 2020
Later among the works it cites.
Does my multimodal model learn cross-modal interactions? It’s harder to tell than you might think!. In
Jack Hessel and Lillian Lee. 2020 · 2020
Later among the works it cites.
Evidential Sparsification of Multimodal Latent Spaces in Conditional Variational Autoencoders
Original
Masha Itkina, B. Ivanovic, Ransalu Senanayake, Mykel J. Kochenderfer, and Marco Pavone. 2020 · 2020
Later among the works it cites.
Text-Image-Video Summary Generation Using Joint Integer Linear Programming. In
Anubhav Jangra, Adam Jatowt, Mohammad Hasanuzzaman, and Sriparna Saha. 2020 · 2020
Later among the works it cites.
Multiplicative Interactions and Where to Find Them. In
Siddhant M. Jayakumar, Wojciech M. Czarnecki, Jacob Menick, Jonathan Schwarz, Jack Rae, Simon Osindero, Yee Whye Teh, Tim Harley, and Razvan Pascanu. 2020 · 2020
Later among the works it cites.
Analyzing and improving the image quality of stylegan. In
Tero Karras, Samuli Laine, Miika Aittala, Janne Hellsten, Jaakko Lehtinen, and Timo Aila. 2020 · 2020
Later among the works it cites.
Concept bottleneck models. In
Pang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann, Emma Pierson, Been Kim, and Percy Liang. 2020 · 2020
Later among the works it cites.
VMSMO: Learning to Generate Multimodal Summary for Video-based News Articles
Original
Mingzhe Li, Xiuying Chen, Shen Gao, Zhangming Chan, Dongyan Zhao, and Rui Yan. 2020 · 2020
Later among the works it cites.
The explanation game: Explaining machine learning models using shapley values. In
Luke Merrick and Ankur Taly. 2020 · 2020
Later among the works it cites.
Nerf: Representing scenes as neural radiance fields for view synthesis. In
Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. 2020 · 2020
Later among the works it cites.
Multi-modal domain adaptation for fine-grained action recognition. In
Jonathan Munro and Dima Damen. 2020 · 2020
Later among the works it cites.
Characterization and classification of semantic image-text relations
Christian Otto, Matthias Springstein, Avishek Anand, and Ralph Ewerth. 2020 · 2020
Later among the works it cites.
Decoding brain representations by multimodal learning of neural activity and visual features
Simone Palazzo, Concetto Spampinato, Isaak Kavasidis, Daniela Giordano, Joseph Schmidt, and Mubarak Shah. 2020 · 2020
Later among the works it cites.
Visualcomet: Reasoning about the dynamic context of a still image. In
Jae Sung Park, Chandra Bhagavatula, Roozbeh Mottaghi, Ali Farhadi, and Yejin Choi. 2020 · 2020
Later among the works it cites.
FairCVtest Demo: Understanding Bias in Multimodal Learning with a Testbed in Fair Automatic Recruitment. In
Alejandro Peña, Ignacio Serna, Aythami Morales, and Julian Fierrez. 2020 · 2020
Later among the works it cites.
Integrating Multimodal Information in Large Pretrained Transformers. In
Wasifur Rahman, Md Kamrul Hasan, Sangwu Lee, AmirAli Bagher Zadeh, Chengfeng Mao, Louis-Philippe Morency, and Ehsan Hoque. 2020 · 2020
Later among the works it cites.
Measuring Social Biases in Grounded Vision and Language Embeddings
Original
Candace Ross, Boris Katz, and Andrei Barbu. 2020 · 2020
Later among the works it cites.
Winogrande: An adversarial winograd schema challenge at scale. In
Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2020 · 2020
Later among the works it cites.
Multimodal graph networks for compositional generalization in visual question answering
Raeid Saqur and Karthik Narasimhan. 2020 · 2020
Later among the works it cites.
Mgat: Multimodal graph attention network for recommendation
Zhulin Tao, Yinwei Wei, Xiang Wang, Xiangnan He, Xianglin Huang, and Tat-Seng Chua. 2020 · 2020
Later among the works it cites.
What makes for good views for contrastive learning?
Yonglong Tian, Chen Sun, Ben Poole, Dilip Krishnan, Cordelia Schmid, and Phillip Isola. 2020b · 2020
Later among the works it cites.
NBDT: Neural-Backed Decision Tree. In
Alvin Wan, Lisa Dunlap, Daniel Ho, Jihan Yin, Scott Lee, Suzanne Petryk, Sarah Adel Bargal, and Joseph E Gonzalez. 2020 · 2020
Later among the works it cites.
Multimodal end-to-end autonomous driving
Yi Xiao, Felipe Codevilla, Akhil Gurram, Onay Urfalioglu, and Antonio M López. 2020 · 2020
Later among the works it cites.
A Novel Graph-based Multi-modal Fusion Encoder for Neural Machine Translation. In
Yongjing Yin, Fandong Meng, Jinsong Su, Chulun Zhou, Zhengyuan Yang, Jie Zhou, and Jiebo Luo. 2020 · 2020
Later among the works it cites.
Foundations of multimodal co-learning
Amir Zadeh, Paul Pu Liang, and Louis-Philippe Morency. 2020 · 2020
Later among the works it cites.
Multimodal feature fusion by relational reasoning and attention for visual question answering
Weifeng Zhang, Jing Yu, Hua Hu, Haiyang Hu, and Zengchang Qin. 2020 · 2020
Later among the works it cites.
Persistent anti-muslim bias in large language models. In
Abubakar Abid, Maheen Farooqi, and James Zou. 2021 · 2021
Later among the works it cites.
Vatt: Transformers for multimodal self-supervised learning from raw video, audio and text
Original
Hassan Akbari, Liangzhe Yuan, Rui Qian, et al · 2021
Later among the works it cites.
On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?. In
Emily M Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021 · 2021
Later among the works it cites.
Multimodal datasets: misogyny, pornography, and malignant stereotypes
Original
Abeba Birhane, Vinay Uday Prabhu, and Emmanuel Kahembwe. 2021 · 2021
Later among the works it cites.
Geometric deep learning: Grids, groups, graphs, geodesics, and gauges
Original
Michael M Bronstein, Joan Bruna, Taco Cohen, and Petar Veličković. 2021 · 2021
Later among the works it cites.
Extracting training data from large language models. In
Nicholas Carlini, Florian Tramer, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, et al · 2021
Later among the works it cites.
Grounding ‘Grounding’in NLP. In
Khyathi Raghavi Chandu, Yonatan Bisk, and Alan W Black. 2021 · 2021
Later among the works it cites.
VisualGPT: Data-efficient adaptation of pretrained language models for image captioning
Original
Jun Chen, Han Guo, Kai Yi, Boyang Li, and Mohamed Elhoseiny. 2021a · 2021
Later among the works it cites.
Compression, Transduction, and Creation: A Unified Framework for Evaluating Natural Language Generation. In
Mingkai Deng, Bowen Tan, Zhengzhong Liu, Eric Xing, and Zhiting Hu. 2021 · 2021
Later among the works it cites.
Multimodal safety-critical scenarios generation for decision-making algorithms evaluation
Wenhao Ding, Baiming Chen, Bo Li, Kim Ji Eun, and Ding Zhao. 2021 · 2021
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, et al · 2021
Later among the works it cites.
Towards Reliable Multimodal Stress Detection under Distribution Shift. In
Andreas Foltyn and Jessica Deuschel. 2021 · 2021
Later among the works it cites.
Perceptual Score: What Data Modalities Does Your Model Perceive?
Itai Gat, Idan Schwartz, and Alex Schwing. 2021 · 2021
Later among the works it cites.
KAT: A Knowledge Augmented Transformer for Vision-and-Language
Original
Liangke Gui, Borui Wang, Qiuyuan Huang, Alex Hauptmann, Yonatan Bisk, and Jianfeng Gao. 2021 · 2021
Later among the works it cites.
The social impact of deepfakes
Jeffrey T Hancock and Jeremy N Bailenson. 2021 · 2021
Later among the works it cites.
Learning by aligning videos in time. In
Sanjay Haresh, Sateesh Kumar, Huseyin Coskun, Shahram N Syed, Andrey Konin, Zeeshan Zia, and Quoc-Huy Tran. 2021 · 2021
Later among the works it cites.
Humor knowledge enriched transformer for understanding multimodal humor. In
Md Kamrul Hasan, Sangwu Lee, Wasifur Rahman, Amir Zadeh, Rada Mihalcea, Louis-Philippe Morency, and Ehsan Hoque. 2021 · 2021
Later among the works it cites.
Decoupling the role of data, attention, and losses in multimodal transformers
Original
Lisa Anne Hendricks, John Mellor, Rosalia Schneider, Jean-Baptiste Alayrac, and Aida Nematzadeh. 2021 · 2021
Later among the works it cites.
Perceiver: General perception with iterative attention
Original
Andrew Jaegle, Felix Gimeno, Andrew Brock, Andrew Zisserman, Oriol Vinyals, and Joao Carreira. 2021 · 2021
Later among the works it cites.
Scaling up visual and vision-language representation learning with noisy text supervision. In
Chao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen, Zarana Parekh, Hieu Pham, et al · 2021
Later among the works it cites.
Text-to-image generation grounded by fine-grained user attention. In
Jing Yu Koh, Jason Baldridge, Honglak Lee, and Yinfei Yang. 2021 · 2021
Later among the works it cites.
The Power of Scale for Parameter-Efficient Prompt Tuning. In
Brian Lester, Rami Al-Rfou, and Noah Constant. 2021 · 2021
Later among the works it cites.
Towards a Unified Foundation Model: Jointly Pre-Training Transformers on Unpaired Images and Text
Original
Qing Li, Boqing Gong, Yin Cui, Dan Kondratyuk, Xianzhi Du, et al · 2021
Later among the works it cites.
Prefix-Tuning: Optimizing Continuous Prompts for Generation. In
Xiang Lisa Li and Percy Liang. 2021 · 2021
Later among the works it cites.
Temporal fusion transformers for interpretable multi-horizon time series forecasting
Bryan Lim, Sercan Ö Arık, Nicolas Loeff, and Tomas Pfister. 2021 · 2021
Later among the works it cites.
Pretrained transformers as universal computation engines
Original
Kevin Lu, Aditya Grover, Pieter Abbeel, and Igor Mordatch. 2021 · 2021
Later among the works it cites.
Smil: Multimodal learning with severely missing modality
Original
Mengmeng Ma, Jian Ren, Long Zhao, Sergey Tulyakov, Cathy Wu, and Xi Peng. 2021 · 2021
Later among the works it cites.
Semi-Supervised Aggregation of Dependent Weak Supervision Sources With Performance Guarantees. In
Alessio Mazzetto, Dylan Sam, Andrew Park, Eli Upfal, and Stephen Bach. 2021 · 2021
Later among the works it cites.
A comprehensive survey on multimodal medical signals fusion for smart healthcare systems
Ghulam Muhammad, Fatima Alshehri, Fakhri Karray, Abdulmotaleb El Saddik, et al · 2021
Later among the works it cites.
Counterfactual vqa: A cause-effect look at language bias. In
Yulei Niu, Kaihua Tang, Hanwang Zhang, Zhiwu Lu, Xian-Sheng Hua, and Ji-Rong Wen. 2021 · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision. In
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, et al · 2021
Later among the works it cites.
Zero-shot text-to-image generation. In
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. 2021 · 2021
Later among the works it cites.
FLAVA: A Foundational Language And Vision Alignment Model
Original
Amanpreet Singh, Ronghang Hu, Vedanuj Goswami, Guillaume Couairon, et al · 2021
Later among the works it cites.
Kvl-bert: Knowledge enhanced visual-and-linguistic bert for visual commonsense reasoning
Dandan Song, Siyi Ma, Zhanchen Sun, Sicheng Yang, and Lejian Liao. 2021 · 2021
Later among the works it cites.
Worst of Both Worlds: Biases Compound in Pre-trained Vision-and-Language Models
Original
Tejas Srinivasan and Yonatan Bisk. 2021 · 2021
Later among the works it cites.
Contrastive learning, multi-view redundancy, and linear models. In
Christopher Tosh, Akshay Krishnamurthy, and Daniel Hsu. 2021 · 2021
Later among the works it cites.
Multimodal few-shot learning with frozen language models
Maria Tsimpoukelli, Jacob Menick, Serkan Cabi, SM Eslami, Oriol Vinyals, and Felix Hill. 2021 · 2021
Later among the works it cites.
M2Lens: Visualizing and explaining multimodal models for sentiment analysis
Xingbo Wang, Jianben He, Zhihua Jin, et al · 2021
Later among the works it cites.
Leveraging sparse linear layers for debuggable deep networks. In
Eric Wong, Shibani Santurkar, and Aleksander Madry. 2021 · 2021
Later among the works it cites.
MUFASA: Multimodal Fusion Architecture Search for Electronic Health Records
Original
Zhen Xu, David R So, and Andrew M Dai. 2021 · 2021
Later among the works it cites.
MTAG: Modal-Temporal Attention Graph for Unaligned Human Multimodal Language Sequences. In
Jianing Yang, Yongxin Wang, Ruitao Yi, Yuying Zhu, Azaan Rehman, Amir Zadeh, et al · 2021
Later among the works it cites.
Vision Guided Generative Pre-trained Language Models for Multimodal Abstractive Summarization. In
Tiezheng Yu, Wenliang Dai, Zihan Liu, and Pascale Fung. 2021 · 2021
Later among the works it cites.
Merlot: Multimodal neural script knowledge models
Rowan Zellers, Ximing Lu, Jack Hessel, Youngjae Yu, Jae Sung Park, Jize Cao, Ali Farhadi, and Yejin Choi. 2021 · 2021
Later among the works it cites.
Arbitrary talking face generation via attentional audio-visual coherence learning. In
Hao Zhu, Huaibo Huang, Yi Li, Aihua Zheng, and Ran He. 2021 · 2021
Later among the works it cites.
VL-InterpreT: An Interactive Visualization Tool for Interpreting Vision-Language Transformers. In
Estelle Aflalo, Meng Du, Shao-Yen Tseng, Yongfei Liu, Chenfei Wu, Nan Duan, and Vasudev Lal. 2022 · 2022
Closest in time.
Flamingo: a visual language model for few-shot learning
Original
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, et al · 2022
Closest in time.
A theory of consciousness from a theoretical computer science perspective: Insights from the Conscious Turing Machine
Lenore Blum and Manuel Blum. 2022 · 2022
Closest in time.
Dall-eval: Probing the reasoning skills and social biases of text-to-image generative transformers
Original
Jaemin Cho, Abhay Zala, and Mohit Bansal. 2022 · 2022
Closest in time.
Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied Agents
Original
Wenlong Huang, Pieter Abbeel, Deepak Pathak, and Igor Mordatch. 2022 · 2022
Closest in time.
Pre-Trained Language Models for Interactive Decision-Making
Original
Shuang Li, Xavier Puig, Yilun Du, Clinton Wang, Ekin Akyurek, Antonio Torralba, Jacob Andreas, and Igor Mordatch. 2022a · 2022
Closest in time.
Brainish: Formalizing A Multimodal Language for Intelligence and Consciousness
Original
Paul Pu Liang. 2022 · 2022
Closest in time.
HighMMT: Towards Modality and Task Generalization for High-Modality Representation Learning
Original
Paul Pu Liang, Yiwei Lyu, Xiang Fan, Shengtong Mo, Dani Yogatama, et al · 2022
Closest in time.
PolyViT: Co-training Vision Transformers on Images, Videos and Audio
Valerii Likhosherstov, Mostafa Dehghani, Anurag Arnab, Krzysztof Marcin Choromanski, Mario Lucic, Yi Tay, and Adrian Weller. 2022 · 2022
Closest in time.
Cross-Modal Discrete Representation Learning. In
Alex Liu, SouYoung Jin, Cheng-I Lai, Andrew Rouditchenko, Aude Oliva, and James Glass. 2022 · 2022
Closest in time.
DIME: Fine-grained Interpretations of Multimodal Models via Disentangled Local Explanations
Original
Yiwei Lyu, Paul Pu Liang, Zihao Deng, Ruslan Salakhutdinov, and Louis-Philippe Morency. 2022 · 2022
Closest in time.
Spoken language interaction with robots: Recommendations for future research
Matthew Marge, Carol Espy-Wilson, Nigel G Ward, Abeer Alwan, Yoav Artzi, Mohit Bansal, et al · 2022
Closest in time.
Learning to Listen: Modeling Non-Deterministic Dyadic Facial Motion. In
Evonne Ng, Hanbyul Joo, Liwen Hu, Hao Li, Trevor Darrell, Angjoo Kanazawa, and Shiry Ginosar. 2022 · 2022
Closest in time.
Multimodal Learning using Optimal Transport for Sarcasm and Humor Detection. In
Shraman Pramanick, Aniket Roy, and Vishal M Patel. 2022 · 2022
Closest in time.
One model to learn them all
Scott Reed, Konrad Zolna, Emilio Parisotto, Sergio Gómez Colmenarejo, Alexander Novikov, Gabriel Barth-Maron, Mai Giménez, Yury Sulsky, et al · 2022
Closest in time.
High-resolution image synthesis with latent diffusion models. In
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. 2022 · 2022
Closest in time.
Make-a-video: Text-to-video generation without text-video data
Original
Uriel Singer, Adam Polyak, Thomas Hayes, Xi Yin, Jie An, Songyang Zhang, Qiyuan Hu, et al · 2022
Closest in time.
Winoground: Probing Vision and Language Models for Visio-Linguistic Compositionality. In
Tristan Thrush, Ryan Jiang, Max Bartolo, Amanpreet Singh, Adina Williams, Douwe Kiela, and Candace Ross. 2022 · 2022
Closest in time.
Multimodal research in vision and language: A review of current and emerging trends
Shagun Uppal, Sarthak Bhagat, Devamanyu Hazarika, Navonil Majumder, et al · 2022
Closest in time.
Analyzing differentiable fuzzy logic operators
Emile van Krieken, Erman Acar, and Frank van Harmelen. 2022 · 2022
Closest in time.
Face-to-Face Contrastive Learning for Social Intelligence Question-Answering
Original
Alex Wilf, Qianli M Ma, Paul Pu Liang, Amir Zadeh, and Louis-Philippe Morency. 2022 · 2022
Closest in time.
Characterizing and overcoming the greedy nature of learning in multi-modal deep neural networks
Original
Nan Wu, Stanisław Jastrzębski, Kyunghyun Cho, and Krzysztof J Geras. 2022 · 2022
Closest in time.
Socratic models: Composing zero-shot multimodal reasoning with language
Original
Andy Zeng, Adrian Wong, Stefan Welker, Krzysztof Choromanski, Federico Tombari, et al · 2022
Closest in time.
Multi-Modal Knowledge Graph Construction and Application: A Survey
Original
Xiangru Zhu, Zhixu Li, Xiaodan Wang, Xueyao Jiang, Penglei Sun, Xuwu Wang, et al · 2022
Closest in time.
MultiViz: Towards Visualizing and Understanding Multimodal Models. In
Paul Pu Liang, Yiwei Lyu, Gunjan Chhablani, Nihal Jain, Zihao Deng, et al · 2023
Closest in time.
Show, attend and tell: Neural image caption generation with visual attention. In
Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhudinov, Rich Zemel, and Yoshua Bengio. 2015a · 2057
Closest in time.
Vokenization: Improving Language Understanding with Contextualized, Visual-Grounded Supervision. In
Hao Tan and Mohit Bansal. 2020 · 2080
Closest in time.
Phrase-based image captioning. In
Rémi Lebret, Pedro Pinheiro, and Ronan Collobert. 2015 · 2094
Closest in time.