Fetching the paper…
Reading the bibliography…
Creating a meaningful representation by fusing single modalities (e.g., text, images, or audio) is the core concept of multimodal learning.
Integration of acoustic and visual speech signals using neural networks
B.P. Yuhas, M.H. Goldstein, and T.J. Sejnowski · 1989
Earlier work this paper cites.
Representation Learning: A Review and New Perspectives
Y. Bengio, A. Courville, and P. Vincent · 2013
Earlier work this paper cites.
Multimodal learning with deep Boltzmann machines
Nitish Srivastava and Ruslan Salakhutdinov · 2014
Earlier work this paper cites.
Ups and Downs: Modeling the Visual Evolution of Fashion Trends with One-Class Collaborative Filtering
Ruining He and Julian McAuley · 2016
Earlier work this paper cites.
The MovieLens Datasets: History and Context
F. Maxwell Harper and Joseph A. Konstan · 2016
Earlier work this paper cites.
Deep Residual Learning for Image Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Multimodal Classification Fusion in Real-World Scenarios
Ignazio Gallo, Alessandro Calefati, and Shah Nawaz · 2017
Earlier work this paper cites.
Multimodal Machine Learning: A Survey and Taxonomy
Tadas Baltrusaitis, Chaitanya Ahuja, and Louis-Philippe Morency · 2018
Earlier work this paper cites.
Efficient Large-Scale Multi-Modal Classification
D. Kiela, E. Grave, A. Joulin, and T. Mikolov · 2018
Earlier work this paper cites.
Local Sensitive Hashing (LSH) and Convolutional Neural Networks (CNNs) for Object Recognition
Mehdi Ghayoumi, Miguel Gomez, Kate E. Baumstein, Narindra Persaud, and Andrew J. Perlowin · 2018
Cited alongside, same era.
Towards VQA Models That Can Read
Amanpreet Singh, Vivek Natarajan, Meet Shah, Yu Jiang, Xinlei Chen, Dhruv Batra, Devi Parikh, and Marcus Rohrbach · 2019
Cited alongside, same era.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
Multi-Label Product Categorization Using Multi-Modal Fusion Models
Pasawee Wirojwatanakul and Artit Wangperawong · 2019
Cited alongside, same era.
A Review of Hashing Methods for Multimodal Retrieval
Wenming Cao, Wenshuo Feng, Qiubin Lin, Guitao Cao, and Zhihai He · 2020
Cited alongside, same era.
I know why you like this movie: Interpretable Efficient Multimodal Recommender
Barbara Rychalska, Dominika Basaj Basaj, Jacek Dabrowski, and Michal Daniluk · 2020
Later among the works it cites.
A comparative study of outfit recommendation methods with a focus on attention-based fusion
Katrien Laenen and Marie-Francine Moens · 2020
Later among the works it cites.
Cornac: A comparative framework for multimodal recommender systems
Aghiles Salah, Quoc-Tuan Truong, and Hady W. Lauw · 2020
Later among the works it cites.
An efficient manifold density estimator for all recommendation systems
Jacek Dabrowski, Barbara Rychalska, Michal Daniluk, Dominika Basaj, Konrad Goluchowski, Piotr Babel, Andrzej Michalowski, and Adam Jakubowski · 2020
Later among the works it cites.
A benchmarking study of classification techniques for behavioral data
Sofie De Cnudde, David Martens, Theodoros Evgeniou, and Foster Provost · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A Survey on Deep Learning for Multimodal Data Fusion
Jing Gao, Peng Li, Zhikui Chen, and Jianing Zhang · 2020
Cited alongside, same era.
Multimodal multitask deep learning model for Alzheimer’s disease progression detection based on time series data
Shaker El-Sappagh, Tamer Abuhmed, S.M. Riazul Islam, and Kyung Sup Kwak · 2020
Cited alongside, same era.
MuSE: a Multimodal Dataset of Stressed Emotion
Mimansa Jaiswal, Cristian-Paul Bara, Yuanhang Luo, Mihai Burzo, Rada Mihalcea, and Emily Mower Provost · 2020
Cited alongside, same era.
Deep multimodal fusion for semantic image segmentation: A survey
Yifei Zhang, Désiré Sidibé, Olivier Morel, and Fabrice Mériaudeau · 2020
Cited alongside, same era.
Later among the works it cites.
Vision-and-Language or Vision-for-Language? On Cross-Modal Influence in Multimodal Transformers
Stella Frank, Emanuele Bugliarello, and Desmond Elliott · 2021
Later among the works it cites.
Cleora: A Simple, Strong and Scalable Graph Embedding Scheme
Barbara Rychalska, Dominika Basaj, Jacek Dabrowski, and Michal Daniluk · 2021
Later among the works it cites.
Trustworthy Machine Learning
Kush R. Varshney · 2022
Closest in time.