Fetching the paper…
Reading the bibliography…
Learning multimodal representations involves integrating information from multiple heterogeneous sources of data.
Learning internal representations by error propagation
David E Rumelhart, Geoffrey E Hinton, and Ronald J Williams · 1985
Earlier work this paper cites.
Vocal quality factors: Analysis, synthesis, and perception
Donald G Childers and CK Lee · 1991
Earlier work this paper cites.
Tidigits speech corpus
R Gary Leonard and George Doddington · 1993
Earlier work this paper cites.
Convolutional networks for images, speech, and time series
Yann LeCun, Yoshua Bengio, et al · 1995
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
Human-computer interaction
Alan Dix, Janet Finlay, Gregory D Abowd, and Russell Beale · 2000
Earlier work this paper cites.
Audio-visual speech modeling for continuous speech recognition
Stéphane Dupont and Juergen Luettin · 2000
Earlier work this paper cites.
Affective computing
Rosalind W Picard · 2000
Earlier work this paper cites.
Neural synergy between kinetic vision and touch
Randolph Blake, Kenith V Sobel, and Thomas W James · 2004
Earlier work this paper cites.
Modeling multimodal human-computer interaction
Zeljko Obrenovic and Dusan Starcevic · 2004
Earlier work this paper cites.
An own-age bias in face recognition for children and older adults
Jeffrey S Anastasi and Matthew G Rhodes · 2005
Earlier work this paper cites.
Large-scale concept ontology for multimedia
Milind Naphade, John R Smith, Jelena Tesic, Shih-Fu Chang, Winston Hsu, Lyndon Kennedy, Alexander Hauptmann, and Jon Curtis · 2006
Earlier work this paper cites.
Probabilistic forecasts, calibration and sharpness
Tilmann Gneiting, Fadoua Balabdaoui, and Adrian E Raftery · 2007
Earlier work this paper cites.
Iemocap: Interactive emotional dyadic motion capture database
Carlos Busso, Murtaza Bulut, Chi-Chun Lee, Abe Kazemzadeh, Emily Mower, Samuel Kim, Jeannette N Chang, Sungbok Lee, and Shrikanth S Narayanan · 2008
Earlier work this paper cites.
Speaker identification on the scotus corpus
Jiahong Yuan and Mark Liberman · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Multimodal interfaces: A survey of principles, models and frameworks
Bruno Dumas, Denis Lalanne, and Sharon Oviatt · 2009
Earlier work this paper cites.
A survey of types of text noise and techniques to handle noisy text
L. Venkata Subramaniam, Shourya Roy, Tanveer A. Faruquie, and Sumit Negi · 2009
Earlier work this paper cites.
On the classification of emotional biosignals evoked while viewing affective pictures: an integrated data-mining-based approach for healthcare applications
Christos A Frantzidis, Charalampos Bratsas, Manousos A Klados, Evdokimos Konstantinidis, Chrysa D Lithari, Ana B Vivas, Christos L Papadelis, Eleni Kaldoudi, Costas Pappas, and Panagiotis D Bamidis · 2010
Earlier work this paper cites.
Emotion recognition and adaptation in spoken dialogue systems
Johannes Pittermann, Angela Pittermann, and Wolfgang Minker · 2010
Earlier work this paper cites.
Joint robust voicing detection and pitch estimation based on residual harmonics
Thomas Drugman and Abeer Alwan · 2011
Earlier work this paper cites.
Accessible ui design and multimodal interaction through hybrid tv platforms: towards a virtual-user centered design framework
Pascal Hamisu, Gregor Heinrich, Christoph Jung, Volker Hahn, Carlos Duarte, Pat Langdon, and Pradipta Biswas · 2011
Earlier work this paper cites.
Robots for use in autism research
Brian Scassellati, Henny Admoni, and Maja Matarić · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
Deep canonical correlation analysis
Galen Andrew, Raman Arora, Jeff Bilmes, and Karen Livescu · 2013
Earlier work this paper cites.
Practicing differential privacy in health care: A review
Fida Kamal Dankar and Khaled El Emam · 2013
Earlier work this paper cites.
Maxout networks
Ian Goodfellow, David Warde-Farley, Mehdi Mirza, Aaron Courville, and Yoshua Bengio · 2013
Earlier work this paper cites.
Wavelet maxima dispersion for breathy to tense voice discrimination
John Kane and Christer Gobl · 2013
Earlier work this paper cites.
Social robots as embedded reinforcers of social behavior in children with autism
Elizabeth S Kim, Lauren D Berkovits, Emily P Bernier, Dan Leyzberg, Frederick Shic, Rhea Paul, and Brian Scassellati · 2013
Earlier work this paper cites.
Zero-shot learning through cross-modal transfer
Richard Socher, Milind Ganjoo, Hamsa Sridhar, Osbert Bastani, Christopher D Manning, and Andrew Y Ng · 2013
Earlier work this paper cites.
Domain adaptation under target and conditional shift
Kun Zhang, Bernhard Schölkopf, Krikamol Muandet, and Zhikun Wang · 2013
Earlier work this paper cites.
Empirical evaluation of gated recurrent neural networks on sequence modeling
Junyoung Chung, Caglar Gulcehre, KyungHyun Cho, and Yoshua Bengio · 2014
Earlier work this paper cites.
Covarep—a collaborative voice analysis repository for speech technologies
Gilles Degottex, John Kane, Thomas Drugman, Tuomo Raitio, and Stefan Scherer · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D Manning · 2014
Earlier work this paper cites.
Intriguing properties of neural networks
Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus · 2014
Earlier work this paper cites.
A review paper: Noise models in digital image processing, 2015
Ajay Kumar Boyat and Brijendra Kumar Joshi · 2015
Earlier work this paper cites.
A survey of current datasets for vision and language research
Francis Ferraro, Nasrin Mostafazadeh, Ting-Hao Huang, Lucy Vanderwende, Jacob Devlin, Michel Galley, and Margaret Mitchell · 2015
Earlier work this paper cites.
Deep unordered composition rivals syntactic methods for text classification
Mohit Iyyer, Varun Manjunatha, Jordan Boyd-Graber, and Hal Daumé III · 2015
Earlier work this paper cites.
librosa: Audio and music signal analysis in python
Brian McFee, Colin Raffel, Dawen Liang, Daniel PW Ellis, Matt McVicar, Eric Battenberg, and Oriol Nieto · 2015
Earlier work this paper cites.
Deception detection using real-life trial data
Verónica Pérez-Rosas, Mohamed Abouelenien, Rada Mihalcea, and Mihai Burzo · 2015
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2015
Earlier work this paper cites.
On deep multi-view representation learning
Weiran Wang, Raman Arora, Karen Livescu, and Jeff Bilmes · 2015
Earlier work this paper cites.
Show, attend and tell: Neural image caption generation with visual attention
Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhudinov, Rich Zemel, and Yoshua Bengio · 2015
Earlier work this paper cites.
Analyzing the behavior of visual question answering models
Aishwarya Agrawal, Dhruv Batra, and Devi Parikh · 2016
Earlier work this paper cites.
Openface: an open source facial behavior analysis toolkit
Tadas Baltrušaitis, Peter Robinson, and Louis-Philippe Morency · 2016
Earlier work this paper cites.
Proceedings of the first workshop on NLP and computational social science
David Bamman, A. Seza Doğruöz, Jacob Eisenstein, Dirk Hovy, David Jurgens, Brendan O’Connor, Alice Oh, Oren Tsur, and Svitlana Volkova · 2016
Earlier work this paper cites.
Big data’s disparate impact
Solon Barocas and Andrew D Selbst · 2016
Earlier work this paper cites.
Man is to computer programmer as woman is to homemaker? Debiasing word embeddings
Tolga Bolukbasi, Kai-Wei Chang, James Y Zou, Venkatesh Saligrama, and Adam T Kalai · 2016
Earlier work this paper cites.
Dropout as a bayesian approximation: Representing model uncertainty in deep learning
Yarin Gal and Zoubin Ghahramani · 2016
Earlier work this paper cites.
David Ha, Andrew Dai, and Quoc V Le · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Mimic-iii, a freely accessible critical care database
Alistair EW Johnson, Tom J Pollard, Lu Shen, H Lehman Li-Wei, Mengling Feng, Mohammad Ghassemi, Benjamin Moody, Peter Szolovits, Leo Anthony Celi, and Roger G Mark · 2016
Earlier work this paper cites.
Being a supercook: Joint food attributes and multimodal content modeling for recipe retrieval and exploration
Weiqing Min, Shuqiang Jiang, Jitao Sang, Huayang Wang, Xinda Liu, and Luis Herranz · 2016
Earlier work this paper cites.
Towards multimodal deep learning for activity recognition on mobile devices
Valentin Radu, Nicholas D Lane, Sourav Bhattacharya, Cecilia Mascolo, Mahesh K Marina, and Fahim Kawsar · 2016
Earlier work this paper cites.
Multimodal research: Addressing the complexity of multimodal environments and the challenges for call
Sabine Tan, Kay O’Halloran, and Peter Wignell · 2016
Earlier work this paper cites.
Wavenet: A generative model for raw audio, 2016
Aaron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu · 2016
Earlier work this paper cites.
Show and tell: Lessons learned from the 2015 mscoco image captioning challenge
Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan · 2016
Earlier work this paper cites.
More than a million ways to be pushed. a high-fidelity experimental dataset of planar pushing
Kuan-Ting Yu, Maria Bauza, Nima Fazeli, and Alberto Rodriguez · 2016
Earlier work this paper cites.
Mosi: multimodal corpus of sentiment intensity and subjectivity analysis in online opinion videos
Amir Zadeh, Rowan Zellers, Eli Pincus, and Louis-Philippe Morency · 2016
Earlier work this paper cites.
VQA: Visual question answering
Aishwarya Agrawal, Jiasen Lu, Stanislaw Antol, Margaret Mitchell, C. Lawrence Zitnick, Devi Parikh, and Dhruv Batra · 2017
Earlier work this paper cites.
Gated multimodal units for information fusion
John Arevalo, Thamar Solorio, Manuel Montes-y Gómez, and Fabio A González · 2017
Cited alongside, same era.
Semantics derived automatically from language corpora contain human-like biases
Aylin Caliskan, Joanna J Bryson, and Arvind Narayanan · 2017
Cited alongside, same era.
Multimodal sentiment analysis with word-level fusion and reinforcement learning
Minghai Chen, Sen Wang, Paul Pu Liang, Tadas Baltrušaitis, Amir Zadeh, and Louis-Philippe Morency · 2017
Cited alongside, same era.
Language modeling with gated convolutional networks
Yann N Dauphin, Angela Fan, Michael Auli, and David Grangier · 2017
Cited alongside, same era.
Rico: A mobile app dataset for building data-driven design applications
Biplab Deka, Zifeng Huang, Chad Franzen, Joshua Hibschman, Daniel Afergan, Yang Li, Jeffrey Nichols, and Ranjitha Kumar · 2017
Cited alongside, same era.
Differentially private federated learning: A client level perspective
Multimodal transformer for unaligned multimodal language sequences
Yao-Hung Hubert Tsai, Shaojie Bai, Paul Pu Liang, J Zico Kolter, Louis-Philippe Morency, and Ruslan Salakhutdinov · 2019
Later among the works it cites.
Learning factorized multimodal representations
Yao-Hung Hubert Tsai, Paul Pu Liang, Amir Zadeh, Louis-Philippe Morency, and Ruslan Salakhutdinov · 2019
Later among the works it cites.
Probabilistic neural symbolic models for interpretable visual question answering
Ramakrishna Vedantam, Karan Desai, Stefan Lee, Marcus Rohrbach, Dhruv Batra, and Devi Parikh · 2019
Later among the works it cites.
Superglue: A stickier benchmark for general-purpose language understanding systems
Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman · 2019
Later among the works it cites.
Adaptive cross-modal few-shot learning
Chen Xing, Negar Rostamzadeh, Boris Oreshkin, and Pedro O O. Pinheiro · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Robin C Geyer, Tassilo Klein, and Moin Nabi · 2017
Cited alongside, same era.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh · 2017
Cited alongside, same era.
Facial expression analysis, 2017
iMotions · 2017
Cited alongside, same era.
The kinetics human action video dataset
Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, et al · 2017
Cited alongside, same era.
A review of affective computing: From unimodal analysis to multimodal fusion
Soujanya Poria, Erik Cambria, Rajiv Bajpai, and Amir Hussain · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Tensor fusion network for multimodal sentiment analysis
Amir Zadeh, Minghai Chen, Soujanya Poria, Erik Cambria, and Louis-Philippe Morency · 2017
Cited alongside, same era.
Keyang Xu, Mike Lam, Jingzhi Pang, Xin Gao, Charlotte Band, Piyush Mathur, Frank Papay, Ashish K Khanna, Jacek B Cywinski, Kamal Maheshwari, et al · 2019
Later among the works it cites.
CMU multimodal SDK
Amir Zadeh · 2019
Later among the works it cites.
Social-iq: A question answering benchmark for artificial social intelligence
Amir Zadeh, Michael Chan, Paul Pu Liang, Edmund Tong, and Louis-Philippe Morency · 2019
Later among the works it cites.
Inherent tradeoffs in learning fair representations
Han Zhao and Geoff Gordon · 2019
Later among the works it cites.
Deep supervised cross-modal retrieval
Liangli Zhen, Peng Hu, Xu Wang, and Dezhong Peng · 2019
Later among the works it cites.
Uncertainty quantification in multimodal ensembles of deep learners
Katherine E Brown, Farzana Ahamed Bhuiyan, and Douglas A Talbert · 2020
Later among the works it cites.
Uniter: Universal image-text representation learning
Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu · 2020
Later among the works it cites.
Beyond pinball loss: Quantile methods for calibrated uncertainty quantification
Youngseog Chung, Willie Neiswanger, Ian Char, and Jeff Schneider · 2020
Later among the works it cites.
Unsupervised natural language inference via decoupled multimodal contrastive learning, 2020
Wanyun Cui, Guangyu Zheng, and Wei Wang · 2020
Later among the works it cites.
See, hear, explore: Curiosity via audio-visual association
Victoria Dean, Shubham Tulsiani, and Abhinav Gupta · 2020
Later among the works it cites.
Clotho: An audio captioning dataset
Konstantinos Drossos, Samuel Lipping, and Tuomas Virtanen · 2020
Later among the works it cites.
Removing bias in multi-modal classifiers: Regularization by maximizing functional entropies
Itai Gat, Idan Schwartz, Alexander Schwing, and Tamir Hazan · 2020
Later among the works it cites.
Multimodal toolkit
Ken Gu · 2020
Later among the works it cites.
Manymodalqa: Modality disambiguation and qa over diverse inputs
Darryl Hannan, Akshay Jain, and Mohit Bansal · 2020
Later among the works it cites.
Does my multimodal model learn cross-modal interactions? it’s harder to tell than you might think!
Jack Hessel and Lillian Lee · 2020
Later among the works it cites.
Xtreme: A massively multilingual multi-task benchmark for evaluating cross-lingual generalisation
Junjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig, Orhan Firat, and Melvin Johnson · 2020
Later among the works it cites.
Open graph benchmark: Datasets for machine learning on graphs
Weihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong, Hongyu Ren, Bowen Liu, Michele Catasta, and Jure Leskovec · 2020
Later among the works it cites.
Multiplicative interactions and where to find them
Siddhant M. Jayakumar, Wojciech M. Czarnecki, Jacob Menick, Jonathan Schwarz, Jack Rae, Simon Osindero, Yee Whye Teh, Tim Harley, and Razvan Pascanu · 2020
Later among the works it cites.
The hateful memes challenge: Detecting hate speech in multimodal memes
Douwe Kiela, Hamed Firooz, Aravind Mohan, Vedanuj Goswami, Amanpreet Singh, Pratik Ringshia, and Davide Testuggine · 2020
Later among the works it cites.
Wilds: A benchmark of in-the-wild distribution shifts
Pang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie, Marvin Zhang, Akshay Balsubramani, Weihua Hu, Michihiro Yasunaga, Richard Lanas Phillips, Sara Beery, et al · 2020
Later among the works it cites.
Multimodal sensor fusion with differentiable filters
Michelle A Lee, Brent Yi, Roberto Martín-Martín, Silvio Savarese, and Jeannette Bohg · 2020
Later among the works it cites.
Making sense of vision and touch: Learning multimodal representations for contact-rich tasks
Michelle A Lee, Yuke Zhu, Peter Zachares, Matthew Tan, Krishnan Srinivasan, Silvio Savarese, Li Fei-Fei, Animesh Garg, and Jeannette Bohg · 2020
Later among the works it cites.
Enrico: A dataset for topic modeling of mobile ui designs
Luis A Leiva, Asutosh Hota, and Antti Oulasvirta · 2020
Later among the works it cites.
Awesome multimodal ml
Paul Liang · 2020
Later among the works it cites.
Towards debiasing sentence representations
Paul Pu Liang, Irene Li, Emily Zheng, Yao Chong Lim, Ruslan Salakhutdinov, and Louis-Philippe Morency · 2020
Later among the works it cites.
Think locally, act globally: Federated learning with local and global representations
Paul Pu Liang, Terrance Liu, Liu Ziyin, Nicholas B. Allen, Randy P. Auerbach, David Brent, Ruslan Salakhutdinov, and Louis-Philippe Morency · 2020
Later among the works it cites.
Cross-modal generalization: Learning in low resource modalities via meta-alignment
Paul Pu Liang, Peter Wu, Liu Ziyin, Louis-Philippe Morency, and Ruslan Salakhutdinov · 2020
Later among the works it cites.
Measuring social biases in grounded vision and language embeddings
Candace Ross, Boris Katz, and Andrei Barbu · 2020
Later among the works it cites.
Vision to language: Methods, metrics and datasets
Naeha Sharif, Uzair Nadeem, Syed Afaq Ali Shah, Mohammed Bennamoun, and Wei Liu · 2020
Later among the works it cites.
Vl-bert: Pre-training of generic visual-linguistic representations
Weijie Su, Xizhou Zhu, Yue Cao, Bin Li, Lewei Lu, Furu Wei, and Jifeng Dai · 2020
Later among the works it cites.
Learning relationships between text, audio, and video via deep canonical correlation for multimodal language analysis
Zhongkai Sun, Prathusha Sarma, William Sethares, and Yingyu Liang · 2020
Later among the works it cites.
Measuring robustness to natural distribution shifts in image classification
Rohan Taori, Achal Dave, Vaishaal Shankar, Nicholas Carlini, Benjamin Recht, and Ludwig Schmidt · 2020
Later among the works it cites.
Contrastive multiview coding
Yonglong Tian, Dilip Krishnan, and Phillip Isola · 2020
Later among the works it cites.
Methods for comparing uncertainty quantifications for material property predictions
Kevin Tran, Willie Neiswanger, Junwoong Yoon, Qingyang Zhang, Eric Xing, and Zachary W Ulissi · 2020
Later among the works it cites.
Multimodal routing: Improving local and global interpretability of multimodal language analysis
Yao-Hung Hubert Tsai, Martin Ma, Muqiao Yang, Ruslan Salakhutdinov, and Louis-Philippe Morency · 2020
Later among the works it cites.
Mimic-extract: A data extraction, preprocessing, and representation pipeline for mimic-iii
Shirly Wang, Matthew BA McDermott, Geeticka Chauhan, Marzyeh Ghassemi, Michael C Hughes, and Tristan Naumann · 2020
Later among the works it cites.
What makes training multi-modal classification networks hard?
Weiyao Wang, Du Tran, and Matt Feiszli · 2020
Later among the works it cites.
Uncertainty-aware multi-view co-training for semi-supervised medical image segmentation and domain adaptation
Yingda Xia, Dong Yang, Zhiding Yu, Fengze Liu, Jinzheng Cai, Lequan Yu, Zhuotun Zhu, Daguang Xu, Alan Yuille, and Holger Roth · 2020
Later among the works it cites.
Multimodal transformer for multimodal machine translation
Shaowei Yao and Xiaojun Wan · 2020
Later among the works it cites.
Foundations of multimodal co-learning
Amir Zadeh, Paul Pu Liang, and Louis-Philippe Morency · 2020
Later among the works it cites.
Moseas: A multimodal language dataset for spanish, portuguese, german and french
AmirAli Bagher Zadeh, Yansheng Cao, Simon Hessner, Paul Pu Liang, Soujanya Poria, and Louis-Philippe Morency · 2020
Later among the works it cites.
Rtfm: Generalising to new environment dynamics via reading
Victor Zhong, Tim Rocktäschel, and Edward Grefenstette · 2020
Later among the works it cites.
https://github.com/Jakobovski/free-spoken-digit-dataset
Free spoken digit dataset (fsdd) · 2021
Closest in time.
https://github.com/uncertainty-toolbox/uncertainty-toolbox , 2021
Uncertainty toolbox · 2021
Closest in time.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2021
Closest in time.
Multimodal Data Collection Made Easy: The EZ-MMLA Toolkit: A Data Collection Website That Provides Educators and Researchers with Easy Access to Multimodal Data Streams
Javaria Hassan, Jovin Leong, and Bertrand Schneider · 2021
Closest in time.
Transformer is all you need: Multimodal multitask learning with a unified transformer
Ronghang Hu and Amanpreet Singh · 2021
Closest in time.
Second opinion needed: communicating uncertainty in medical machine learning
Benjamin Kompa, Jasper Snoek, and Andrew L Beam · 2021
Closest in time.
Detect, reject, correct: Crossmodal compensation of corrupted sensors
Michelle A Lee, Matthew Tan, Yuke Zhu, and Jeannette Bohg · 2021
Closest in time.
Smil: Multimodal learning with severely missing modality
Mengmeng Ma, Jian Ren, Long Zhao, Sergey Tulyakov, Cathy Wu, and Xi Peng · 2021
Closest in time.
Zero-shot text-to-image generation
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever · 2021
Closest in time.
Multimodal fusion refiner networks
Sethuraman Sankaran, David Yang, and Ser-Nam Lim · 2021
Closest in time.
Worst of both worlds: Biases compound in pre-trained vision-and-language models
Tejas Srinivasan and Yonatan Bisk · 2021
Closest in time.
Multimodal{qa}: complex question answering over text, tables and images
Alon Talmor, Ori Yoran, Amnon Catav, Dan Lahav, Yizhong Wang, Akari Asai, Gabriel Ilharco, Hannaneh Hajishirzi, and Jonathan Berant · 2021
Closest in time.
Long range arena: A benchmark for efficient transformers
Yi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen, Dara Bahri, Philip Pham, Jinfeng Rao, Liu Yang, Sebastian Ruder, and Donald Metzler · 2021
Closest in time.
Mufasa: Multimodal fusion architecture search for electronic health records
Zhen Xu, David R So, and Andrew M Dai · 2021
Closest in time.