Fetching the paper…
Reading the bibliography…
Multimodal language analysis is a burgeoning field of NLP that aims to simultaneously model a speaker's words, acoustical annotations, and facial expressions.
“Iemocap: Interactive emotional dyadic motion capture database,”
Carlos Busso, Murtaza Bulut, Chi-Chun Lee, Abe Kazemzadeh, Emily Mower, Samuel Kim, Jeannette N Chang, Sungbok Lee, and Shrikanth S Narayanan, · 2008
Earlier work this paper cites.
“Towards Multimodal Sentiment Analysis: Harvesting Opinions from The Web,”
Louis-Philippe Morency, Rada Mihalcea, and Payal Doshi, · 2011
Earlier work this paper cites.
“Neural machine translation by jointly learning to align and translate,”
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio, · 2014
Earlier work this paper cites.
“Librispeech: an asr corpus based on public domain audio books,”
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur, · 2015
Earlier work this paper cites.
“Mosi: multimodal corpus of sentiment intensity and subjectivity analysis in online opinion videos,”
Amir Zadeh, Rowan Zellers, Eli Pincus, and Louis-Philippe Morency, · 2016
Earlier work this paper cites.
“Gaussian error linear units (gelus),”
Dan Hendrycks and Kevin Gimpel, · 2016
Earlier work this paper cites.
“A survey of multimodal sentiment analysis,”
Mohammad Soleymani, David Garcia, Brendan Jou, Björn Schuller, Shih-Fu Chang, and Maja Pantic, · 2017
Earlier work this paper cites.
“Multimodal sentiment analysis with word-level fusion and reinforcement learning,”
Minghai Chen, Sen Wang, Paul Pu Liang, Tadas Baltrušaitis, Amir Zadeh, and Louis-Philippe Morency, · 2017
Earlier work this paper cites.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin, · 2017
Earlier work this paper cites.
“Tensor fusion network for multimodal sentiment analysis,”
Amir Zadeh, Minghai Chen, Soujanya Poria, Erik Cambria, and Louis-Philippe Morency, · 2017
Earlier work this paper cites.
“Decoupled weight decay regularization,”
Ilya Loshchilov and Frank Hutter, · 2017
Earlier work this paper cites.
“Multimodal language analysis in the wild: CMU-MOSEI dataset and interpretable dynamic fusion graph,”
AmirAli Bagher Zadeh, Paul Pu Liang, Soujanya Poria, Erik Cambria, and Louis-Philippe Morency, · 2018
Earlier work this paper cites.
“Bert: Pre-training of deep bidirectional transformers for language understanding,”
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova, · 2018
Earlier work this paper cites.
“Multimodal language analysis in the wild: Cmu-mosei dataset and interpretable dynamic fusion graph,”
AmirAli Bagher Zadeh, Paul Pu Liang, Soujanya Poria, Erik Cambria, and Louis-Philippe Morency, · 2018
Earlier work this paper cites.
“Multimodal machine learning: A survey and taxonomy,”
Tadas Baltrušaitis, Chaitanya Ahuja, and Louis-Philippe Morency, · 2018
Earlier work this paper cites.
“Efficient low-rank multimodal fusion with modality-specific factors,”
Zhun Liu, Ying Shen, Varun Bharadhwaj Lakshminarasimhan, Paul Pu Liang, Amir Zadeh, and Louis-Philippe Morency, · 2018
Earlier work this paper cites.
“Multi-attention recurrent network for human communication comprehension,”
Amir Zadeh, Paul Pu Liang, Soujanya Poria, Prateek Vij, Erik Cambria, and Louis-Philippe Morency, · 2018
Cited alongside, same era.
“Memory fusion network for multi-view sequential learning,”
Amir Zadeh, Paul Pu Liang, Navonil Mazumder, Soujanya Poria, Erik Cambria, and Louis-Philippe Morency, · 2018
Cited alongside, same era.
“Voxceleb2: Deep speaker recognition,”
Joon Son Chung, Arsha Nagrani, and Andrew Zisserman, · 2018
Cited alongside, same era.
“Glue: A multi-task benchmark and analysis platform for natural language understanding,”
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman, · 2018
Cited alongside, same era.
“Learning factorized multimodal representations,”
Yao-Hung Hubert Tsai, Paul Pu Liang, Amir Zadeh, Louis-Philippe Morency, and Ruslan Salakhutdinov, · 2018
“Huggingface’s transformers: State-of-the-art natural language processing,”
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, et al., · 2019
Later among the works it cites.
Seattle, USA, July 2020, Association for Computational Linguistics
“Second grand-challenge and workshop on multimodal language (challenge-hml),” · 2020
Later among the works it cites.
“Language models are few-shot learners,”
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al., · 2020
Later among the works it cites.
“wav2vec 2.0: A framework for self-supervised learning of speech representations,”
Alexei Baevski, Henry Zhou, Abdelrahman Mohamed, and Michael Auli, · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Multimodal transformer for unaligned multimodal language sequences,”
Yao-Hung Hubert Tsai, Shaojie Bai, Paul Pu Liang, J. Zico Kolter, Louis-Philippe Morency, and Ruslan Salakhutdinov, · 2019
Cited alongside, same era.
“Roberta: A robustly optimized bert pretraining approach,”
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov, · 2019
Cited alongside, same era.
“Language models are unsupervised multitask learners,”
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al., · 2019
Cited alongside, same era.
“Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks,”
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee, · 2019
Cited alongside, same era.
“Ur-funny: A multimodal language dataset for understanding humor,”
Md Kamrul Hasan, Wasifur Rahman, Amir Zadeh, Jianyuan Zhong, Md Iftekhar Tanveer, Louis-Philippe Morency, et al., · 2019
Cited alongside, same era.
“Towards multimodal sarcasm detection (an _obviously_ perfect paper),”
Santiago Castro, Devamanyu Hazarika, Verónica Pérez-Rosas, Roger Zimmermann, Rada Mihalcea, and Soujanya Poria, · 2019
Cited alongside, same era.
“Words can shift: Dynamically adjusting word representations using nonverbal behaviors,”
Yansen Wang, Ying Shen, Zhun Liu, Paul Pu Liang, Amir Zadeh, and Louis-Philippe Morency, · 2019
Cited alongside, same era.
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al., · 2020
Later among the works it cites.
“Learning relationships between text, audio, and video via deep canonical correlation for multimodal language analysis,”
Zhongkai Sun, Prathusha Sarma, William Sethares, and Yingyu Liang, · 2020
Later among the works it cites.
Jianing Yang, Yongxin Wang, Ruitao Yi, Yuying Zhu, Azaan Rehman, Amir Zadeh, Soujanya Poria, and Louis-Philippe Morency, · 2020
Later among the works it cites.
“Integrating multimodal information in large pretrained transformers,”
Wasifur Rahman, Md Kamrul Hasan, Sangwu Lee, Amir Zadeh, Chengfeng Mao, Louis-Philippe Morency, and Ehsan Hoque, · 2020
Later among the works it cites.
Shamane Siriwardhana, Andrew Reis, Rivindu Weerasekera, and Suranga Nanayakkara, · 2020
Later among the works it cites.
“Misa: Modality-invariant and-specific representations for multimodal sentiment analysis,”
Devamanyu Hazarika, Roger Zimmermann, and Soujanya Poria, · 2020
Later among the works it cites.
“Learning modality-specific representations with self-supervised multi-task learning for multimodal sentiment analysis,”
Wenmeng Yu, Hua Xu, Yuan Ziqi, and Wu Jiele, · 2021
Closest in time.
“Self-supervised learning with cross-modal transformers for emotion recognition,”
Aparna Khare, Srinivas Parthasarathy, and Shiva Sundaram, · 2021
Closest in time.
“Humor knowledge enriched transformer for understanding multimodal humor,”
Md Kamrul Hasan, Sangwu Lee, Wasifur Rahman, Amir Zadeh, Rada Mihalcea, Louis-Philippe Morency, and Ehsan Hoque, · 2021
Closest in time.
“Pretrained transformers as universal computation engines,”
Kevin Lu, Aditya Grover, Pieter Abbeel, and Igor Mordatch, · 2021
Closest in time.
“Multimodal few-shot learning with frozen language models,”
Maria Tsimpoukelli, Jacob Menick, Serkan Cabi, SM Eslami, Oriol Vinyals, and Felix Hill, · 2021
Closest in time.