Fetching the paper…
Reading the bibliography…
Deep Learning has implemented a wide range of applications and has become increasingly popular in recent years.
An analysis of visual question answering algorithms. In Proceedings of the IEEE International Conference on Computer Vision . 1965–1973
Kushal Kafle and Christopher Kanan. 2017 · 1973
Earlier work this paper cites.
Multimodal behavior therapy: Treating the" BASIC ID."
Arnold A Lazarus. 1973 · 1973
Earlier work this paper cites.
Multimodal signal detection: Independent decisions vs. integration
Robert M Mulligan and Marilyn L Shaw. 1980 · 1980
Earlier work this paper cites.
Boltzmann machines: Constraint satisfaction networks that learn
D Ackley, G Hinton, and T Sejnowski. 1985 · 1985
Earlier work this paper cites.
Automatic Lipreading to Enhance Speech Recognition (Speech Reading)
Eric David Petajan. 1985 · 1985
Earlier work this paper cites.
Information processing in dynamical systems: Foundations of harmony theory
Paul Smolensky. 1986 · 1986
Earlier work this paper cites.
Backpropagation applied to handwritten zip code recognition
Yann LeCun, Bernhard Boser, John S Denker, Donnie Henderson, Richard E Howard, Wayne Hubbard, and Lawrence D Jackel. 1989 · 1989
Earlier work this paper cites.
Integration of acoustic and visual speech signals using neural networks
Ben P Yuhas, Moise H Goldstein, and Terrence J Sejnowski. 1989 · 1989
Earlier work this paper cites.
Hidden Markov models for speech recognition
Biing Hwang Juang and Laurence R Rabiner. 1991 · 1991
Earlier work this paper cites.
WordNet: a lexical database for English
George A Miller. 1995 · 1995
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Bidirectional recurrent neural networks
Mike Schuster and Kuldip K Paliwal. 1997 · 1997
Earlier work this paper cites.
Murel: Multimodal relational reasoning for visual question answering. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . 1989–1998
Remi Cadene, Hedi Ben-Younes, Matthieu Cord, and Nicolas Thome. 2019 · 1998
Earlier work this paper cites.
ConceptNet—a practical commonsense reasoning tool-kit
Hugo Liu and Push Singh. 2004 · 2004
Earlier work this paper cites.
The AMI meeting corpus: A pre-announcement. In International workshop on machine learning for multimodal interaction . Springer, 28–39
Jean Carletta, Simone Ashby, Sebastien Bourban, Mike Flynn, Mael Guillemot, Thomas Hain, Jaroslav Kadlec, Vasilis Karaiskos, Wessel Kraaij, Melissa Kronenthal, et al · 2005
Earlier work this paper cites.
A bimodal face and body gesture database for automatic analysis of human nonverbal affective behavior. In 18th International conference on pattern recognition (ICPR’06) , Vol. 1. IEEE, 1148–1153
Hatice Gunes and Massimo Piccardi. 2006 · 2006
Earlier work this paper cites.
A fast learning algorithm for deep belief nets
Geoffrey E Hinton, Simon Osindero, and Yee-Whye Teh. 2006 · 2006
Earlier work this paper cites.
The eNTERFACE’05 audio-visual emotion database. In 22nd International Conference on Data Engineering Workshops (ICDEW’06) . IEEE, 8–8
Olivier Martin, Irene Kotsia, Benoit Macq, and Ioannis Pitas. 2006 · 2006
Earlier work this paper cites.
Dbpedia: A nucleus for a web of open data
Sören Auer, Christian Bizer, Georgi Kobilarov, Jens Lehmann, Richard Cyganiak, and Zachary Ives. 2007 · 2007
Earlier work this paper cites.
Freebase: a collaboratively created graph database for structuring human knowledge. In Proceedings of the 2008 ACM SIGMOD international conference on Management of data . 1247–1250
Kurt Bollacker, Colin Evans, Praveen Paritosh, Tim Sturge, and Jamie Taylor. 2008 · 2008
Earlier work this paper cites.
IEMOCAP: Interactive emotional dyadic motion capture database
Carlos Busso, Murtaza Bulut, Chi-Chun Lee, Abe Kazemzadeh, Emily Mower, Samuel Kim, Jeannette N Chang, Sungbok Lee, and Shrikanth S Narayanan. 2008 · 2008
Earlier work this paper cites.
The CALO meeting speech recognition and understanding system. In 2008 IEEE Spoken Language Technology Workshop . IEEE, 69–72
Gokhan Tur, Andreas Stolcke, Lynn Voss, John Dowding, Benoît Favre, Raquel Fernández, Matthew Frampton, Michael Frandsen, Clint Frederickson, Martin Graciarena, et al · 2008
Earlier work this paper cites.
Deep boltzmann machines. In Artificial intelligence and statistics . 448–455
Ruslan Salakhutdinov and Geoffrey Hinton. 2009 · 2009
Earlier work this paper cites.
Computers in the human interaction loop
Alex Waibel, Hartwig Steusloff, Rainer Stiefelhagen, and Kym Watson. 2009 · 2009
Earlier work this paper cites.
Collecting image annotations using amazon’s mechanical turk. In Proceedings of the NAACL HLT 2010 Workshop on Creating Speech and Language Data with Amazon’s Mechanical Turk . 139–147
Cyrus Rashtchian, Peter Young, Micah Hodosh, and Julia Hockenmaier. 2010 · 2010
Earlier work this paper cites.
Collecting highly parallel data for paraphrase evaluation. In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies . 190–200
David Chen and William B Dolan. 2011 · 2011
Earlier work this paper cites.
The semaine database: Annotated multimodal records of emotionally colored conversations between a person and a limited agent
Gary McKeown, Michel Valstar, Roddy Cowie, Maja Pantic, and Marc Schroder. 2011 · 2011
Earlier work this paper cites.
Im2text: Describing images using 1 million captioned photographs
Vicente Ordonez, Girish Kulkarni, and Tamara Berg. 2011 · 2011
Earlier work this paper cites.
Framing image description as a ranking task: Data, models and evaluation metrics
Micah Hodosh, Peter Young, and Julia Hockenmaier. 2013 · 2013
Earlier work this paper cites.
Grounding action descriptions in videos
Michaela Regneri, Marcus Rohrbach, Dominikus Wetzel, Stefan Thater, Bernt Schiele, and Manfred Pinkal. 2013 · 2013
Earlier work this paper cites.
Introducing the RECOLA multimodal corpus of remote collaborative and affective interactions. In 2013 10th IEEE international conference and workshops on automatic face and gesture recognition (FG) . IEEE, 1–8
Fabien Ringeval, Andreas Sonderegger, Juergen Sauer, and Denis Lalanne. 2013 · 2013
Earlier work this paper cites.
Empirical evaluation of gated recurrent neural networks on sequence modeling
Junyoung Chung, Caglar Gulcehre, KyungHyun Cho, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Generative adversarial nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context. In European conference on computer vision . Springer, 740–755
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. 2014 · 2014
Earlier work this paper cites.
A multi-world approach to question answering about real-world scenes based on uncertain input. In Advances in neural information processing systems . 1682–1690
Mateusz Malinowski and Mario Fritz. 2014 · 2014
Earlier work this paper cites.
Coherent multi-sentence video description with variable level of detail. In German conference on pattern recognition . Springer, 184–195
Anna Rohrbach, Marcus Rohrbach, Wei Qiu, Annemarie Friedrich, Manfred Pinkal, and Bernt Schiele. 2014 · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman. 2014 · 2014
Earlier work this paper cites.
Webchild: Harvesting and organizing commonsense knowledge from the web. In Proceedings of the 7th ACM international conference on Web search and data mining . 523–532
Niket Tandon, Gerard De Melo, Fabian Suchanek, and Gerhard Weikum. 2014 · 2014
Earlier work this paper cites.
Learning Spatiotemporal Features with 3D Convolutional Networks. ArXiv e-prints
D Tran, L Bourdev, R Fergus, L Torresani, and M Paluri. 2014 · 2014
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Peter Young, Alice Lai, Micah Hodosh, and Julia Hockenmaier. 2014 · 2014
Earlier work this paper cites.
Fuzzy deep belief networks for semi-supervised sentiment classification
Shusen Zhou, Qingcai Chen, and Xiaolong Wang. 2014 · 2014
Earlier work this paper cites.
Vqa: Visual question answering. In Proceedings of the IEEE international conference on computer vision . 2425–2433
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh. 2015 · 2015
Earlier work this paper cites.
Activitynet: A large-scale video benchmark for human activity understanding. In Proceedings of the ieee conference on computer vision and pattern recognition . 961–970
Fabian Caba Heilbron, Victor Escorcia, Bernard Ghanem, and Juan Carlos Niebles. 2015 · 2015
Earlier work this paper cites.
Fuzzy restricted Boltzmann machine for the enhancement of deep learning
CL Philip Chen, Chun-Yang Zhang, Long Chen, and Min Gan. 2015b · 2015
Earlier work this paper cites.
Microsoft coco captions: Data collection and evaluation server
Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Dollár, and C Lawrence Zitnick. 2015a · 2015
Earlier work this paper cites.
Librispeech: an asr corpus based on public domain audio books. In 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 5206–5210
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur. 2015 · 2015
Earlier work this paper cites.
Mutlimodal learning with deep boltzmann machine for emotion prediction in user generated videos. In Proceedings of the 5th ACM on International Conference on Multimedia Retrieval . 619–622
Lei Pang and Chong-Wah Ngo. 2015 · 2015
Earlier work this paper cites.
Exploring models and data for image question answering
Mengye Ren, Ryan Kiros, and Richard Zemel. 2015b · 2015
Earlier work this paper cites.
Shaoqing Ren, Kaiming He, Ross B Girshick, and Jian Sun. 2015a · 2015
Earlier work this paper cites.
Going deeper with convolutions. In Proceedings of the IEEE conference on computer vision and pattern recognition . 1–9
Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich. 2015 · 2015
Earlier work this paper cites.
Using descriptive video services to create a large data source for video annotation research
Atousa Torabi, Christopher Pal, Hugo Larochelle, and Aaron Courville. 2015 · 2015
Earlier work this paper cites.
MSP-IMPROV: An acted corpus of dyadic interactions to study emotion perception
Carlos Busso, Srinivas Parthasarathy, Alec Burmania, Mohammed AbdelWahab, Najmeh Sadoughi, and Emily Mower Provost. 2016 · 2016
Earlier work this paper cites.
Xception: Deep learning with depthwise separable convolutions, 2016
François Chollet. 2016 · 2016
Earlier work this paper cites.
Multimodal sparse coding for event detection
Youngjune Gwon, William Campbell, Kevin Brady, Douglas Sturim, Miriam Cha, and HT Kung. 2016 · 2016
Earlier work this paper cites.
Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition . 770–778
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
Adding chinese captions to images. In Proceedings of the 2016 ACM on International Conference on Multimedia Retrieval . 271–275
Xirong Li, Weiyu Lan, Jianfeng Dong, and Hailong Liu. 2016 · 2016
Earlier work this paper cites.
Senticap: Generating image descriptions with sentiments. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 30
Alexander Mathews, Lexing Xie, and Xuming He. 2016 · 2016
Earlier work this paper cites.
You only look once: Unified, real-time object detection. In Proceedings of the IEEE conference on computer vision and pattern recognition . 779–788
Joseph Redmon, Santosh Divvala, Ross Girshick, and Ali Farhadi. 2016 · 2016
Earlier work this paper cites.
Hollywood in homes: Crowdsourcing data collection for activity understanding. In European Conference on Computer Vision . Springer, 510–526
Gunnar A Sigurdsson, Gül Varol, Xiaolong Wang, Ali Farhadi, Ivan Laptev, and Abhinav Gupta. 2016 · 2016
Cited alongside, same era.
Support vector machine
Shan Suthaharan. 2016 · 2016
Cited alongside, same era.
Inception-v4, inception-resnet and the impact of residual connections on learning
Christian Szegedy, Sergey Ioffe, Vincent Vanhoucke, and Alex Alemi. 2016a · 2016
Cited alongside, same era.
Superseded-cstr vctk corpus: English multi-speaker corpus for cstr voice cloning toolkit
Christophe Veaux, Junichi Yamagishi, Kirsten MacDonald, et al · 2016
Cited alongside, same era.
Msr-vtt: A large video description dataset for bridging video and language. In Proceedings of the IEEE conference on computer vision and pattern recognition . 5288–5296
Jun Xu, Tao Mei, Ting Yao, and Yong Rui. 2016 · 2016
Competitive deep-belief networks for underwater acoustic target recognition
Honghui Yang, Sheng Shen, Xiaohui Yao, Meiping Sheng, and Chen Wang. 2018 · 2018
Later among the works it cites.
Video description: A survey of methods, datasets, and evaluation metrics
Nayyer Aafaq, Ajmal Mian, Wei Liu, Syed Zulqarnain Gilani, and Mubarak Shah. 2019b · 2019
Later among the works it cites.
An efficient normalized restricted Boltzmann machine for solving multiclass classification problems
Muhammad Aamir, Fazli Wahid, Hairulnizam Mahdin, and Nazri Mohd Nawi. 2019 · 2019
Later among the works it cites.
VQA-Med: Overview of the Medical Visual Question Answering Task at ImageCLEF 2019.. In CLEF (Working Notes)
Asma Ben Abacha, Sadid A Hasan, Vivek Datla, Joey Liu, Dina Demner-Fushman, and Henning Müller. 2019 · 2019
Later among the works it cites.
Image Captioning with Bidirectional Semantic Attention-Based Guiding of Long Short-Term Memory
Pengfei Cao, Zhongyi Yang, Liang Sun, Yanchun Liang, Mary Qu Yang, and Renchu Guan. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Video paragraph captioning using hierarchical recurrent neural networks. In Proceedings of the IEEE conference on computer vision and pattern recognition . 4584–4593
Haonan Yu, Jiang Wang, Zhiheng Huang, Yi Yang, and Wei Xu. 2016 · 2016
Cited alongside, same era.
Visual7w: Grounded question answering in images. In Proceedings of the IEEE conference on computer vision and pattern recognition . 4995–5004
Yuke Zhu, Oliver Groth, Michael Bernstein, and Li Fei-Fei. 2016 · 2016
Cited alongside, same era.
Deep voice: Real-time neural text-to-speech
Sercan O Arik, Mike Chrzanowski, Adam Coates, Gregory Diamos, Andrew Gibiansky, Yongguo Kang, Xian Li, John Miller, Andrew Ng, Jonathan Raiman, et al · 2017
Cited alongside, same era.
Mutan: Multimodal tucker fusion for visual question answering. In Proceedings of the IEEE international conference on computer vision . 2612–2620
Hedi Ben-Younes, Rémi Cadene, Matthieu Cord, and Nicolas Thome. 2017 · 2017
Cited alongside, same era.
Stylenet: Generating attractive visual captions with styles. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . 3137–3146
Chuang Gan, Zhe Gan, Xiaodong He, Jianfeng Gao, and Li Deng. 2017 · 2017
Cited alongside, same era.
Event classification in microblogs via social tracking
Yue Gao, Hanwang Zhang, Xibin Zhao, and Shuicheng Yan. 2017 · 2017
Cited alongside, same era.
Deep voice 2: Multi-speaker neural text-to-speech. In Advances in neural information processing systems . 2962–2970
Andrew Gibiansky, Sercan Arik, Gregory Diamos, John Miller, Kainan Peng, Wei Ping, Jonathan Raiman, and Yanqi Zhou. 2017 · 2017
Cited alongside, same era.
EmoChat: Bringing multimodal emotion detection to mobile conversation. In 2019 5th International Conference on Big Data Computing and Communications (BIGCOM) . IEEE, 213–221
Luyao Chong, Meng Jin, and Yuan He. 2019 · 2019
Later among the works it cites.
Unsupervised image captioning. In Proceedings of the IEEE conference on computer vision and pattern recognition . 4125–4134
Yang Feng, Lin Ma, Wei Liu, and Jiebo Luo. 2019 · 2019
Later among the works it cites.
Hidden markov models
Monica Franzese and Antonella Iuliano. 2019 · 2019
Later among the works it cites.
Deep multimodal representation learning: A survey
Wenzhong Guo, Jianwen Wang, and Shiping Wang. 2019b · 2019
Later among the works it cites.
Image caption generation with part of speech guidance
Xinwei He, Baoguang Shi, Xiang Bai, Gui-Song Xia, Zhaoxiang Zhang, and Weisheng Dong. 2019 · 2019
Later among the works it cites.
Muse-ing on the impact of utterance ordering on crowdsourced emotion annotations. In ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 7415–7419
Mimansa Jaiswal, Zakaria Aldeneh, Cristian-Paul Bara, Yuanhang Luo, Mihai Burzo, Rada Mihalcea, and Emily Mower Provost. 2019b · 2019
Later among the works it cites.
Controlling for confounders in multimodal emotion classification via adversarial learning. In 2019 International Conference on Multimodal Interaction . 174–184
Mimansa Jaiswal, Zakaria Aldeneh, and Emily Mower Provost. 2019a · 2019
Later among the works it cites.
Relation-aware graph attention network for visual question answering. In Proceedings of the IEEE International Conference on Computer Vision . 10313–10322
Linjie Li, Zhe Gan, Yu Cheng, and Jingjing Liu. 2019 · 2019
Later among the works it cites.
End-to-end video captioning with multitask reinforcement learning. In 2019 IEEE Winter Conference on Applications of Computer Vision (WACV) . IEEE, 339–348
Lijun Li and Boqing Gong. 2019 · 2019
Later among the works it cites.
Fuzzy Removing Redundancy Restricted Boltzmann Machine: improving learning speed and classification accuracy
Xueqin Lü, Lingzheng Meng, Chao Chen, and Peisong Wang. 2019 · 2019
Later among the works it cites.
Ok-vqa: A visual question answering benchmark requiring external knowledge. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . 3195–3204
Kenneth Marino, Mohammad Rastegari, Ali Farhadi, and Roozbeh Mottaghi. 2019 · 2019
Later among the works it cites.
Trends in integration of vision and language research: A survey of tasks, datasets, and methods
Aditya Mogadala, Marimuthu Kalimuthu, and Dietrich Klakow. 2019 · 2019
Later among the works it cites.
Streamlined dense video captioning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . 6588–6597
Jonghwan Mun, Linjie Yang, Zhou Ren, Ning Xu, and Bohyung Han. 2019 · 2019
Later among the works it cites.
Expression Classification in Children Using Mean Supervised Deep Boltzmann Machine. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops . 0–0
Shruti Nagpal, Maneet Singh, Mayank Vatsa, Richa Singh, and Afzel Noore. 2019 · 2019
Later among the works it cites.
Memory-attended recurrent network for video captioning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . 8347–8356
Wenjie Pei, Jiyuan Zhang, Xiangrong Wang, Lei Ke, Xiaoyong Shen, and Yu-Wing Tai. 2019 · 2019
Later among the works it cites.
A parallel gaussian–bernoulli restricted boltzmann machine for mining area classification with hyperspectral imagery
Kun Tan, Fuyu Wu, Qian Du, Peijun Du, and Yu Chen. 2019 · 2019
Later among the works it cites.
Shared multi-view data representation for multi-domain event detection
Zhenguo Yang, Qing Li, Wenyin Liu, and Jianming Lv. 2019 · 2019
Later among the works it cites.
Deep modular co-attention networks for visual question answering. In Proceedings of the IEEE conference on computer vision and pattern recognition . 6281–6290
Zhou Yu, Jun Yu, Yuhao Cui, Dacheng Tao, and Qi Tian. 2019 · 2019
Later among the works it cites.
Multimodal Representation Learning: Advances, Trends and Challenges. In 2019 International Conference on Machine Learning and Cybernetics (ICMLC) . IEEE, 1–6
Su-Fang Zhang, Jun-Hai Zhai, Bo-Jun Xie, Yan Zhan, and Xin Wang. 2019b · 2019
Later among the works it cites.
Reconstruct and represent video contents for captioning via reinforcement learning
Wei Zhang, Bairui Wang, Lin Ma, and Wei Liu. 2019a · 2019
Later among the works it cites.
Aqua: Asp-based visual question answering. In International Symposium on Practical Aspects of Declarative Languages . Springer, 57–72
Kinjal Basu, Farhad Shakerin, and Gopal Gupta. 2020 · 2020
Later among the works it cites.
Yonatan Bisk, Ari Holtzman, Jesse Thomason, Jacob Andreas, Yoshua Bengio, Joyce Chai, Mirella Lapata, Angeliki Lazaridou, Jonathan May, Aleksandr Nisnevich, et al · 2020
Later among the works it cites.
Stack-VS: Stacked Visual-Semantic Attention for Image Caption Generation
Ling Cheng, Wei Wei, Xianling Mao, Yong Liu, and Chunyan Miao. 2020 · 2020
Later among the works it cites.
Cross-subject multimodal emotion recognition based on hybrid fusion
Yucel Cimtay, Erhan Ekmekcioglu, and Seyma Caglar-Ozhan. 2020 · 2020
Later among the works it cites.
Parallel Tacotron: Non-Autoregressive and Controllable TTS
Isaac Elias, Heiga Zen, Jonathan Shen, Yu Zhang, Ye Jia, Ron Weiss, and Yonghui Wu. 2020 · 2020
Later among the works it cites.
Video2commonsense: Generating commonsense descriptions to enrich video captioning
Zhiyuan Fang, Tejas Gokhale, Pratyay Banerjee, Chitta Baral, and Yezhou Yang. 2020 · 2020
Later among the works it cites.
A survey on deep learning for multimodal data fusion
Jing Gao, Peng Li, Zhikui Chen, and Jianing Zhang. 2020 · 2020
Later among the works it cites.
Re-Attention for Visual Question Answering. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 34. 91–98
Wenya Guo, Ying Zhang, Xiaoping Wu, Jufeng Yang, Xiangrui Cai, and Xiaojie Yuan. 2020 · 2020
Later among the works it cites.
Video multimodal emotion recognition based on Bi-GRU and attention fusion
Ruo-Hong Huan, Jia Shu, Sheng-Lin Bao, Rong-Hua Liang, Peng Chen, and Kai-Kai Chi. 2020 · 2020
Later among the works it cites.
Adaptive attention model for image captioning
LU Jiasen, Caiming Xiong, and Richard Socher. 2020 · 2020
Later among the works it cites.
Different Contextual Window Sizes Based RNNs for Multimodal Emotion Detection in Interactive Conversations
Helang Lai, Hongying Chen, and Shuangyan Wu. 2020 · 2020
Later among the works it cites.
Multistep Deep System for Multimodal Emotion Detection With Invalid Data in the Internet of Things
Minjia Li, Lun Xie, Zeping Lv, Juan Li, and Zhiliang Wang. 2020 · 2020
Later among the works it cites.
Chinese Image Caption Generation via Visual Attention and Topic Modeling
Maofu Liu, Huijun Hu, Lingjun Li, Yan Yu, and Weili Guan. 2020a · 2020
Later among the works it cites.
Image caption generation with dual attention mechanism
Maofu Liu, Lingjun Li, Huijun Hu, Weili Guan, and Jing Tian. 2020b · 2020
Later among the works it cites.
Sibnet: Sibling convolutional encoder for video captioning
Sheng Liu, Zhou Ren, and Junsong Yuan. 2020c · 2020
Later among the works it cites.
RSVQA: Visual Question Answering for Remote Sensing Data
Sylvain Lobry, Diego Marcos, Jesse Murray, and Devis Tuia. 2020 · 2020
Later among the works it cites.
M3ER: Multiplicative Multimodal Emotion Recognition using Facial, Textual, and Speech Cues.. In AAAI . 1359–1367
Trisha Mittal, Uttaran Bhattacharya, Rohan Chandra, Aniket Bera, and Dinesh Manocha. 2020 · 2020
Later among the works it cites.
Robust Explanations for Visual Question Answering. In The IEEE Winter Conference on Applications of Computer Vision . 1577–1586
Badri Patro, Shivansh Patel, and Vinay Namboodiri. 2020 · 2020
Later among the works it cites.
Semantically Sensible Video Captioning (SSVC)
Md Rahman, Thasin Abedin, Khondokar SS Prottoy, Ayana Moshruba, Fazlul Hasan Siddiqui, et al · 2020
Later among the works it cites.
A review of deep learning with special emphasis on architectures, applications and recent trends
Saptarshi Sengupta, Sanchita Basak, Pallabi Saikia, Sayak Paul, Vasilios Tsalavoutis, Frederick Atiah, Vadlamani Ravi, and Alan Peters. 2020 · 2020
Later among the works it cites.
Cross-Lingual Image Caption Generation Based on Visual Attention Model
Bin Wang, Cungang Wang, Qian Zhang, Ying Su, Yang Wang, and Yanyan Xu. 2020 · 2020
Later among the works it cites.
Exploiting the local temporal information for video captioning
Ran Wei, Li Mi, Yaosi Hu, and Zhenzhong Chen. 2020a · 2020
Later among the works it cites.
Multi-Attention Generative Adversarial Network for image captioning
Yiwei Wei, Leiquan Wang, Haiwen Cao, Mingwen Shao, and Chunlei Wu. 2020b · 2020
Later among the works it cites.
Visual question answering model based on visual relationship detection
Yuling Xi, Yanning Zhang, Songtao Ding, and Shaohua Wan. 2020 · 2020
Later among the works it cites.
Deep Reinforcement Polishing Network for Video Captioning
Wanru Xu, Jian Yu, Zhenjiang Miao, Lili Wan, Yi Tian, and Qiang Ji. 2020 · 2020
Later among the works it cites.
Cross-modal knowledge reasoning for knowledge-based visual question answering
Jing Yu, Zihao Zhu, Yujing Wang, Weifeng Zhang, Yue Hu, and Jianlong Tan. 2020 · 2020
Later among the works it cites.
Multimodal intelligence: Representation learning, information fusion, and applications
Chao Zhang, Zichao Yang, Xiaodong He, and Li Deng. 2020b · 2020
Later among the works it cites.
Robust spike-and-slab deep Boltzmann machines for face denoising
Nan Zhang, Shifei Ding, Jian Zhang, and Xingyu Zhao. 2020a · 2020
Later among the works it cites.
HEU Emotion: a large-scale database for multimodal emotion recognition in the wild
Jing Chen, Chenhui Wang, Kejun Wang, Chaoqun Yin, Cong Zhao, Tao Xu, Xinyi Zhang, Ziqiang Huang, Meichen Liu, and Tao Yang. 2021 · 2021
Closest in time.
Video Captioning of Future Frames. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision . 980–989
Mehrdad Hosseinzadeh and Yang Wang. 2021 · 2021
Closest in time.
Cptr: Full transformer network for image captioning
Wei Liu, Sihan Chen, Longteng Guo, Xinxin Zhu, and Jing Liu. 2021 · 2021
Closest in time.
Online multimedia retrieval on CPU–GPU platforms with adaptive work partition
Rafael Souza, André Fernandes, Thiago SFX Teixeira, George Teodoro, and Renato Ferreira. 2021 · 2021
Closest in time.