Fetching the paper…
Reading the bibliography…
Designing effective neural networks is fundamentally important in deep multimodal learning.
Integration of acoustic and visual speech signals using neural networks
Ben P Yuhas, Moise H Goldstein, and Terrence J Sejnowski. 1989 · 1989
Earlier work this paper cites.
Audio-visual speech modeling for continuous speech recognition
Stéphane Dupont and Juergen Luettin. 2000 · 2000
Earlier work this paper cites.
Referitgame: Referring to objects in photographs of natural scenes. In Conference on Empirical Methods in Natural Language Processing (EMNLP) . 787–798
Sahar Kazemzadeh, Vicente Ordonez, Mark Matten, and Tamara Berg. 2014 · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context. In European Conference on Computer Vision (ECCV) . 740–755
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. 2014 · 2014
Earlier work this paper cites.
Glove: Global Vectors for Word Representation.. In Conference on Empirical Methods in Natural Language Processing (EMNLP) , Vol. 14. 1532–1543
Jeffrey Pennington, Richard Socher, and Christopher D Manning. 2014 · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman. 2014 · 2014
Earlier work this paper cites.
Vqa: Visual question answering. In IEEE International Conference on Computer Vision (ICCV) . 2425–2433
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh. 2015 · 2015
Earlier work this paper cites.
Microsoft coco captions: Data collection and evaluation server
Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Dollár, and C Lawrence Zitnick. 2015 · 2015
Earlier work this paper cites.
Fast r-cnn. In IEEE International Conference on Computer Vision (ICCV) . 1440–1448
Ross Girshick. 2015 · 2015
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR) . 3128–3137
Andrej Karpathy and Li Fei-Fei. 2015 · 2015
Earlier work this paper cites.
Flickr30k entities: Collecting region-to-phrase correspondences for richer image-to-sentence models. In Proceedings of the IEEE international conference on computer vision . 2641–2649
Bryan A Plummer, Liwei Wang, Chris M Cervantes, Juan C Caicedo, Julia Hockenmaier, and Svetlana Lazebnik. 2015 · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks. In Advances in neural information processing systems (NIPS) . 91–99
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. 2015 · 2015
Earlier work this paper cites.
Show, Attend and Tell: Neural Image Caption Generation with Visual Attention.. In International Conference on Machine Learning (ICML) , Vol. 14. 77–81
Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron C Courville, Ruslan Salakhutdinov, Richard S Zemel, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. 2016 · 2016
Earlier work this paper cites.
Multimodal Compact Bilinear Pooling for Visual Question Answering and Visual Grounding
Akira Fukui, Dong Huk Park, Daylen Yang, Anna Rohrbach, Trevor Darrell, and Marcus Rohrbach. 2016 · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
Ssd: Single shot multibox detector. In European Conference on Computer Vision (ECCV) . Springer, 21–37
Wei Liu, Dragomir Anguelov, Dumitru Erhan, Christian Szegedy, Scott Reed, Cheng-Yang Fu, and Alexander C Berg. 2016 · 2016
Earlier work this paper cites.
Generation and comprehension of unambiguous object descriptions. In IEEE International Conference on Computer Vision (ICCV) . 11–20
Junhua Mao, Jonathan Huang, Alexander Toshev, Oana Camburu, Alan L Yuille, and Kevin Murphy. 2016 · 2016
Earlier work this paper cites.
Grounding of textual phrases in images by reconstruction. In European Conference on Computer Vision (ECCV) . 817–834
Anna Rohrbach, Marcus Rohrbach, Ronghang Hu, Trevor Darrell, and Bernt Schiele. 2016 · 2016
Earlier work this paper cites.
Generalized deep transfer networks for knowledge propagation in heterogeneous domains
Jinhui Tang, Xiangbo Shu, Zechao Li, Guo-Jun Qi, and Jingdong Wang. 2016 · 2016
Earlier work this paper cites.
Modeling context in referring expressions. In European Conference on Computer Vision (ECCV) . Springer, 69–85
Licheng Yu, Patrick Poirson, Shan Yang, Alexander C Berg, and Tamara L Berg. 2016 · 2016
Earlier work this paper cites.
Neural architecture search with reinforcement learning
Barret Zoph and Quoc V Le. 2016 · 2016
Cited alongside, same era.
Mutan: Multimodal tucker fusion for visual question answering. In IEEE International Conference on Computer Vision (ICCV)
Hedi Ben-Younes, Rémi Cadene, Matthieu Cord, and Nicolas Thome. 2017 · 2017
Cited alongside, same era.
Vse++: Improving visual-semantic embeddings with hard negatives
Fartash Faghri, David J Fleet, Jamie Ryan Kiros, and Sanja Fidler. 2017 · 2017
Cited alongside, same era.
Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. 2017 · 2017
Cited alongside, same era.
Mask r-cnn. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR) . 2961–2969
Kaiming He, Georgia Gkioxari, Piotr Dollár, and Ross Girshick. 2017 · 2017
Darts: Differentiable architecture search
Hanxiao Liu, Karen Simonyan, and Yiming Yang. 2018 · 2018
Later among the works it cites.
Improved Fusion of Visual and Language Representations by Dense Symmetric Co-Attention for Visual Question Answering
Duy-Kien Nguyen and Takayuki Okatani. 2018 · 2018
Later among the works it cites.
Efficient Neural Architecture Search via Parameters Sharing. In International Conference on Machine Learning (ICML) . 4095–4104
Hieu Pham, Melody Guan, Barret Zoph, Quoc Le, and Jeff Dean. 2018 · 2018
Later among the works it cites.
Beyond Bilinear: Generalized Multimodal Factorized High-Order Pooling for Visual Question Answering
Zhou Yu, Jun Yu, Chenchao Xiang, Jianping Fan, and Dacheng Tao. 2018b · 2018
Later among the works it cites.
Rethinking Diversified and Discriminative Proposal Generation for Visual Grounding
Zhou Yu, Jun Yu, Chenchao Xiang, Zhou Zhao, Qi Tian, and Dacheng Tao. 2018c · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Modeling relationships in referential expressions with compositional modular networks. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR) . 1115–1124
Ronghang Hu, Marcus Rohrbach, Jacob Andreas, Trevor Darrell, and Kate Saenko. 2017 · 2017
Cited alongside, same era.
Hadamard Product for Low-rank Bilinear Pooling. In International Conference on Learning Representation (ICLR)
Jin-Hwa Kim, Kyoung Woon On, Woosang Lim, Jeonghee Kim, Jung-Woo Ha, and Byoung-Tak Zhang. 2017 · 2017
Cited alongside, same era.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, et al · 2017
Cited alongside, same era.
Dual attention networks for multimodal reasoning and matching. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR) . 299–307
Hyeonseob Nam, Jung-Woo Ha, and Jeonghee Kim. 2017 · 2017
Cited alongside, same era.
Tips and tricks for visual question answering: Learnings from the 2017 challenge. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR . 4223–4232
Damien Teney, Peter Anderson, Xiaodong He, and Anton Van Den Hengel. 2018 · 2017
Cited alongside, same era.
Attention is all you need. In Advances in Neural Information Processing Systems (NIPS) . 6000–6010
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Multi-modal Factorized Bilinear Pooling with Co-Attention Learning for Visual Question Answering
Zhou Yu, Jun Yu, Jianping Fan, and Dacheng Tao. 2017b · 2017
Cited alongside, same era.
Later among the works it cites.
Taskonomy: Disentangling task transfer learning. In Proceedings of the IEEE conference on computer vision and pattern recognition . 3712–3722
Amir R Zamir, Alexander Sax, William Shen, Leonidas J Guibas, Jitendra Malik, and Silvio Savarese. 2018 · 2018
Later among the works it cites.
Grounding Referring Expressions in Images by Variational Context. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
Hanwang Zhang, Yulei Niu, and Shih-Fu Chang. 2018 · 2018
Later among the works it cites.
Learning transferable architectures for scalable image recognition. In IEEE conference on Computer Vision and Pattern Recognition (CVPR) . 8697–8710
Barret Zoph, Vijay Vasudevan, Jonathon Shlens, and Quoc V Le. 2018 · 2018
Later among the works it cites.
Uniter: Learning universal image-text representations
Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu. 2019 · 2019
Later among the works it cites.
Fairnas: Rethinking evaluation fairness of weight sharing neural architecture search
Xiangxiang Chu, Bo Zhang, Ruijun Xu, and Jixiang Li. 2019 · 2019
Later among the works it cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Conference of the NAACL-HLT . 4171–4186
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
Dynamic fusion with intra-and inter-modality attention flow for visual question answering. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR) . 6639–6648
Peng Gao, Zhengkai Jiang, Haoxuan You, Pan Lu, Steven CH Hoi, Xiaogang Wang, and Hongsheng Li. 2019 · 2019
Later among the works it cites.
Nas-fpn: Learning scalable feature pyramid architecture for object detection. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR) . 7036–7045
Golnaz Ghiasi, Tsung-Yi Lin, and Quoc V Le. 2019 · 2019
Later among the works it cites.
Visualbert: A simple and performant baseline for vision and language
Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang. 2019 · 2019
Later among the works it cites.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks. In Advances in Neural Information Processing Systems (NIPS) . 13–23
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. 2019 · 2019
Later among the works it cites.
Mfas: Multimodal fusion architecture search. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR) . 6966–6975
Juan-Manuel Pérez-Rúa, Valentin Vielzeuf, Stéphane Pateux, Moez Baccouche, and Frédéric Jurie. 2019 · 2019
Later among the works it cites.
David R So, Chen Liang, and Quoc V Le. 2019 · 2019
Later among the works it cites.
LXMERT: Learning Cross-Modality Encoder Representations from Transformers. In Conference on Empirical Methods in Natural Language Processing (EMNLP) . 5103–5114
Hao Tan and Mohit Bansal. 2019 · 2019
Later among the works it cites.
CAMP: Cross-Modal Adaptive Message Passing for Text-Image Retrieval. In IEEE International Conference on Computer Vision (ICCV) . 5764–5773
Zihao Wang, Xihui Liu, Hongsheng Li, Lu Sheng, Junjie Yan, Xiaogang Wang, and Jing Shao. 2019 · 2019
Later among the works it cites.
Pc-darts: Partial channel connections for memory-efficient differentiable architecture search
Yuhui Xu, Lingxi Xie, Xiaopeng Zhang, Xin Chen, Guo-Jun Qi, Qi Tian, and Hongkai Xiong. 2019 · 2019
Later among the works it cites.
Deep Modular Co-Attention Networks for Visual Question Answering. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR) . 6281–6290
Zhou Yu, Jun Yu, Yuhao Cui, Dacheng Tao, and Qi Tian. 2019 · 2019
Later among the works it cites.