Fetching the paper…
Reading the bibliography…
Several multi-modality representation learning approaches such as LXMERT and ViLBERT have been proposed recently.
“Revisiting self-supervised visual representation learning,”
Alexander Kolesnikov, Xiaohua Zhai, and Lucas Beyer, · 1929
Earlier work this paper cites.
“Imagenet: A large-scale hierarchical image database,”
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei, · 2009
Earlier work this paper cites.
“Imagenet classification with deep convolutional neural networks,”
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton, · 2012
Earlier work this paper cites.
“Microsoft coco: Common objects in context,”
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick, · 2014
Earlier work this paper cites.
“Adam: A method for stochastic optimization,”
Diederik P Kingma and Jimmy Ba, · 2014
Earlier work this paper cites.
“Vqa: Visual question answering,”
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh, · 2015
Earlier work this paper cites.
“Faster r-cnn: Towards real-time object detection with region proposal networks,”
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun, · 2015
Earlier work this paper cites.
“Context encoders: Feature learning by inpainting,”
Deepak Pathak, Philipp Krahenbuhl, Jeff Donahue, Trevor Darrell, and Alexei A Efros, · 2016
Earlier work this paper cites.
“Unsupervised learning of visual representations by solving jigsaw puzzles,”
Mehdi Noroozi and Paolo Favaro, · 2016
Earlier work this paper cites.
“Stacked attention networks for image question answering,”
Zichao Yang, Xiaodong He, Jianfeng Gao, Li Deng, and Alex Smola, · 2016
Earlier work this paper cites.
“Multimodal compact bilinear pooling for visual question answering and visual grounding,”
Akira Fukui, Dong Huk Park, Daylen Yang, Anna Rohrbach, Trevor Darrell, and Marcus Rohrbach, · 2016
Earlier work this paper cites.
“Hadamard product for low-rank bilinear pooling,”
Jin-Hwa Kim, Kyoung-Woon On, Woosang Lim, Jeonghee Kim, Jung-Woo Ha, and Byoung-Tak Zhang, · 2016
Earlier work this paper cites.
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, et al., · 2016
Earlier work this paper cites.
“Visual genome: Connecting language and vision using crowdsourced dense image annotations,”
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, et al., · 2017
Earlier work this paper cites.
“A corpus of natural language for visual reasoning,”
Alane Suhr, Mike Lewis, James Yeh, and Yoav Artzi, · 2017
Earlier work this paper cites.
“Attention is all you need,”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin, · 2017
Cited alongside, same era.
“Mutan: Multimodal tucker fusion for visual question answering,”
Hedi Ben-Younes, Rémi Cadene, Matthieu Cord, and Nicolas Thome, · 2017
Cited alongside, same era.
“Making the V in VQA matter: Elevating the role of image understanding in Visual Question Answering,”
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh, · 2017
Cited alongside, same era.
“Bert: Pre-training of deep bidirectional transformers for language understanding,”
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova, · 2018
Cited alongside, same era.
“Improving language understanding by generative pre-training,”
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever, · 2018
Cited alongside, same era.
“Dynamic fusion with intra-and inter-modality attention flow for visual question answering,”
Peng Gao, Zhengkai Jiang, Haoxuan You, Pan Lu, Steven CH Hoi, Xiaogang Wang, and Hongsheng Li, · 2019
Later among the works it cites.
“Deep modular co-attention networks for visual question answering,”
Zhou Yu, Jun Yu, Yuhao Cui, Dacheng Tao, and Qi Tian, · 2019
Later among the works it cites.
“Gqa: A new dataset for real-world visual reasoning and compositional question answering,”
Drew A Hudson and Christopher D Manning, · 2019
Later among the works it cites.
“Momentum contrast for unsupervised visual representation learning,”
Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick, · 2019
Later among the works it cites.
“Reducing transformer depth on demand with structured dropout,”
Angela Fan, Edouard Grave, and Armand Joulin, · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Improved fusion of visual and language representations by dense symmetric co-attention for visual question answering,”
Duy-Kien Nguyen and Takayuki Okatani, · 2018
Cited alongside, same era.
“Bilinear attention networks,”
Jin-Hwa Kim, Jaehyun Jun, and Byoung-Tak Zhang, · 2018
Cited alongside, same era.
“Unsupervised feature learning via non-parametric instance discrimination,”
Zhirong Wu, Yuanjun Xiong, Stella X Yu, and Dahua Lin, · 2018
Cited alongside, same era.
“Glue: A multi-task benchmark and analysis platform for natural language understanding,”
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman, · 2018
Cited alongside, same era.
“Question-guided hybrid convolution for visual question answering,”
Peng Gao, Hongsheng Li, Shuang Li, Pan Lu, Yikang Li, Steven CH Hoi, and Xiaogang Wang, · 2018
Cited alongside, same era.
“Bottom-up and top-down attention for image captioning and visual question answering,”
Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang, · 2018
Cited alongside, same era.
“A corpus for reasoning about natural language grounded in photographs,”
Alane Suhr, Stephanie Zhou, Ally Zhang, Iris Zhang, Huajun Bai, and Yoav Artzi, · 2018
Cited alongside, same era.
“Contrastive multiview coding,”
Yonglong Tian, Dilip Krishnan, and Phillip Isola, · 2019
Later among the works it cites.
“Multi-modality latent interaction network for visual question answering,”
Peng Gao, Haoxuan You, Zhanpeng Zhang, Xiaogang Wang, and Hongsheng Li, · 2019
Later among the works it cites.
“Visualbert: A simple and performant baseline for vision and language,”
Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang, · 2019
Later among the works it cites.
“Uniter: Learning universal image-text representations,”
Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu, · 2019
Later among the works it cites.
“A simple framework for contrastive learning of visual representations,”
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton, · 2020
Closest in time.
“Improved baselines with momentum contrastive learning,”
Xinlei Chen, Haoqi Fan, Ross Girshick, and Kaiming He, · 2020
Closest in time.
“Multi-layer content interaction through quaternion product for visual question answering,”
Lei Shi, Shijie Geng, Kai Shuang, Chiori Hori, Songxiang Liu, Peng Gao, and Sen Su, · 2020
Closest in time.
“Character matters: Video story understanding with character-aware relations,”
Shijie Geng, Ji Zhang, Zuohui Fu, Peng Gao, Hang Zhang, and Gerard de Melo, · 2020
Closest in time.
“Spatio-temporal scene graphs for video dialog,”
Shijie Geng, Peng Gao, Chiori Hori, Jonathan Le Roux, and Anoop Cherian, · 2020
Closest in time.
“Show, attend and tell: Neural image caption generation with visual attention,”
Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhudinov, Rich Zemel, and Yoshua Bengio, · 2057
Closest in time.