Fetching the paper…
Reading the bibliography…
Pre-trained language-vision models have shown remarkable performance on the visual question answering (VQA) task.
Visualbert: A simple and performant baseline for vision and language
Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang. 2019 · 1908
Earlier work this paper cites.
Analyzing the behavior of visual question answering models
Aishwarya Agrawal, Dhruv Batra, and Devi Parikh. 2016 · 1960
Earlier work this paper cites.
Duelling languages: Grammatical structure in codeswitching
Carol Myers-Scotton. 1997 · 1997
Earlier work this paper cites.
BLEU: A Method for Automatic Evaluation of Machine Translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Oscar: Object-semantics aligned pre-training for vision-language tasks
Xiujun Li, Xi Yin, Chunyuan Li, Xiaowei Hu, Pengchuan Zhang, Lei Zhang, Lijuan Wang, Houdong Hu, Li Dong, Furu Wei, et al. 2020 · 2004
Earlier work this paper cites.
ROUGE: A Package for Automatic Evaluation of Summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
A Study of Translation Edit Rate with Targeted Human Annotation
Matthew Snover, Bonnie Dorr, Richard Schwartz, Linnea Micciulla, and John Makhoul. 2006 · 2006
Earlier work this paper cites.
Visualizing data using t-sne
Laurens Van der Maaten and Geoffrey Hinton. 2008 · 2008
Earlier work this paper cites.
Indoor segmentation and support inference from rgbd images
Nathan Silberman, Derek Hoiem, Pushmeet Kohli, and Rob Fergus. 2012 · 2012
Earlier work this paper cites.
A simple, fast, and effective reparameterization of ibm model 2
Chris Dyer, Victor Chahuneau, and Noah A Smith. 2013 · 2013
Earlier work this paper cites.
On measuring the complexity of code-mixing
Björn Gambäck and Amitava Das. 2014 · 2014
Earlier work this paper cites.
VQA: Visual Question Answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh. 2015 · 2015
Earlier work this paper cites.
Are You Talking to a Machine? Dataset and Methods for Multilingual Image Question Answering
Haoyuan Gao, Junhua Mao, Jie Zhou, Zhiheng Huang, Lei Wang, and Wei Xu. 2015 · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. 2015 · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba. 2015 · 2015
Earlier work this paper cites.
Answer ka type kya he? Learning to Classify Questions in Code-Mixed Language
Khyathi Chandu Raghavi, Manoj Kumar Chinnakotla, and Manish Shrivastava. 2015 · 2015
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. 2016 · 2016
Cited alongside, same era.
Multimodal Compact Bilinear Pooling for Visual Question Answering and Visual Grounding
Akira Fukui, Dong Huk Park, Daylen Yang, Anna Rohrbach, Trevor Darrell, and Marcus Rohrbach. 2016 · 2016
Cited alongside, same era.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, et al. 2016 · 2016
Cited alongside, same era.
Visual7w: Grounded Question Answering in Images
Yuke Zhu, Oliver Groth, Michael Bernstein, and Li Fei-Fei. 2016 · 2016
Cited alongside, same era.
Complexity metric for code-mixed social media text
Souvick Ghosh, Satanu Ghosh, and Dipankar Das. 2017 · 2017
Bilinear Attention Networks
Jin-Hwa Kim, Jaehyun Jun, and Byoung-Tak Zhang. 2018 · 2018
Later among the works it cites.
Word translation without parallel data
Guillaume Lample, Alexis Conneau, Marc’Aurelio Ranzato, Ludovic Denoyer, and Hervé Jégou. 2018 · 2018
Later among the works it cites.
Learning to specialize with knowledge distillation for visual question answering
Jonghwan Mun, Kimin Lee, Jinwoo Shin, and Bohyung Han. 2018 · 2018
Later among the works it cites.
Word Embeddings for Code-Mixed Language Processing
Adithya Pratapa, Monojit Choudhury, and Sunayana Sitaram. 2018b · 2018
Later among the works it cites.
Code-Switching Language Modeling using Syntax-Aware Multi-Task Learning
Genta Indra Winata, Andrea Madotto, Chien-Sheng Wu, and Pascale Fung. 2018 · 2018
Later among the works it cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. 2017 · 2017
Cited alongside, same era.
SMPOST: Parts of Speech Tagger for Code-Mixed Indic Social Media Text
Deepak Gupta, Shubham Tripathi, Asif Ekbal, and Pushpak Bhattacharyya. 2017 · 2017
Cited alongside, same era.
Hadamard product for low-rank bilinear pooling
Jin-Hwa Kim, Kyoung Woon On, Woosang Lim, Jeonghee Kim, Jung-Woo Ha, and Byoung-Tak Zhang. 2017 · 2017
Cited alongside, same era.
SGDR: stochastic gradient descent with warm restarts
Ilya Loshchilov and Frank Hutter. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Multi-Modal Factorized Bilinear Pooling With Co-Attention Learning for Visual Question Answering
Zhou Yu, Jun Yu, Jianping Fan, and Dacheng Tao. 2017 · 2017
Cited alongside, same era.
Bottom-Up and Top-Down Attention for Image Captioning and Visual Question Answering
Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang. 2018 · 2018
Cited alongside, same era.
Later among the works it cites.
Compact trilinear interaction for visual question answering
Tuong Do, Thanh-Toan Do, Huy Tran, Erman Tjiputra, and Quang D Tran. 2019 · 2019
Later among the works it cites.
Language modeling for code-switching: Evaluation, integration of monolingual data, and discriminative training
Hila Gonen and Yoav Goldberg. 2019 · 2019
Later among the works it cites.
Gqa: A new dataset for real-world visual reasoning and compositional question answering
Drew A. Hudson and Christopher D. Manning. 2019 · 2019
Later among the works it cites.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Jiasen Lu, Dhruv Batra, Devi Parikh, and Stefan Lee. 2019 · 2019
Later among the works it cites.
LXMERT: Learning cross-modality encoder representations from transformers
Hao Tan and Mohit Bansal. 2019 · 2019
Later among the works it cites.
A semi-supervised approach to generate the code-mixed text using pre-trained encoder and transfer learning
Deepak Gupta, Asif Ekbal, and Pushpak Bhattacharyya. 2020a · 2020
Later among the works it cites.
Vl-bert: Pre-training of generic visual-linguistic representations
Weijie Su, Xizhou Zhu, Yue Cao, Bin Li, Lewei Lu, Furu Wei, and Jifeng Dai. 2020 · 2020
Later among the works it cites.
Unified vision-language pre-training for image captioning and vqa
Luowei Zhou, Hamid Palangi, Lei Zhang, Houdong Hu, Jason J Corso, and Jianfeng Gao. 2020 · 2020
Later among the works it cites.
Dealing with missing modalities in the visual question answer-difference prediction task through knowledge distillation
Jae Won Cho, Dong-Jin Kim, Jinsoo Choi, Yunjae Jung, and In So Kweon. 2021 · 2021
Closest in time.