Fetching the paper…
Reading the bibliography…
Visual question answering (VQA) is one of the crucial vision-and-language tasks.
Unifying vision-and-language tasks via text generation
Jaemin Cho, Jie Lei, Hao Tan, and Mohit Bansal. 2021 · 1942
Earlier work this paper cites.
Analyzing the behavior of visual question answering models
Aishwarya Agrawal, Dhruv Batra, and Devi Parikh. 2016 · 1960
Earlier work this paper cites.
Open-domain question-answering
John M. Prager. 2006 · 2006
Earlier work this paper cites.
Typing candidate answers using type coercion
J. William Murdock, Aditya Kalyanpur, Chris Welty, James Fan, David A. Ferrucci, David Gondek, Lei Zhang, and Hiroshi Kanayama. 2012 · 2012
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Andrew M. Saxe, James L. McClelland, and Surya Ganguli. 2014 · 2014
Earlier work this paper cites.
Are you talking to a machine? dataset and methods for multilingual image question answering
Haoyuan Gao, Junhua Mao, Jie Zhou, Zhiheng Huang, Lei Wang, and Wei Xu. 2015 · 2015
Earlier work this paper cites.
Multi30K: Multilingual English-German image descriptions
Desmond Elliott, Stella Frank, Khalil Sima’an, and Lucia Specia. 2016 · 2016
Earlier work this paper cites.
Gaussian error linear units (GELUs)
Dan Hendrycks and Kevin Gimpel. 2016 · 2016
Earlier work this paper cites.
Findings of the second shared task on multimodal machine translation and multilingual image description
Desmond Elliott, Stella Frank, Loïc Barrault, Fethi Bougares, and Lucia Specia. 2017 · 2017
Earlier work this paper cites.
Don’t just assume; look and answer: Overcoming priors for visual question answering
Aishwarya Agrawal, Dhruv Batra, Devi Parikh, and Aniruddha Kembhavi. 2018 · 2018
Earlier work this paper cites.
Bottom-up and top-down attention for image captioning and visual question answering
Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang. 2018 · 2018
Earlier work this paper cites.
Findings of the third shared task on multimodal machine translation
Loïc Barrault, Fethi Bougares, Lucia Specia, Chiraag Lala, Desmond Elliott, and Stella Frank. 2018 · 2018
Earlier work this paper cites.
Annotation artifacts in natural language inference data
Suchin Gururangan, Swabha Swayamdipta, Omer Levy, Roy Schwartz, Samuel Bowman, and Noah A. Smith. 2018 · 2018
Earlier work this paper cites.
maskrnn-benchmark: Fast, modular reference implementation of Instance Segmentation and Object Detection algorithms in PyTorch
Francisco Massa and Ross Girshick. 2018 · 2018
Earlier work this paper cites.
Hypothesis only baselines in natural language inference
Adam Poliak, Jason Naradowsky, Aparajita Haldar, Rachel Rudinger, and Benjamin Van Durme. 2018 · 2018
Earlier work this paper cites.
On the limitations of unsupervised bilingual dictionary induction
Anders Søgaard, Sebastian Ruder, and Ivan Vulić. 2018 · 2018
Earlier work this paper cites.
GQA: A new dataset for real-world visual reasoning and compositional question answering
Drew A. Hudson and Christopher D. Manning. 2019 · 2019
Cited alongside, same era.
How multilingual is multilingual BERT?
Telmo Pires, Eva Schlinger, and Dan Garrette. 2019 · 2019
Cited alongside, same era.
LXMERT: Learning cross-modality encoder representations from transformers
Hao Tan and Mohit Bansal. 2019 · 2019
Cited alongside, same era.
The secret is in the spectra: Predicting cross-lingual task performance with spectral similarity measures
Haim Dubossarsky, Ivan Vulić, Roi Reichart, and Anna Korhonen. 2020 · 2020
Cited alongside, same era.
XTREME: A massively multilingual multi-task benchmark for evaluating cross-lingual generalisation
Junjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig, Orhan Firat, and Melvin Johnson. 2020 · 2020
Cited alongside, same era.
Exploiting cloze-questions for few-shot text classification and natural language inference
Timo Schick and Hinrich Schütze. 2021 · 2021
Later among the works it cites.
WIT: Wikipedia-Based Image Text Dataset for Multimodal Multilingual Machine Learning , page 2443–2449. Association for Computing Machinery, New York, NY, USA
Krishna Srinivasan, Karthik Raman, Jiecao Chen, Michael Bendersky, and Marc Najork. 2021 · 2021
Later among the works it cites.
GEM: A general evaluation benchmark for multimodal tasks
Lin Su, Nan Duan, Edward Cui, Lei Ji, Chenfei Wu, Huaishao Luo, Yongfei Liu, Ming Zhong, Taroon Bharti, and Arun Sacheti. 2021 · 2021
Later among the works it cites.
A closer look at few-shot crosslingual transfer: The choice of shots matters
Mengjie Zhao, Yi Zhu, Ehsan Shareghi, Ivan Vulić, Roi Reichart, Anna Korhonen, and Hinrich Schütze. 2021 · 2021
Later among the works it cites.
UC2: Universal cross-lingual cross-modal vision-and-language pre-training
Mingyang Zhou, Luowei Zhou, Shuohang Wang, Yu Cheng, Linjie Li, Zhou Yu, and Jingjing Liu. 2021 · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Anne Lauscher, Vinit Ravishankar, Ivan Vulić, and Goran Glavaš. 2020 · 2020
Cited alongside, same era.
Oscar: Object-semantics aligned pre-training for vision-language tasks
Xiujun Li, Xi Yin, Chunyuan Li, Pengchuan Zhang, Xiaowei Hu, Lei Zhang, Lijuan Wang, Houdong Hu, Li Dong, Furu Wei, Yejin Choi, and Jianfeng Gao. 2020 · 2020
Cited alongside, same era.
AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts
Taylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace, and Sameer Singh. 2020 · 2020
Cited alongside, same era.
A negative case analysis of visual grounding methods for VQA
Robik Shrestha, Kushal Kafle, and Christopher Kanan. 2020 · 2020
Cited alongside, same era.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush. 2020 · 2020
Cited alongside, same era.
Vision-and-language or vision-for-language? on cross-modal influence in multimodal transformers
Stella Frank, Emanuele Bugliarello, and Desmond Elliott. 2021 · 2021
Cited alongside, same era.
MDETR - modulated detection for end-to-end multi-modal understanding
Aishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve, Ishan Misra, and Nicolas Carion. 2021 · 2021
Cited alongside, same era.
Later among the works it cites.
IGLUE: A benchmark for transfer learning across modalities, tasks, and languages
Emanuele Bugliarello, Fangyu Liu, Jonas Pfeiffer, Siva Reddy, Desmond Elliott, Edoardo Maria Ponti, and Ivan Vulic. 2022 · 2022
Closest in time.
Towards multi-lingual visual question answering
Soravit Changpinyo, Linting Xue, Idan Szpektor, Ashish V. Thapliyal, Julien Amelot, Xi Chen, and Radu Soricut. 2022 · 2022
Closest in time.
Retrieve fast, rerank smart: Cooperative and joint approaches for improved cross-modal retrieval
Gregor Geigle, Jonas Pfeiffer, Nils Reimers, Ivan Vulić, and Iryna Gurevych. 2022 · 2022
Closest in time.
Fine-tuning can distort pretrained features and underperform out-of-distribution
Ananya Kumar, Aditi Raghunathan, Robbie Matthew Jones, Tengyu Ma, and Percy Liang. 2022 · 2022
Closest in time.
xGQA: Cross-lingual visual question answering
Jonas Pfeiffer, Gregor Geigle, Aishwarya Kamath, Jan-Martin O. Steitz, Stefan Roth, Ivan Vulić, and Iryna Gurevych. 2022 · 2022
Closest in time.
Don’t stop fine-tuning: On training regimes for few-shot cross-lingual transfer with multilingual language models
Fabian David Schmidt, Ivan Vulić, and Goran Glavaš. 2022 · 2022
Closest in time.
How much can CLIP benefit vision-and-language tasks?
Sheng Shen, Liunian Harold Li, Hao Tan, Mohit Bansal, Anna Rohrbach, Kai-Wei Chang, Zhewei Yao, and Kurt Keutzer. 2022 · 2022
Closest in time.
Multisubs: A large-scale multimodal and multilingual dataset
Josiah Wang, Josiel Figueiredo, and Lucia Specia. 2022 · 2022
Closest in time.
Parameter-efficient tuning makes a good classification head
Zhuoyi Yang, Ming Ding, Yanhui Guo, Qingsong Lv, and Jie Tang. 2022 · 2022
Closest in time.
LiT: Zero-shot transfer with locked-image text tuning
Xiaohua Zhai, Xiao Wang, Basil Mustafa, Andreas Steiner, Daniel Keysers, Alexander Kolesnikov, and Lucas Beyer. 2022 · 2022
Closest in time.