Fetching the paper…
Reading the bibliography…
Visual Question Answering with Natural Language Explanation (VQA-NLE) task is challenging due to its high demand for reasoning-based inference.
“Bleu: a method for automatic evaluation of machine translation,”
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu, · 2002
Earlier work this paper cites.
“Rouge: A package for automatic evaluation of summaries,”
Chin-Yew Lin, · 2004
Earlier work this paper cites.
“Meteor: An automatic metric for mt evaluation with improved correlation with human judgments,”
Satanjeev Banerjee and Alon Lavie, · 2005
Earlier work this paper cites.
“A multi-world approach to question answering about real-world scenes based on uncertain input,”
Mateusz Malinowski and Mario Fritz, · 2014
Earlier work this paper cites.
“Cider: Consensus-based image description evaluation,”
Ramakrishna Vedantam, C Lawrence Zitnick, and Devi Parikh, · 2015
Earlier work this paper cites.
“Generating visual explanations,”
Lisa Anne Hendricks, Zeynep Akata, Marcus Rohrbach, Jeff Donahue, Bernt Schiele, and Trevor Darrell, · 2016
Earlier work this paper cites.
“Spice: Semantic propositional image caption evaluation,”
Peter Anderson, Basura Fernando, Mark Johnson, and Stephen Gould, · 2016
Earlier work this paper cites.
“Grad-cam: Visual explanations from deep networks via gradient-based localization,”
Ramprasaath R Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra, · 2017
Earlier work this paper cites.
“Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments,”
Peter Anderson, Qi Wu, Damien Teney, Jake Bruce, Mark Johnson, Niko Sünderhauf, Ian Reid, Stephen Gould, and Anton Van Den Hengel, · 2018
Earlier work this paper cites.
“Bottom-up and top-down attention for image captioning and visual question answering,”
Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang, · 2018
Earlier work this paper cites.
“Visual question answering with memory-augmented networks,”
Chao Ma, Chunhua Shen, Anthony Dick, Qi Wu, Peng Wang, Anton Van den Hengel, and Ian Reid, · 2018
Earlier work this paper cites.
“Multimodal explanations: Justifying decisions and pointing to the evidence,”
Dong Huk Park, Lisa Anne Hendricks, Zeynep Akata, Anna Rohrbach, Bernt Schiele, Trevor Darrell, and Marcus Rohrbach, · 2018
Earlier work this paper cites.
“Language models are unsupervised multitask learners,”
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al., · 2019
Cited alongside, same era.
“Faithful multimodal explanation for visual question answering,”
Jialin Wu and Raymond J Mooney, · 2019
Cited alongside, same era.
“Natural language rationales with full-stack visual reasoning: From pixels to semantic frames to commonsense graphs,”
Ana Marasović, Chandra Bhagavatula, Jae Sung Park, Ronan Le Bras, Noah A Smith, and Yejin Choi, · 2020
Cited alongside, same era.
“Uniter: Universal image-text representation learning,”
Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu, · 2020
Cited alongside, same era.
“Retrieval augmented language model pre-training,”
Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Mingwei Chang, · 2020
Cited alongside, same era.
“Retrieval-augmented generation for knowledge-intensive nlp tasks,”
“Nlx-gpt: A model for natural language explanations in vision and vision-language tasks,”
Fawaz Sammani, Tanmoy Mukherjee, and Nikos Deligiannis, · 2022
Later among the works it cites.
“Improving language models by retrieving from trillions of tokens,”
Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Rutherford, Katie Millican, George Bm Van Den Driessche, Jean-Baptiste Lespiau, Bogdan Damoc, Aidan Clark, et al., · 2022
Later among the works it cites.
Qian Yang, Yunxin Li, Baotian Hu, Lin Ma, Yuxin Ding, and Min Zhang, · 2022
Later among the works it cites.
“Nle-dm: Natural-language explanations for decision making of autonomous driving based on semantic scene understanding,”
Yuchao Feng, Wei Hua, and Yuxiang Sun, · 2023
Later among the works it cites.
“S3c: Semi-supervised vqa natural language explanation via self-critical learning,”
Wei Suo, Mengyang Sun, Weisong Liu, Yiqi Gao, Peng Wang, Yanning Zhang, and Qi Wu, · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al., · 2020
Cited alongside, same era.
“Bertscore: Evaluating text generation with bert,”
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi, · 2020
Cited alongside, same era.
“e-vil: A dataset and benchmark for natural language explanations in vision-language tasks,”
Maxime Kayser, Oana-Maria Camburu, Leonard Salewski, Cornelius Emde, Virginie Do, Zeynep Akata, and Thomas Lukasiewicz, · 2021
Cited alongside, same era.
“Learning transferable visual models from natural language supervision,”
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al., · 2021
Cited alongside, same era.
“Learn to explain: Multimodal reasoning via thought chains for science question answering,”
Pan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu, Kai-Wei Chang, Song-Chun Zhu, Oyvind Tafjord, Peter Clark, and Ashwin Kalyan, · 2022
Cited alongside, same era.
“Retrieval-augmented transformer for image captioning,”
Sara Sarto, Marcella Cornia, Lorenzo Baraldi, and Rita Cucchiara, · 2022
Cited alongside, same era.
Later among the works it cites.
“From wrong to right: A recursive approach towards vision-language explanation,”
Jiaxin Ge, Sanjay Subramanian, Trevor Darrell, and Boyi Li, · 2023
Later among the works it cites.
“Retrieving-to-answer: Zero-shot video question answering with frozen large language models,”
Junting Pan, Ziyi Lin, Yuying Ge, Xiatian Zhu, Renrui Zhang, Yi Wang, Yu Qiao, and Hongsheng Li, · 2023
Later among the works it cites.
“Smallcap: lightweight image captioning prompted with retrieval augmentation,”
Rita Ramos, Bruno Martins, Desmond Elliott, and Yova Kementchedjhieva, · 2023
Later among the works it cites.
“Towards a unified model for generating answers and explanations in visual question answering,”
Chenxi Whitehouse, Tillman Weyde, and Pranava Madhyastha, · 2023
Later among the works it cites.
“Towards more faithful natural language explanation using multi-level contrastive learning in vqa,”
Chengen Lai, Shengli Song, Shiqi Meng, Jingyang Li, Sitong Yan, and Guangneng Hu, · 2024
Closest in time.
“Do you remember? dense video captioning with cross-modal memory retrieval,”
Minkuk Kim, Hyeon Bae Kim, Jinyoung Moon, Jinwoo Choi, and Seong Tae Kim, · 2024
Closest in time.