Fetching the paper…
Reading the bibliography…
8 years after the visual question answering (VQA) task was proposed, accuracy remains the primary metric for automatic evaluation.
Language models are few-shot learners
Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J. D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. 2020 · 1901
Earlier work this paper cites.
On Faithfulness and Factuality in Abstractive Summarization
Maynez, J.; Narayan, S.; Bohnet, B.; and McDonald, R. 2020 · 1919
Earlier work this paper cites.
METEOR: An automatic metric for MT evaluation with improved correlation with human judgments
Banerjee, S.; and Lavie, A. 2005 · 2005
Earlier work this paper cites.
Vqa: Visual question answering
Antol, S.; Agrawal, A.; Lu, J.; Mitchell, M.; Batra, D.; Zitnick, C. L.; and Parikh, D. 2015 · 2015
Earlier work this paper cites.
Cider: Consensus-based image description evaluation
Vedantam, R.; Lawrence Zitnick, C.; and Parikh, D. 2015 · 2015
Earlier work this paper cites.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Goyal, Y.; Khot, T.; Summers-Stay, D.; Batra, D.; and Parikh, D. 2017 · 2017
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Krishna, R.; Zhu, Y.; Groth, O.; Johnson, J.; Hata, K.; Kravitz, J.; Chen, S.; Kalantidis, Y.; Li, L.-J.; Shamma, D. A.; et al. 2017 · 2017
Earlier work this paper cites.
Evaluating question answering evaluation
Chen, A.; Stanovsky, G.; Singh, S.; and Gardner, M. 2019 · 2019
Earlier work this paper cites.
Ok-vqa: A visual question answering benchmark requiring external knowledge
Marino, K.; Rastegari, M.; Farhadi, A.; and Mottaghi, R. 2019 · 2019
Earlier work this paper cites.
Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
Reimers, N.; and Gurevych, I. 2019 · 2019
Earlier work this paper cites.
Transformers: State-of-the-art natural language processing
Wolf, T.; Debut, L.; Sanh, V.; Chaumond, J.; Delangue, C.; Moi, A.; Cistac, P.; Rault, T.; Louf, R.; Funtowicz, M.; et al. 2020 · 2020
Earlier work this paper cites.
BERTScore: Evaluating Text Generation with BERT
Zhang, T.; Kishore, V.; Wu, F.; Weinberger, K. Q.; and Artzi, Y. 2020 · 2020
Cited alongside, same era.
KPQA: A Metric for Generative Question Answering Using Keyphrase Weights
Lee, H.; Yoon, S.; Dernoncourt, F.; Kim, D. S.; Bui, T.; Shin, J.; and Jung, K. 2021 · 2021
Cited alongside, same era.
‘Just because you are right, doesn’t mean I am wrong’: Overcoming a bottleneck in development and evaluation of Open-Ended VQA tasks
Luo, M.; Sampat, S. K.; Tallman, R.; Zeng, Y.; Vancha, M.; Sajja, A.; and Baral, C. 2021 · 2021
Cited alongside, same era.
Semantic Answer Similarity for Evaluating Question Answering Models
Risch, J.; Möller, T.; Gutsch, J.; and Pietsch, M. 2021 · 2021
Cited alongside, same era.
What’s in a Name? Answer Equivalence For Open-Domain Question Answering
Si, C.; Zhao, C.; and Boyd-Graber, J. 2021 · 2021
Cited alongside, same era.
Opt: Open pre-trained transformer language models
Zhang, S.; Roller, S.; Goyal, N.; Artetxe, M.; Chen, M.; Chen, S.; Dewan, C.; Diab, M.; Li, X.; Lin, X. V.; et al. 2022 · 2022
Later among the works it cites.
Reassessing Evaluation Practices in Visual Question Answering: A Case Study on Out-of-Distribution Generalization
Agrawal, A.; Kajic, I.; Bugliarello, E.; Davoodi, E.; Gergely, A.; Blunsom, P.; and Nematzadeh, A. 2023 · 2023
Closest in time.
Introducing Claude
Anthropic. 2023 · 2023
Closest in time.
Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90%* ChatGPT Quality
Chiang, W.-L.; Li, Z.; Lin, Z.; Sheng, Y.; Wu, Z.; Zhang, H.; Zheng, L.; Zhuang, S.; Zhuang, Y.; Gonzalez, J. E.; Stoica, I.; and Xing, E. P. 2023 · 2023
Closest in time.
Gptscore: Evaluate as you desire
Fu, J.; Ng, S.-K.; Jiang, Z.; and Liu, P. 2023 · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tomayto, Tomahto. Beyond Token-level Answer Equivalence for Question Answering Evaluation
Bulian, J.; Buck, C.; Gajewski, W.; Boerschinger, B.; and Schuster, T. 2022 · 2022
Cited alongside, same era.
Scaling instruction-finetuned language models
Chung, H. W.; Hou, L.; Longpre, S.; Zoph, B.; Tay, Y.; Fedus, W.; Li, E.; Wang, X.; Dehghani, M.; Brahma, S.; et al. 2022 · 2022
Cited alongside, same era.
PromptCap: Prompt-Guided Task-Aware Image Captioning
Hu, Y.; Hua, H.; Yang, Z.; Shi, W.; Smith, N. A.; and Luo, J. 2022 · 2022
Cited alongside, same era.
Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Li, J.; Li, D.; Xiong, C.; and Hoi, S. 2022 · 2022
Cited alongside, same era.
Introducing ChatGPT
OpenAI. 2022 · 2022
Cited alongside, same era.
Super-naturalinstructions: Generalization via declarative instructions on 1600+ nlp tasks
Wang, Y.; Mishra, S.; Alipoormolabashi, P.; Kordi, Y.; Mirzaei, A.; Naik, A.; Ashok, A.; Dhanasekaran, A. S.; Arunkumar, A.; Stap, D.; et al. 2022 · 2022
Cited alongside, same era.
Mimic-it: Multi-modal in-context instruction tuning
Li, B.; Zhang, Y.; Chen, L.; Wang, J.; Pu, F.; Yang, J.; Li, C.; and Liu, Z. 2023a
Cited in the paper.
Evaluating Open-Domain Question Answering in the Era of Large Language Models
Kamalloo, E.; Dziri, N.; Clarke, C. L.; and Rafiei, D. 2023 · 2023
Closest in time.
Gpteval: Nlg evaluation using gpt-4 with better human alignment
Liu, Y.; Iter, D.; Xu, Y.; Wang, S.; Xu, R.; and Zhu, C. 2023 · 2023
Closest in time.
GPT-4 technical report
OpenAI, R. 2023 · 2023
Closest in time.
Can foundation models label data like humans?
Rajani, N.; Lambert, N.; Han, S.; Wang, J.; Nitski, O.; Beeching, E.; and Tunstall, L. 2023 · 2023
Closest in time.
Judging LLM-as-a-judge with MT-Bench and Chatbot Arena
Zheng, L.; Chiang, W.-L.; Sheng, Y.; Zhuang, S.; Wu, Z.; Zhuang, Y.; Lin, Z.; Li, Z.; Li, D.; Xing, E. P.; Zhang, H.; Gonzalez, J. E.; and Stoica, I. 2023 · 2023
Closest in time.
LIMA: Less Is More for Alignment
Zhou, C.; Liu, P.; Xu, P.; Iyer, S.; Sun, J.; Mao, Y.; Ma, X.; Efrat, A.; Yu, P.; Yu, L.; Zhang, S.; Ghosh, G.; Lewis, M.; Zettlemoyer, L.; and Levy, O. 2023 · 2023
Closest in time.