Fetching the paper…
Reading the bibliography…
Numerous benchmarks have been established to assess the performance of foundation models on open-ended question answering, which serves as a comprehensive test of a model's ability to understand and generate language in a manner similar to humans.
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu, “Bleu: a method for automatic evaluation of machine translation,” in Proceedings of the 40th annual meeting of the Association for Computational Linguistics
2002
Earlier work this paper cites.
C.-Y. Lin, “Rouge: A package for automatic evaluation of summaries,” in Text summarization branches out
2004
Earlier work this paper cites.
S. Banerjee and A. Lavie, “Meteor: An automatic metric for mt evaluation with improved correlation with human judgments,” in Proceedings of the acl workshop on intrinsic and extrinsic evaluation measures for machine translation and/or summarization
2005
Earlier work this paper cites.
J. Berant, A. Chou, R. Frostig, and P. Liang, “Semantic parsing on freebase from question-answer pairs,” in Proceedings of the 2013 conference on empirical methods in natural language processing
2013
Earlier work this paper cites.
T. Nguyen, M. Rosenberg, X. Song, J. Gao, S. Tiwary, R. Majumder, and L. Deng, “Ms marco: A human generated machine reading comprehension dataset,” choice
2016
Earlier work this paper cites.
P. Rajpurkar, J. Zhang, K. Lopyrev, and P. Liang, “Squad: 100,000+ questions for machine comprehension of text,” in Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing
2016
Earlier work this paper cites.
P. Rajpurkar, R. Jia, and P. Liang, “Know what you don’t know: Unanswerable questions for squad,” in Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers)
2018
Earlier work this paper cites.
T. Mihaylov, P. Clark, T. Khot, and A. Sabharwal, “Can a suit of armor conduct electricity? a new dataset for open book question answering,” in Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing
2018
Earlier work this paper cites.
E. Choi, H. He, M. Iyyer, M. Yatskar, W.-t. Yih, Y. Choi, P. Liang, and L. Zettlemoyer, “Quac: Question answering in context,” in Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing
2018
Earlier work this paper cites.
A. Wang, Y. Pruksachatkun, N. Nangia, A. Singh, J. Michael, F. Hill, O. Levy, and S. Bowman, “Superglue: A stickier benchmark for general-purpose language understanding systems,” Advances in neural information processing systems
2019
Earlier work this paper cites.
A. Chen, G. Stanovsky, S. Singh, and M. Gardner, “Evaluating question answering evaluation,” in Proceedings of the 2nd workshop on machine reading for question answering
2019
Earlier work this paper cites.
T. Kwiatkowski, J. Palomaki, O. Redfield, M. Collins, A. Parikh, C. Alberti, D. Epstein, I. Polosukhin, J. Devlin, K. Lee, et al
2019
Earlier work this paper cites.
A. Fan, Y. Jernite, E. Perez, D. Grangier, J. Weston, and M. Auli, “Eli5: Long form question answering,” in Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics
2019
Earlier work this paper cites.
W. Zhao, M. Peyrard, F. Liu, Y. Gao, C. M. Meyer, and S. Eger, “Moverscore: Text generation evaluating with contextualized embeddings and earth mover distance,” in Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing (EMNLP)
2019
Earlier work this paper cites.
S. Reddy, D. Chen, and C. D. Manning, “Coqa: A conversational question answering challenge,” Transactions of the Association for Computational Linguistics
2019
Earlier work this paper cites.
T. Zhang, V. Kishore, F. Wu, K. Q. Weinberger, and Y. Artzi, “Bertscore: Evaluating text generation with bert,” in International Conference on Learning Representations
2020
Earlier work this paper cites.
A. Chen, G. Stanovsky, S. Singh, and M. Gardner, “Mocha: A dataset for training and evaluating generative reading comprehension metrics,” in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)
2020
Earlier work this paper cites.
T. Sellam, D. Das, and A. Parikh, “Bleurt: Learning robust metrics for text generation,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics
2020
Earlier work this paper cites.
R. Rei, C. Stewart, A. C. Farinha, and A. Lavie, “Comet: A neural framework for mt evaluation,” in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)
2020
Earlier work this paper cites.
longman, 2020
B. S. Bloom and D. R. Krathwohl, Taxonomy of educational objectives: The classification of educational goals. Book 1, Cognitive domain · 2020
Earlier work this paper cites.
J. A. Campos, A. Otegi, A. Soroa, J. M. Deriu, M. Cieliebak, and E. Agirre, “Doqa-accessing domain-specific faqs via conversational qa,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics
2020
Earlier work this paper cites.
2020
Cited alongside, same era.
2021
Cited alongside, same era.
D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt, “Measuring massive multitask language understanding,” in International Conference on Learning Representations
2021
Cited alongside, same era.
K. Krishna, A. Roy, and M. Iyyer, “Hurdles to progress in long-form question answering,” in Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies
2021
Cited alongside, same era.
J. Yu, X. Wang, S. Tu, S. Cao, D. Zhang-Li, X. Lv, H. Peng, Z. Yao, X. Zhang, H. Li, et al
2023
Closest in time.
2023
Closest in time.
A. Rogers, M. Gardner, and I. Augenstein, “Qa dataset explosion: A taxonomy of nlp resources for question answering and reading comprehension,” ACM Computing Surveys
2023
Closest in time.
2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
C. Si, C. Zhao, and J. Boyd-Graber, “What’s in a name? answer equivalence for open-domain question answering,” in Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing
2021
Cited alongside, same era.
W. Yuan, G. Neubig, and P. Liu, “Bartscore: Evaluating generated text as text generation,” Advances in Neural Information Processing Systems
2021
Cited alongside, same era.
A. R. Fabbri, W. Kryściński, B. McCann, C. Xiong, R. Socher, and D. Radev, “Summeval: Re-evaluating summarization evaluation,” Transactions of the Association for Computational Linguistics
2021
Cited alongside, same era.
OpenAI, “Introducing chatgpt,” 2022
2022
Cited alongside, same era.
2022
Cited alongside, same era.
2022
Cited alongside, same era.
A. B. Sai, A. K. Mohankumar, and M. M. Khapra, “A survey of evaluation metrics used for nlg systems,” ACM Computing Surveys (CSUR)
2022
Cited alongside, same era.
2022
Cited alongside, same era.
2023
Closest in time.
2023
Closest in time.
OpenAI, “Openai: Gpt-4,” 2023
2023
Closest in time.
W.-L. Chiang, Z. Li, Z. Lin, Y. Sheng, Z. Wu, H. Zhang, L. Zheng, S. Zhuang, Y. Zhuang, J. E. Gonzalez, I. Stoica, and E. P. Xing, “Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,” March 2023
2023
Closest in time.
2023
Closest in time.
W. Yidong, Y. Zhuohao, Z. Zhengran, Y. Linyi, H. Qiang, W. Cunxiang, C. Hao, J. Chaoya, X. Rui, W. Jindong, X. Xing, Y. Wei, Z. Shikun, and Z. Yue, “Pandalm: Reproducible and automated language model assessment.” https://github.com/WeOpenML/PandaLM , 2023
2023
Closest in time.
2023
Closest in time.
H. Chen, D. M. Vo, H. Takamura, Y. Miyao, and H. Nakayama, “Storyer: Automatic story evaluation via ranking, rating and reasoning,” Journal of Natural Language Processing
2023
Closest in time.
OpenAI, “Gpt-4 technical report,” arXiv preprint arXiv:2303.08774
2023
Closest in time.
2023
Closest in time.
Anthropic, “Anthropic: Claude,” 2023
2023
Closest in time.
Google, “Google: Bard,” 2023
2023
Closest in time.
A. Zeng, X. Liu, Z. Du, Z. Wang, H. Lai, M. Ding, Z. Yang, Y. Xu, W. Zheng, X. Xia, W. L. Tam, Z. Ma, Y. Xue, J. Zhai, W. Chen, Z. Liu, P. Zhang, Y. Dong, and J. Tang, “GLM-130b: An open bilingual pre-trained model,” in The Eleventh International Conference on Learning Representations
2023
Closest in time.
LMSYS, “Lmsys org: Chatbot arena,” 2023
2023
Closest in time.
R. Taori, I. Gulrajani, T. Zhang, Y. Dubois, X. Li, C. Guestrin, P. Liang, and T. B. Hashimoto, “Stanford alpaca: An instruction-following llama model.” https://github.com/tatsu-lab/stanford_alpaca , 2023
2023
Closest in time.