Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have demonstrated impressive capabilities across various domains, prompting a surge in their practical applications.
A coefficient of agreement for nominal scales
Jacob Cohen · 1960
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin · 2004
Earlier work this paper cites.
Meteor: An automatic metric for mt evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie · 2005
Earlier work this paper cites.
The probabilistic relevance framework: Bm25 and beyond
Stephen Robertson, Hugo Zaragoza, et al · 2009
Earlier work this paper cites.
Semantic parsing on freebase from question-answer pairs
Jonathan Berant, Andrew Chou, Roy Frostig, and Percy Liang · 2013
Earlier work this paper cites.
Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension
Mandar Joshi, Eunsol Choi, Daniel S Weld, and Luke Zettlemoyer · 2017
Earlier work this paper cites.
Hotpotqa: A dataset for diverse, explainable multi-hop question answering
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W Cohen, Ruslan Salakhutdinov, and Christopher D Manning · 2018
Earlier work this paper cites.
Natural questions: a benchmark for question answering research
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, et al · 2019
Earlier work this paper cites.
Are red roses red? evaluating consistency of question-answering models
Marco Tulio Ribeiro, Carlos Guestrin, and Sameer Singh · 2019
Earlier work this paper cites.
Bertscore: Evaluating text generation with bert
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi · 2019
Earlier work this paper cites.
Logic-guided data augmentation and regularization for consistent question answering
Akari Asai and Hannaneh Hajishirzi · 2020
Earlier work this paper cites.
Dense passage retrieval for open-domain question answering
Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih · 2020
Cited alongside, same era.
How much knowledge can you pack into the parameters of a language model?
Adam Roberts, Colin Raffel, and Noam Shazeer · 2020
Cited alongside, same era.
Bartscore: Evaluating generated text as text generation
Weizhe Yuan, Graham Neubig, and Pengfei Liu · 2021
Cited alongside, same era.
Scaling instruction-finetuned language models
Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al · 2022
Cited alongside, same era.
Evaluating factuality in text simplification
Ashwin Devaraj, William Sheffield, Byron C Wallace, and Junyi Jessy Li · 2022
Cited alongside, same era.
Lm vs lm: Detecting factual errors via cross examination
Roi Cohen, May Hamri, Mor Geva, and Amir Globerson · 2023
Later among the works it cites.
Gptscore: Evaluate as you desire
Jinlan Fu, See-Kiong Ng, Zhengbao Jiang, and Pengfei Liu · 2023
Later among the works it cites.
Improving Sequential Model Editing with Fact Retrieval
Xiaoqi Han, Ru Li, Hongye Tan, Wang Yuanlong, Qinghua Chai, and Jeff Z. Pan · 2023
Later among the works it cites.
Multi-view Contrastive Learning for Entity Typing over Knowledge Graphs
Zhiwei Hu, Victor Basulto, Zhiliang Xiang, Ru Li, and Jeff Z. Pan · 2023
Later among the works it cites.
Challenges and applications of large language models
Jean Kaddour, Joshua Harris, Maximilian Mozes, Herbie Bradley, Roberta Raileanu, and Robert McHardy · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Transformer-based Entity Typing in Knowledge Graphs
Zhiwei Hu, Victor Basulto, Zhiliang Xiang, Ru Li, and Jeff Z. Pan · 2022
Cited alongside, same era.
Becel: Benchmark for consistency evaluation of language models
Myeongjun Jang, Deuk Sin Kwon, and Thomas Lukasiewicz · 2022
Cited alongside, same era.
Alex Mallen, Akari Asai, Victor Zhong, Rajarshi Das, Hannaneh Hajishirzi, and Daniel Khashabi · 2022
Cited alongside, same era.
Evaluating the factual consistency of large language models through summarization
Derek Tam, Anisha Mascarenhas, Shiyue Zhang, Sarah Kwan, Mohit Bansal, and Colin Raffel · 2022
Cited alongside, same era.
Multilingual summarization with factual consistency evaluation
Roee Aharoni, Shashi Narayan, Joshua Maynez, Jonathan Herzig, Elizabeth Clark, and Mirella Lapata · 2023
Cited alongside, same era.
Sparks of artificial general intelligence: Early experiments with gpt-4
Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, et al · 2023
Cited alongside, same era.
Menli: Robust evaluation metrics from natural language inference
Yanran Chen and Steffen Eger · 2023
Cited alongside, same era.
Later among the works it cites.
Trustworthy llms: a survey and guideline for evaluating large language models’ alignment
Yang Liu, Yuanshun Yao, Jean-Francois Ton, Xiaoying Zhang, Ruocheng Guo Hao Cheng, Yegor Klochkov, Muhammad Faaiz Taufiq, and Hang Li · 2023
Later among the works it cites.
Large Language Models and Knowledge Graphs: Opportunities and Challenges
Jeff Z. Pan, Simon Razniewski, Jan-Christoph Kalo, Sneha Singhania, Jiaoyan Chen, Stefan Dietze, Hajira Jabeen, Janna Omeliyanenko, Wen Zhang, Matteo Lissandrini, Russa Biswas, Gerard de Melo, Angela Bonifati, Edlira Vakaj, Mauro Dragoni, , and Damien Graux · 2023
Later among the works it cites.
Is chatgpt a general-purpose natural language processing task solver?
Chengwei Qin, Aston Zhang, Zhuosheng Zhang, Jiaao Chen, Michihiro Yasunaga, and Diyi Yang · 2023
Later among the works it cites.
Context generation improves open domain question answering
Dan Su, Mostofa Patwary, Shrimai Prabhumoye, Peng Xu, Ryan Prenger, Mohammad Shoeybi, Pascale Fung, Animashree Anandkumar, and Bryan Catanzaro · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Later among the works it cites.
Korc: Knowledge oriented reading comprehension benchmark for deep text understanding
Zijun Yao, Yantao Liu, Xin Lv, Shulin Cao, Jifan Yu, Juanzi Li, and Lei Hou · 2023
Later among the works it cites.