Fetching the paper…
Reading the bibliography…
This paper presents ICAT, an evaluation framework for measuring coverage of diverse factual information in long-form text generation.
Information retrieval. 2nd. newton, ma
Cornelius Joost Van Rijsbergen. 1979 · 1979
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Meteor: An automatic metric for mt evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie. 2005 · 2005
Earlier work this paper cites.
Novelty and diversity in information retrieval evaluation
Charles L. A. Clarke, Maheedhar Kolla, Gordon V. Cormack, Olga Vechtomova, Azin Ashkan, Stefan Büttcher, and Ian MacKinnon. 2008 · 2008
Earlier work this paper cites.
Overview of the trec 2009 web track
Charles L. A. Clarke, Nick Craswell, and Ian Soboroff. 2009 · 2009
Earlier work this paper cites.
Overview of the trec 2010 web track
Charles L. A. Clarke, Nick Craswell, Ian Soboroff, and Gordon V. Cormack. 2010 · 2010
Earlier work this paper cites.
Overview of the trec 2011 web track
Charles L. A. Clarke, Nick Craswell, Ian Soboroff, and Ellen M. Voorhees. 2011 · 2011
Earlier work this paper cites.
Efficient and effective spam filtering and re-ranking for large web datasets
Gordon V. Cormack, Mark D. Smucker, and Charles L. A. Clarke. 2011 · 2011
Earlier work this paper cites.
Overview of the trec 2012 web track
Charles L. A. Clarke, Nick Craswell, and Ellen M. Voorhees. 2012 · 2012
Earlier work this paper cites.
A diversity-promoting objective function for neural conversation models
Jiwei Li, Michel Galley, Chris Brockett, Jianfeng Gao, and Bill Dolan. 2016 · 2016
Earlier work this paper cites.
FEVER: a large-scale dataset for fact extraction and VERification
James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal. 2018 · 2018
Earlier work this paper cites.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel Bowman. 2018 · 2018
Earlier work this paper cites.
Texygen: A benchmarking platform for text generation models
Yaoming Zhu, Sidi Lu, Lei Zheng, Jiaxian Guo, Weinan Zhang, Jun Wang, and Yong Yu. 2018 · 2018
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Billion-scale similarity search with GPUs
Jeff Johnson, Matthijs Douze, and Hervé Jégou. 2019 · 2019
Cited alongside, same era.
Bertscore: Evaluating text generation with bert
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. 2019 · 2019
Cited alongside, same era.
USR: An unsupervised and reference free evaluation metric for dialog generation
Shikib Mehri and Maxine Eskenazi. 2020 · 2020
Cited alongside, same era.
Adversarial NLI: A new benchmark for natural language understanding
Yixin Nie, Adina Williams, Emily Dinan, Mohit Bansal, Jason Weston, and Douwe Kiela. 2020 · 2020
G-eval: NLG evaluation using gpt-4 with better human alignment
Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023 · 2023
Later among the works it cites.
FActScore: Fine-grained atomic evaluation of factual precision in long form text generation
Sewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis, Wen-tau Yih, Pang Koh, Mohit Iyyer, Luke Zettlemoyer, and Hannaneh Hajishirzi. 2023 · 2023
Later among the works it cites.
Openchat: Advancing open-source language models with mixed-quality data
Guan Wang, Sijie Cheng, Xianyuan Zhan, Xiangang Li, Sen Song, and Yang Liu. 2023 · 2023
Later among the works it cites.
An exam-based evaluation approach beyond traditional relevance judgments
Naghmeh Farzi and Laura Dietz. 2024 · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Bleurt: Learning robust metrics for text generation
Thibault Sellam, Dipanjan Das, and Ankur Parikh. 2020 · 2020
Cited alongside, same era.
Automatic machine translation evaluation in many languages via zero-shot paraphrasing
Brian Thompson and Matt Post. 2020 · 2020
Cited alongside, same era.
Assessing reference-free peer evaluation for machine translation
Sweta Agrawal, George Foster, Markus Freitag, and Colin Cherry. 2021 · 2021
Cited alongside, same era.
Pengcheng He, Jianfeng Gao, and Weizhu Chen. 2021 · 2021
Cited alongside, same era.
Revisiting open domain query facet extraction and generation
Chris Samarinas, Arkin Dharawat, and Hamed Zamani. 2022 · 2022
Cited alongside, same era.
Qlora: Efficient finetuning of quantized llms
Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer. 2023 · 2023
Cited alongside, same era.
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, and Akhil Mathur et al. 2024 · 2024
Later among the works it cites.
Llms-as-judges: A comprehensive survey on llm-based evaluation methods
Haitao Li, Qian Dong, Junjie Chen, Huixue Su, Yujia Zhou, Qingyao Ai, Ziyi Ye, and Yiqun Liu. 2024 · 2024
Later among the works it cites.
Arctic-embed: Scalable, efficient, and accurate text embedding models
Luke Merrick, Danmei Xu, Gaurav Nuti, and Daniel Campos. 2024 · 2024
Later among the works it cites.
Initial nugget evaluation results for the trec 2024 rag track with the autonuggetizer framework
Ronak Pradeep, Nandan Thakur, Shivani Upadhyay, Daniel Campos, Nick Craswell, and Jimmy Lin. 2024 · 2024
Later among the works it cites.
Simulating task-oriented dialogues with state transition graphs and large language models
Chris Samarinas, Pracha Promthaw, Atharva Nijasure, Hansi Zeng, Julian Killingback, and Hamed Zamani. 2024 · 2024
Later among the works it cites.
VeriScore: Evaluating the factuality of verifiable claims in long-form text generation
Yixiao Song, Yekyung Kim, and Mohit Iyyer. 2024 · 2024
Later among the works it cites.
The ClueWeb09 dataset
The Lemur Project. 2009 · 2024
Later among the works it cites.
Towards a unified multi-dimensional evaluator for text generation
Ming Zhong, Yang Liu, Da Yin, Yuning Mao, Yizhu Jiao, Pengfei Liu, Chenguang Zhu, Heng Ji, and Jiawei Han. 2022 · 2038
Closest in time.