Fetching the paper…
Reading the bibliography…
Large Language Models have found application in various mundane and repetitive tasks including Human Resource (HR) support.
Bertscore: Evaluating text generation with bert
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. 2019 · 1904
Earlier work this paper cites.
A technique for the measurement of attitudes
Rensis Likert. 1932 · 1932
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
When do generative query and document expansions fail? a comprehensive study across methods, retrievers, and datasets
Orion Weller, Kyle Lo, David Wadden, Dawn Lawrie, Benjamin Van Durme, Arman Cohan, and Luca Soldaini. 2024 · 2003
Earlier work this paper cites.
Dense passage retrieval for open-domain question answering
Vladimir Karpukhin, Barlas Oğuz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020a · 2004
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Spearman correlation coefficients, differences between
Leann Myers and Maria J Sirois. 2004 · 2004
Earlier work this paper cites.
Simple bm25 extension to multiple weighted fields
Stephen Robertson, Hugo Zaragoza, and Michael Taylor. 2004 · 2004
Earlier work this paper cites.
Comparing automatic and human evaluation of nlg systems
Anja Belz and Ehud Reiter. 2006 · 2006
Earlier work this paper cites.
The kendall rank correlation coefficient
Hervé Abdi. 2007 · 2007
Earlier work this paper cites.
The probabilistic relevance framework: Bm25 and beyond
Stephen Robertson, Hugo Zaragoza, et al. 2009 · 2009
Earlier work this paper cites.
Cqadupstack: A benchmark data set for community question-answering research
Doris Hoogeveen, Karin Verspoor, and Timothy Baldwin. 2015 · 2015
Cited alongside, same era.
Rasa: Open source language understanding and dialogue management
Tom Bocklisch, Joey Faulkner, Nick Pawlowski, and Alan Nichol. 2017 · 2017
Cited alongside, same era.
Why we need new evaluation metrics for nlg
Jekaterina Novikova, Ondřej Dušek, Amanda Cercas Curry, and Verena Rieser. 2017 · 2017
Cited alongside, same era.
Deep metric learning using triplet network
Elad Hoffer and Nir Ailon. 2018 · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Towards a unified multi-dimensional evaluator for text generation
Ming Zhong, Yang Liu, Da Yin, Yuning Mao, Yizhu Jiao, Pengfei Liu, Chenguang Zhu, Heng Ji, and Jiawei Han. 2022 · 2022
Later among the works it cites.
Benchmarking large language models in retrieval-augmented generation
Jiawei Chen, Hongyu Lin, Xianpei Han, and Le Sun. 2023 · 2023
Later among the works it cites.
Prometheus: Inducing fine-grained evaluation capability in language models
Seungone Kim, Jamin Shin, Yejin Cho, Joel Jang, Shayne Longpre, Hwaran Lee, Sangdoo Yun, Seongjin Shin, Sungdong Kim, James Thorne, et al. 2023 · 2023
Later among the works it cites.
Gpteval: Nlg evaluation using gpt-4 with better human alignment
Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. 2021 · 2021
Cited alongside, same era.
Retrieval augmentation reduces hallucination in conversation
Kurt Shuster, Spencer Poff, Moya Chen, Douwe Kiela, and Jason Weston. 2021 · 2021
Cited alongside, same era.
The first workshop on evaluations and assessments of neural conversation systems
Wei Wei, Bo Dai, Tuo Zhao, Lihong Li, Diyi Yang, Yun-Nung Chen, Y-Lan Boureau, Asli Celikyilmaz, Alborz Geramifard, Aman Ahuja, et al. 2021 · 2021
Cited alongside, same era.
Precise zero-shot dense retrieval without relevance labels
Luyu Gao, Xueguang Ma, Jimmy Lin, and Jamie Callan. 2022 · 2022
Cited alongside, same era.
LongT5: Efficient text-to-text transformer for long sequences
Mandy Guo, Joshua Ainslie, David Uthus, Santiago Ontanon, Jianmo Ni, Yun-Hsuan Sung, and Yinfei Yang. 2022 · 2022
Cited alongside, same era.
Reciprocal rank fusion outperforms condorcet and individual rank learning methods
Gordon V. Cormack, Charles L A Clarke, and Stefan Buettcher. 2009a
Cited in the paper.
Reciprocal rank fusion outperforms condorcet and individual rank learning methods
Gordon V. Cormack, Charles L A Clarke, and Stefan Buettcher. 2009b
Cited in the paper.
Query rewriting in retrieval-augmented large language models
Xinbei Ma, Yeyun Gong, Pengcheng He, hai zhao, and Nan Duan. 2023 · 2023
Later among the works it cites.
Ares: An automated evaluation framework for retrieval-augmented generation systems
Jon Saad-Falcon, Omar Khattab, Christopher Potts, and Matei Zaharia. 2023 · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023 · 2023
Later among the works it cites.
Empower large language model to perform better on industrial domain-specific question answering
Zezhong Wang, Fangkai Yang, Pu Zhao, Lu Wang, Jue Zhang, Mohit Garg, Qingwei Lin, and Dongmei Zhang. 2023 · 2023
Later among the works it cites.
RAGAs: Automated evaluation of retrieval augmented generation
Shahul Es, Jithin James, Luis Espinosa Anke, and Steven Schockaert. 2024 · 2024
Closest in time.
Leveraging large language models for nlg evaluation: A survey
Zhen Li, Xiaohan Xu, Tao Shen, Can Xu, Jia-Chen Gu, and Chongyang Tao. 2024 · 2024
Closest in time.