Fetching the paper…
Reading the bibliography…
As large language models (LLMs) advance, it becomes more challenging to reliably evaluate their output due to the high costs of human evaluation.
SemEval-2017 task 1: Semantic textual similarity multilingual and crosslingual focused evaluation
D. Cer, M. Diab, E. Agirre, I. Lopez-Gazpio, and L. Specia · 2001
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu · 2002
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
C.-Y. Lin · 2004
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. Kingma and J. Ba · 2015
Earlier work this paper cites.
First Quora Dataset release: Question pairs, 2017
S. Iyer, N. Dandekar, and K. Csernai · 2017
Earlier work this paper cites.
e-snli: Natural language inference with natural language explanations
O.-M. Camburu, T. Rocktäschel, T. Lukasiewicz, and P. Blunsom · 2018
Earlier work this paper cites.
A broad-coverage challenge corpus for sentence understanding through inference
A. Williams, N. Nangia, and S. Bowman · 2018
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2019
Earlier work this paper cites.
GLTR: Statistical detection and visualization of generated text
S. Gehrmann, H. Strobelt, and A. Rush · 2019
Earlier work this paper cites.
Neural network acceptability judgments
A. Warstadt, A. Singh, and S. R. Bowman · 2019
Earlier work this paper cites.
PAWS: Paraphrase adversaries from word scrambling
Y. Zhang, J. Baldridge, and L. He · 2019
Earlier work this paper cites.
MoverScore: Text generation evaluating with contextualized embeddings and earth mover distance
W. Zhao, M. Peyrard, F. Liu, Y. Gao, C. M. Meyer, and S. Eger · 2019
Earlier work this paper cites.
MOCHA: A dataset for training and evaluating generative reading comprehension metrics
A. Chen, G. Stanovsky, S. Singh, and M. Gardner · 2020
Earlier work this paper cites.
FEQA: A question answering evaluation framework for faithfulness assessment in abstractive summarization
E. Durmus, H. He, and M. Diab · 2020
Earlier work this paper cites.
On faithfulness and factuality in abstractive summarization
J. Maynez, S. Narayan, B. Bohnet, and R. McDonald · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu · 2020
Earlier work this paper cites.
COMET: A neural framework for MT evaluation
R. Rei, C. Stewart, A. C. Farinha, and A. Lavie · 2020
Earlier work this paper cites.
BLEURT: Learning robust metrics for text generation
T. Sellam, D. Das, and A. Parikh · 2020
Earlier work this paper cites.
Learning to summarize with human feedback
N. Stiennon, L. Ouyang, J. Wu, D. Ziegler, R. Lowe, C. Voss, A. Radford, D. Amodei, and P. F. Christiano · 2020
Earlier work this paper cites.
Automatic machine translation evaluation in many languages via zero-shot paraphrasing
B. Thompson and M. Post · 2020
Earlier work this paper cites.
Asking and answering questions to evaluate the factual consistency of summaries
A. Wang, K. Cho, and M. Lewis · 2020
Earlier work this paper cites.
Bertscore: Evaluating text generation with bert
T. Zhang, V. Kishore, F. Wu, K. Q. Weinberger, and Y. Artzi · 2020
Earlier work this paper cites.
A general language assistant as a laboratory for alignment
A. Askell, Y. Bai, A. Chen, D. Drain, D. Ganguli, T. Henighan, A. Jones, N. Joseph, B. Mann, N. DasSarma, et al · 2021
Earlier work this paper cites.
Evaluating large language models trained on code
M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. D. O. Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman, et al · 2021
Earlier work this paper cites.
Towards question-answering as an automatic metric for evaluating the content quality of a summary
D. Deutsch, T. Bedrax-Weiss, and D. Roth · 2021
Earlier work this paper cites.
SummEval: Re-evaluating summarization evaluation
A. R. Fabbri, W. Kryściński, B. McCann, C. Xiong, R. Socher, and D. Radev · 2021
Earlier work this paper cites.
Experts, errors, and context: A large-scale study of human evaluation for machine translation
M. Freitag, G. Foster, D. Grangier, V. Ratnakar, Q. Tan, and W. Macherey · 2021
Earlier work this paper cites.
Annotating and modeling fine-grained factuality in summarization
T. Goyal and G. Durrett · 2021
Earlier work this paper cites.
q 2 q^{2} : Evaluating factual consistency in knowledge-grounded dialogues via question generation and question answering
O. Honovich, L. Choshen, R. Aharoni, E. Neeman, I. Szpektor, and O. Abend · 2021
Earlier work this paper cites.
The perils of using Mechanical Turk to evaluate open-ended text generation
M. Karpinska, N. Akoury, and M. Iyyer · 2021
Earlier work this paper cites.
Hurdles to progress in long-form question answering
K. Krishna, A. Roy, and M. Iyyer · 2021
Earlier work this paper cites.
Datasets: A community library for natural language processing
Q. Lhoest, A. Villanova del Moral, Y. Jernite, A. Thakur, P. von Platen, S. Patil, J. Chaumond, M. Drame, J. Plu, L. Tunstall, J. Davison, M. Šaško, G. Chhablani, B. Malik, S. Brandeis, T. Le Scao, V. Sanh, C. Xu, N. Patry, A. McMillan-Major, P. Schmid, S. Gugger, C. Delangue, T. Matussière, L. Debut, S. Bekman, P. Cistac, T. Goehringer, V. Mustar, F. Lagunas, A. Rush, and T. Wolf · 2021
Earlier work this paper cites.
Webgpt: Browser-assisted question-answering with human feedback
R. Nakano, J. Hilton, S. Balaji, J. Wu, L. Ouyang, C. Kim, C. Hesse, S. Jain, V. Kosaraju, W. Saunders, et al · 2021
Earlier work this paper cites.
Understanding factuality in abstractive summarization with FRANK: A benchmark for factuality metrics
A. Pagnoni, V. Balachandran, and Y. Tsvetkov · 2021
Earlier work this paper cites.
Crisscrossed captions: Extended intramodal and intermodal semantic similarity judgments for MS-COCO
Z. Parekh, J. Baldridge, D. Cer, A. Waters, and Y. Yang · 2021
Earlier work this paper cites.
Mauve: Measuring the gap between neural text and human text using divergence frontiers
K. Pillutla, S. Swayamdipta, R. Zellers, J. Thickstun, S. Welleck, Y. Choi, and Z. Harchaoui · 2021
Earlier work this paper cites.
Get your vitamin C! robust fact verification with contrastive evidence
T. Schuster, A. Fisch, and R. Barzilay · 2021
Earlier work this paper cites.
Bartscore: Evaluating generated text as text generation
W. Yuan, G. Neubig, and P. Liu · 2021
Earlier work this paper cites.
Help me write a poem - instruction tuning as a vehicle for collaborative poetry writing
T. Chakrabarty, V. Padmakumar, and H. He · 2022
Earlier work this paper cites.
PaLM: Scaling language modeling with pathways
A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann, et al · 2022
Cited alongside, same era.
Is GPT-3 text indistinguishable from human text? scarecrow: A framework for scrutinizing machine text
Y. Dou, M. Forbes, R. Koncel-Kedziorski, N. A. Smith, and Y. Choi · 2022
Cited alongside, same era.
Improving large-scale paraphrase acquisition and generation
Y. Dou, C. Jiang, and W. Xu · 2022
Cited alongside, same era.
FaithDial: A faithful benchmark for information-seeking dialogue
N. Dziri, E. Kamalloo, S. Milton, O. Zaiane, M. Yu, E. M. Ponti, and S. Reddy · 2022
Cited alongside, same era.
Evaluating attribution in dialogue systems: The BEGIN benchmark
N. Dziri, H. Rashkin, T. Linzen, and D. Reitter · 2022
Cited alongside, same era.
HaluEval: A large-scale hallucination evaluation benchmark for large language models
J. Li, X. Cheng, X. Zhao, J.-Y. Nie, and J.-R. Wen · 2023
Later among the works it cites.
G-eval: NLG evaluation using gpt-4 with better human alignment
Y. Liu, D. Iter, Y. Xu, S. Wang, R. Xu, and C. Zhu · 2023
Later among the works it cites.
The flan collection: Designing data and methods for effective instruction tuning
S. Longpre, L. Hou, T. Vu, A. Webson, H. W. Chung, Y. Tay, D. Zhou, Q. V. Le, B. Zoph, J. Wei, and A. Roberts · 2023
Later among the works it cites.
LENS: A learnable evaluation metric for text simplification
M. Maddela, Y. Dou, D. Heineman, and W. Xu · 2023
Later among the works it cites.
SelfCheckGPT: Zero-resource black-box hallucination detection for generative large language models
P. Manakul, A. Liusie, and M. Gales · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Understanding dataset difficulty with 𝒱 \mathcal{V} -usable information
K. Ethayarajh, Y. Choi, and S. Swayamdipta · 2022
Cited alongside, same era.
Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned
D. Ganguli, L. Lovitt, J. Kernion, A. Askell, Y. Bai, S. Kadavath, B. Mann, E. Perez, N. Schiefer, K. Ndousse, et al · 2022
Cited alongside, same era.
News summarization and evaluation in the era of gpt-3
T. Goyal, J. J. Li, and G. Durrett · 2022
Cited alongside, same era.
DialFact: A benchmark for fact-checking in dialogue
P. Gupta, C.-S. Wu, W. Liu, and C. Xiong · 2022
Cited alongside, same era.
GENIE: Toward reproducible and standardized human evaluation for text generation
D. Khashabi, G. Stanovsky, J. Bragg, N. Lourie, J. Kasai, Y. Choi, N. A. Smith, and D. Weld · 2022
Cited alongside, same era.
RankGen: Improving text generation with large ranking models
K. Krishna, Y. Chang, J. Wieting, and M. Iyyer · 2022
Cited alongside, same era.
Using interactive feedback to improve the accuracy and explainability of question answering systems post-deployment
Z. Li, P. Sharma, X. H. Lu, J. Cheung, and S. Reddy · 2022
Cited alongside, same era.
S. Min, K. Krishna, X. Lyu, M. Lewis, W.-t. Yih, P. Koh, M. Iyyer, L. Zettlemoyer, and H. Hajishirzi · 2023
Later among the works it cites.
Coffee: Boost your code llms by fixing bugs with feedback
S. Moon, Y. Song, H. Chae, D. Kang, T. Kwon, K. T.-i. Ong, S.-w. Hwang, and J. Yeo · 2023
Later among the works it cites.
Octopack: Instruction tuning code large language models
N. Muennighoff, Q. Liu, A. Zebaze, Q. Zheng, B. Hui, T. Y. Zhuo, S. Singh, X. Tang, L. Von Werra, and S. Longpre · 2023
Later among the works it cites.
Codegen: An open large language model for code with multi-turn program synthesis
E. Nijkamp, B. Pang, H. Hayashi, L. Tu, H. Wang, Y. Zhou, S. Savarese, and C. Xiong · 2023
Later among the works it cites.
Scaling up models and data with t5x and seqio
A. Roberts, H. W. Chung, G. Mishra, A. Levskaya, J. Bradbury, D. Andor, S. Narang, B. Lester, C. Gaffney, A. Mohiuddin, C. Hawthorne, A. Lewkowycz, A. Salcianu, M. van Zee, J. Austin, S. Goodman, L. B. Soares, H. Hu, S. Tsvyashchenko, A. Chowdhery, J. Bastings, J. Bulian, X. Garcia, J. Ni, A. Chen, K. Kenealy, K. Han, M. Casbon, J. H. Clark, S. Lee, D. Garrette, J. Lee-Thorp, C. Raffel, N. Shazeer, M. Ritter, M. Bosma, A. Passos, J. Maitin-Shepard, N. Fiedel, M. Omernick, B. Saeta, R. Sepassi, A. Spiridonov, J. Newlan, and A. Gesmundo · 2023
Later among the works it cites.
Towards better evaluation of instruction-following: A case-study in summarization
O. Skopek, R. Aralikatte, S. Gooding, and V. Carbune · 2023
Later among the works it cites.
Freshllms: Refreshing large language models with search engine augmentation
T. Vu, M. Iyyer, X. Wang, N. Constant, J. Wei, J. Wei, C. Tar, Y.-H. Sung, D. Zhou, Q. Le, et al · 2023
Later among the works it cites.
Is ChatGPT a good NLG evaluator? a preliminary study
J. Wang, Y. Liang, F. Meng, Z. Sun, H. Shi, Z. Li, J. Xu, J. Qu, and J. Zhou · 2023
Later among the works it cites.
Fine-grained human feedback gives better rewards for language model training
Z. Wu, Y. Hu, W. Shi, N. Dziri, A. Suhr, P. Ammanabrolu, N. A. Smith, M. Ostendorf, and H. Hajishirzi · 2023
Later among the works it cites.
A critical evaluation of evaluations for long-form question answering
F. Xu, Y. Song, M. Iyyer, and E. Choi · 2023
Later among the works it cites.
INSTRUCTSCORE: Towards explainable text generation evaluation with automatic feedback
W. Xu, D. Wang, L. Pan, Z. Song, M. Freitag, W. Wang, and L. Li · 2023
Later among the works it cites.
CREPE: Open-domain question answering with false presuppositions
X. Yu, S. Min, L. Zettlemoyer, and H. Hajishirzi · 2023
Later among the works it cites.
Slic-hf: Sequence likelihood calibration with human feedback
Y. Zhao, R. Joshi, T. Liu, M. Khalman, M. Saleh, and P. J. Liu · 2023
Later among the works it cites.
Judging llm-as-a-judge with mt-bench and chatbot arena
L. Zheng, W.-L. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, E. Xing, H. Zhang, J. E. Gonzalez, and I. Stoica · 2023
Later among the works it cites.
Lima: Less is more for alignment
C. Zhou, P. Liu, P. Xu, S. Iyer, J. Sun, Y. Mao, X. Ma, A. Efrat, P. Yu, L. YU, S. Zhang, G. Ghosh, M. Lewis, L. Zettlemoyer, and O. Levy · 2023
Later among the works it cites.
Introducing the next generation of claude, 2024
A. Anthropic · 2024
Closest in time.
Chatbot arena: An open platform for evaluating llms by human preference
W.-L. Chiang, L. Zheng, Y. Sheng, A. N. Angelopoulos, T. Li, D. Li, H. Zhang, B. Zhu, M. Jordan, J. E. Gonzalez, et al · 2024
Closest in time.
Scaling instruction-finetuned language models
H. W. Chung, L. Hou, S. Longpre, B. Zoph, Y. Tay, W. Fedus, Y. Li, X. Wang, M. Dehghani, S. Brahma, A. Webson, S. S. Gu, Z. Dai, M. Suzgun, X. Chen, A. Chowdhery, A. Castro-Ros, M. Pellat, K. Robinson, D. Valter, S. Narang, G. Mishra, A. Yu, V. Zhao, Y. Huang, A. Dai, H. Yu, S. Petrov, E. H. Chi, J. Dean, J. Devlin, A. Roberts, D. Zhou, Q. V. Le, and J. Wei · 2024
Closest in time.
Length-controlled alpacaeval: A simple way to debias automatic evaluators
Y. Dubois, B. Galambosi, P. Liang, and T. B. Hashimoto · 2024
Closest in time.
GPTScore: Evaluate as you desire
J. Fu, S.-K. Ng, Z. Jiang, and P. Liu · 2024
Closest in time.
One thousand and one pairs: A" novel" challenge for long-context language models
M. Karpinska, K. Thai, K. Lo, T. Goyal, and M. Iyyer · 2024
Closest in time.
Rewardbench: Evaluating reward models for language modeling
N. Lambert, V. Pyatkin, J. Morrison, L. Miranda, B. Y. Lin, K. Chandu, N. Dziri, S. Kumar, T. Zick, Y. Choi, et al · 2024
Closest in time.
PRD: Peer rank and discussion improve large language model based evaluations
R. Li, T. Patel, and X. Du · 2024
Closest in time.
Let’s verify step by step
H. Lightman, V. Kosaraju, Y. Burda, H. Edwards, B. Baker, T. Lee, J. Leike, J. Schulman, I. Sutskever, and K. Cobbe · 2024
Closest in time.
Wildbench: Benchmarking llms with challenging tasks from real users in the wild
B. Y. Lin, Y. Deng, K. Chandu, F. Brahman, A. Ravichander, V. Pyatkin, N. Dziri, R. L. Bras, and Y. Choi · 2024
Closest in time.
Benchmarking generation and evaluation capabilities of large language models for instruction controllable summarization
Y. Liu, A. Fabbri, J. Chen, Y. Zhao, S. Han, S. Joty, P. Liu, D. Radev, C.-S. Wu, and A. Cohan · 2024
Closest in time.
Introducing meta llama 3: The most capable openly available llm to date
A. Meta · 2024
Closest in time.
Llm evaluators recognize and favor their own generations
A. Panickssery, S. R. Bowman, and S. Feng · 2024
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
R. Rafailov, A. Sharma, E. Mitchell, C. D. Manning, S. Ermon, and C. Finn · 2024
Closest in time.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
M. Reid, N. Savinov, D. Teplyashin, D. Lepikhin, T. Lillicrap, J.-b. Alayrac, R. Soricut, A. Lazaridou, O. Firat, J. Schrittwieser, et al · 2024
Closest in time.
Minicheck: Efficient fact-checking of llms on grounding documents
L. Tang, P. Laban, and G. Durrett · 2024
Closest in time.
Helpsteer2: Open-source dataset for training top-performing reward models
Z. Wang, Y. Dong, O. Delalleau, J. Zeng, G. Shen, D. Egert, J. J. Zhang, M. N. Sreedhar, and O. Kuchaiev · 2024
Closest in time.
Long-form factuality in large language models
J. Wei, C. Yang, X. Song, Y. Lu, N. Hu, D. Tran, D. Peng, R. Liu, D. Huang, C. Du, et al · 2024
Closest in time.