Fetching the paper…
Reading the bibliography…
It is now common to evaluate Large Language Models (LLMs) by having humans manually vote to evaluate model outputs, in contrast to typical benchmarks that evaluate knowledge or skill at some particular task.
A vector space model for automatic indexing
Gerard Salton, Anita Wong, and Chung-Shu Yang · 1975
Earlier work this paper cites.
Term-weighting approaches in automatic text retrieval
Gerard Salton and Christopher Buckley · 1988
Earlier work this paper cites.
The eigentrust algorithm for reputation management in p2p networks
Sepandar D Kamvar, Mario T Schlosser, and Hector Garcia-Molina · 2003
Earlier work this paper cites.
Mm algorithms for generalized bradley-terry models
David R Hunter · 2004
Earlier work this paper cites.
Correlation between rouge and human evaluation of extractive meeting summaries
Feifan Liu and Yang Liu · 2008
Earlier work this paper cites.
A survey of attack and defense techniques for reputation systems
Kevin Hoffman, David Zage, and Cristina Nita-Rotaru · 2009
Earlier work this paper cites.
A multi-faceted approach to anonymity online: Examining the relations between anonymity and antisocial behaviour
Rebecca Chui · 2014
Earlier work this paper cites.
Revisiting summarization evaluation for scientific articles
Arman Cohan and Nazli Goharian · 2016
Earlier work this paper cites.
{ \{ AnonRep } \} : Towards { \{ Tracking-Resistant } \} anonymous reputation
Ennan Zhai, David Isaac Wolinsky, Ruichuan Chen, Ewa Syta, Chao Teng, and Bryan Ford · 2016
Earlier work this paper cites.
Drop: A reading comprehension benchmark requiring discrete reasoning over paragraphs
Dheeru Dua, Yizhong Wang, Pradeep Dasigi, Gabriel Stanovsky, Sameer Singh, and Matt Gardner · 2019
Earlier work this paper cites.
HellaSwag: Can a machine really finish your sentence?
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi · 2019
Earlier work this paper cites.
Fltrust: Byzantine-robust federated learning via trust bootstrapping
Xiaoyu Cao, Minghong Fang, Jia Liu, and Neil Zhenqiang Gong · 2020
Earlier work this paper cites.
De-anonymizing text by fingerprinting language generation
Zhen Sun, Roei Schuster, and Vitaly Shmatikov · 2020
Earlier work this paper cites.
Reverse engineering configurations of neural text generation models
Yi Tay, Dara Bahri, Che Zheng, Clifford Brunk, Donald Metzler, and Andrew Tomkins · 2020
Earlier work this paper cites.
Program synthesis with large language models
Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, et al · 2021
Earlier work this paper cites.
Evaluating large language models trained on code
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde De Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al · 2021
Earlier work this paper cites.
Summeval: Re-evaluating summarization evaluation
Alexander R Fabbri, Wojciech Kryściński, Bryan McCann, Caiming Xiong, Richard Socher, and Dragomir Radev · 2021
Earlier work this paper cites.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt · 2021
Cited alongside, same era.
The perils of using Mechanical Turk to evaluate open-ended text generation
Marzena Karpinska, Nader Akoury, and Mohit Iyyer · 2021
Cited alongside, same era.
Through the looking glass: Learning to attribute synthetic text generated by language models
Shaoor Munir, Brishna Batool, Zubair Shafiq, Padmini Srinivasan, and Fareed Zaffar · 2021
Cited alongside, same era.
Data contamination: From memorization to exploitation
Inbal Magar and Roy Schwartz · 2022
Cited alongside, same era.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Cited alongside, same era.
Chatbot arena: An open platform for evaluating LLMs by human preference
Wei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos, Tianle Li, Dacheng Li, Banghua Zhu, Hao Zhang, Michael Jordan, Joseph E. Gonzalez, and Ion Stoica · 2024
Later among the works it cites.
Undetectable watermarks for language models
Miranda Christ, Sam Gunn, and Or Zamir · 2024
Later among the works it cites.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al · 2024
Later among the works it cites.
Authorship attribution in the era of llms: Problems, methodologies, and challenges
Baixiang Huang, Canyu Chen, and Kai Shu · 2024
Later among the works it cites.
Openassistant conversations-democratizing large language model alignment
Andreas Köpf, Yannic Kilcher, Dimitri von Rütte, Sotiris Anagnostidis, Zhi Rui Tam, Keith Stevens, Abdullah Barhoum, Duc Nguyen, Oliver Stanley, Richárd Nagyfi, et al · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A study of { \{ Multi-Factor } \} and { \{ Risk-Based } \} authentication availability
Anthony Gavazzi, Ryan Williams, Engin Kirda, Long Lu, Andre King, Andy Davis, and Tim Leek · 2023
Cited alongside, same era.
Prometheus: Inducing fine-grained evaluation capability in language models
Seungone Kim, Jamin Shin, Yejin Cho, Joel Jang, Shayne Longpre, Hwaran Lee, Sangdoo Yun, Seongjin Shin, Sungdong Kim, James Thorne, et al · 2023
Cited alongside, same era.
A watermark for large language models
John Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz, Ian Miers, and Tom Goldstein · 2023
Cited alongside, same era.
Okapi: Instruction-tuned large language models in multiple languages with reinforcement learning from human feedback
Viet Lai, Chien Nguyen, Nghia Ngo, Thuat Nguyen, Franck Dernoncourt, Ryan Rossi, and Thien Nguyen · 2023
Cited alongside, same era.
Evaluation of real-world risk-based authentication at online services revisited: Complexity wins
Jan-Phillip Makowski and Daniela Pöhn · 2023
Cited alongside, same era.
Gpqa: A graduate-level google-proof q&a benchmark
David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, and Samuel R Bowman · 2023
Cited alongside, same era.
SolidGoldMagikarp (plus, prompt generation)
Jessica Rumbelow and Matthew Watkins · 2023
Cited alongside, same era.
Later among the works it cites.
Robust distortion-free watermarks for language models
Rohith Kuditipudi, John Thickstun, Tatsunori Hashimoto, and Percy Liang · 2024
Later among the works it cites.
Why aren’t we using passkeys? obstacles companies face deploying FIDO2 passwordless authentication
Leona Lassak, Elleen Pan, Blase Ur, and Maximilian Golla · 2024
Later among the works it cites.
Talk arena: Interactive evaluation of large audio models, 2024
Minzhi Li, Will Held, Michael J. Ryan, Kunat Pipatanakul, Potsawee Manakul, Hao Zhu, and Diyi Yang · 2024
Later among the works it cites.
Wildvision: Evaluating vision-language models in the wild with human preferences
Yujie Lu, Dongfu Jiang, Wenhu Chen, William Yang Wang, Yejin Choi, and Bill Yuchen Lin · 2024
Later among the works it cites.
Proving test set contamination in black box language models
Yonatan Oren, Nicole Meister, Niladri Chatterji, Faisal Ladhak, and Tatsunori B Hashimoto · 2024
Later among the works it cites.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Machel Reid, Nikolay Savinov, Denis Teplyashin, Dmitry Lepikhin, Timothy Lillicrap, Jean-baptiste Alayrac, Radu Soricut, Angeliki Lazaridou, Orhan Firat, Julian Schrittwieser, et al · 2024
Later among the works it cites.
Detecting pretraining data from large language models
Weijia Shi, Anirudh Ajith, Mengzhou Xia, Yangsibo Huang, Daogao Liu, Terra Blevins, Danqi Chen, and Luke Zettlemoyer · 2024
Later among the works it cites.
Results reports
VirusTotal · 2024
Later among the works it cites.
M4gt-bench: Evaluation benchmark for black-box machine-generated text detection
Yuxia Wang, Jonibek Mansurov, Petar Ivanov, Jinyan Su, Artem Shelmanov, Akim Tsvigun, Osama Mohanned Afzal, Tarek Mahmoud, Giovanni Puccetti, Thomas Arnold, et al · 2024
Later among the works it cites.
Conceptmix: A compositional image generation benchmark with controllable difficulty
Xindi Wu, Dingli Yu, Yangsibo Huang, Olga Russakovsky, and Sanjeev Arora · 2024
Later among the works it cites.
On memorization of large language models in logical reasoning
Chulin Xie, Yangsibo Huang, Chiyuan Zhang, Da Yu, Xinyun Chen, Bill Yuchen Lin, Bo Li, Badih Ghazi, and Ravi Kumar · 2024
Later among the works it cites.
Benchmarking benchmark leakage in large language models
Ruijie Xu, Zengzhi Wang, Run-Ze Fan, and Pengfei Liu · 2024
Later among the works it cites.
Challenges in trustworthy human evaluation of chatbots, 2024
Wenting Zhao, Alexander M. Rush, and Tanya Goyal · 2024
Later among the works it cites.