Fetching the paper…
Reading the bibliography…
We introduce the GEM (Generative Estimator for Mutual Information), an evaluation metric for assessing language generation by Large Language Models (LLMs), particularly in generating informative judgments, without the need for a gold standard reference.
A mathematical theory of communication
Claude Elwood Shannon · 1948
Earlier work this paper cites.
Verification of forecasts expressed in terms of probability
Glenn W Brier · 1950
Earlier work this paper cites.
Comparison of experiments
David Blackwell · 1951
Earlier work this paper cites.
Rational decisions
Irving John Good · 1952
Earlier work this paper cites.
Equivalent comparisons of experiments
David Blackwell · 1953
Earlier work this paper cites.
Transmission of information: A statistical theory of communications
Robert M Fano and David Hawkins · 1961
Earlier work this paper cites.
A general class of coefficients of divergence of one distribution from another
Syed Mumtaz Ali and Samuel D Silvey · 1966
Earlier work this paper cites.
Proper scores for probability forecasters
Arlo D Hendrickson and Robert J Buehler · 1971
Earlier work this paper cites.
Statistical power analysis for the behavioral sciences
J Cohen · 1988
Earlier work this paper cites.
Word association norms, mutual information, and lexicography
Kenneth Church and Patrick Hanks · 1990
Earlier work this paper cites.
Development of the review quality instrument (rqi) for assessing peer reviews of manuscripts
Susan Van Rooyen, Nick Black, and Fiona Godlee · 1999
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin · 2004
Earlier work this paper cites.
Evaluating content selection in summarization: The pyramid method
Ani Nenkova and Rebecca J Passonneau · 2004
Earlier work this paper cites.
A bayesian truth serum for subjective data
Drazen Prelec · 2004
Earlier work this paper cites.
Eliciting informative feedback: The peer-prediction method
Nolan Miller, Paul Resnick, and Richard Zeckhauser · 2005
Earlier work this paper cites.
Elements of information theory
MTCAJ Thomas and A Thomas Joy · 2006
Earlier work this paper cites.
Crowdsourced judgement elicitation with endogenous proficiency
Anirban Dasgupta and Arpita Ghosh · 2013
Earlier work this paper cites.
Mine: mutual information neural estimation
Mohamed Ishmael Belghazi, Aristide Baratin, Sai Rajeswar, Sherjil Ozair, Yoshua Bengio, Aaron Courville, and R Devon Hjelm · 2018
Cited alongside, same era.
Learning deep representations by mutual information estimation and maximization
R Devon Hjelm, Alex Fedorov, Samuel Lavoie-Marchildon, Karan Grewal, Phil Bachman, Adam Trischler, and Yoshua Bengio · 2018
Cited alongside, same era.
Eliciting expertise without verification
Yuqing Kong and Grant Schoenebeck · 2018
Cited alongside, same era.
Water from two rocks: Maximizing the mutual information
Yuqing Kong and Grant Schoenebeck · 2018
Cited alongside, same era.
Max-mig: an information theoretic approach for joint learning from crowds
Peng Cao, Yilun Xu, Yuqing Kong, and Yizhou Wang · 2019
Bartscore: Evaluating generated text as text generation
Weizhe Yuan, Graham Neubig, and Pengfei Liu · 2021
Later among the works it cites.
The dangers of using large language models for peer review
Tjibbe Donker · 2023
Later among the works it cites.
Gptscore: Evaluate as you desire
Jinlan Fu, See-Kiong Ng, Zhengbao Jiang, and Pengfei Liu · 2023
Later among the works it cites.
Time travel in llms: Tracing data contamination in large language models
Shahriar Golchin and Mihai Surdeanu · 2023
Later among the works it cites.
Peer reviews of peer reviews: A randomized controlled trial and other experiments
Alexander Goldberg, Ivan Stelmakh, Kyunghyun Cho, Alice Oh, Alekh Agarwal, Danielle Belgrave, and Nihar B Shah · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
On the measure of intelligence
François Chollet · 2019
Cited alongside, same era.
An information theoretic framework for designing information elicitation mechanisms that reward truth-telling
Yuqing Kong and Grant Schoenebeck · 2019
Cited alongside, same era.
Natural questions: a benchmark for question answering research
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, et al · 2019
Cited alongside, same era.
Tools used to assess the quality of peer review reports: a methodological systematic review
Cecilia Superchi, José Antonio González, Ivan Solà, Erik Cobo, Darko Hren, and Isabelle Boutron · 2019
Cited alongside, same era.
L_dmi: A novel information-theoretic loss function for training deep nets robust to label noise
Yilun Xu, Peng Cao, Yuqing Kong, and Yizhou Wang · 2019
Cited alongside, same era.
Hellaswag: Can a machine really finish your sentence?
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi · 2019
Cited alongside, same era.
Bertscore: Evaluating text generation with bert
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi · 2019
Cited alongside, same era.
Pei Ke, Bosi Wen, Zhuoer Feng, Xiao Liu, Xuanyu Lei, Jiale Cheng, Shengyuan Wang, Aohan Zeng, Yuxiao Dong, Hongning Wang, et al · 2023
Later among the works it cites.
Can large language models provide useful feedback on research papers? a large-scale empirical analysis, 2023
Weixin Liang, Yuhui Zhang, Hancheng Cao, Binglu Wang, Daisy Ding, Xinyu Yang, Kailas Vodrahalli, Siyu He, Daniel Smith, Yian Yin, Daniel McFarland, and James Zou · 2023
Later among the works it cites.
Reviewergpt? an exploratory study on using large language models for paper reviewing, 2023
Ryan Liu and Nihar B. Shah · 2023
Later among the works it cites.
Proving test set contamination in black box language models
Yonatan Oren, Nicole Meister, Niladri Chatterji, Faisal Ladhak, and Tatsunori B Hashimoto · 2023
Later among the works it cites.
Gpt4 is slightly helpful for peer-review assistance: A pilot study, 2023
Zachary Robertson · 2023
Later among the works it cites.
Nlp evaluation in trouble: On the need to measure llm data contamination for each benchmark
Oscar Sainz, Jon Ander Campos, Iker García-Ferrero, Julen Etxaniz, Oier Lopez de Lacalle, and Eneko Agirre · 2023
Later among the works it cites.
Benchmarking foundation models with language-model-as-an-examiner
Yushi Bai, Jiahao Ying, Yixin Cao, Xin Lv, Yuze He, Xiaozhi Wang, Jifan Yu, Kaisheng Zeng, Yijia Xiao, Haozhe Lyu, et al · 2024
Closest in time.
What can natural language processing do for peer review?
Ilia Kuznetsov, Osama Mohammed Afzal, Koen Dercksen, Nils Dycke, Alexander Goldberg, Tom Hope, Dirk Hovy, Jonathan K Kummerfeld, Anne Lauscher, Kevin Leyton-Brown, et al · 2024
Closest in time.
Eliciting informative text evaluations with large language models
Yuxuan Lu, Shengwei Xu, Yichi Zhang, Yuqing Kong, and Grant Schoenebeck · 2024
Closest in time.
Llm evaluators recognize and favor their own generations
Arjun Panickssery, Samuel R Bowman, and Shi Feng · 2024
Closest in time.
Ai models collapse when trained on recursively generated data
Ilia Shumailov, Zakhar Shumaylov, Yiren Zhao, Nicolas Papernot, Ross Anderson, and Yarin Gal · 2024
Closest in time.
Spot check equivalence: an interpretable metric for information elicitation mechanisms
Shengwei Xu, Yichi Zhang, Paul Resnick, and Grant Schoenebeck · 2024
Closest in time.
stella_en_400m_v5
Dun Zhang · 2024
Closest in time.