Fetching the paper…
Reading the bibliography…
The training data for many Large Language Models (LLMs) is contaminated with test data.
Sentence-bert: Sentence embeddings using siamese bert-networks
N. Reimers and I. Gurevych · 1908
Earlier work this paper cites.
Modified Randomization Tests for Nonparametric Hypotheses
M. Dwass · 1957
Earlier work this paper cites.
The Design of Experiments
R. A. Fisher · 1974
Earlier work this paper cites.
Problems of monetary management: the UK experience
C. A. Goodhart · 1984
Earlier work this paper cites.
‘improving ratings’: audit in the british university system
M. Strathern · 1997
Earlier work this paper cites.
Language models are few-shot learners, 2020
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei · 2005
Earlier work this paper cites.
The elements of statistical learning: data mining, inference and prediction
T. Hastie, R. Tibshirani, and J. Friedman · 2009
Earlier work this paper cites.
Modern Two-Sample Tests, July 2012
normaldeviate · 2012
Earlier work this paper cites.
Understanding Machine Learning - From Theory to Algorithms
S. Shalev-Shwartz and S. Ben-David · 2014
Earlier work this paper cites.
Machine Learning Yearning
A. Ng · 2017
Earlier work this paper cites.
Axiomatic attribution for deep networks, 2017
M. Sundararajan, A. Taly, and Q. Yan · 2017
Earlier work this paper cites.
UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction, Feb. 2018
L. McInnes, J. Healy, and J. Melville · 2018
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding, May 2019
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2019
Earlier work this paper cites.
OpenIntro Statistics
D. Diez, M. Çetinkaya-Rundel, and C. Barr · 2019
Earlier work this paper cites.
Specification gaming: the flip side of AI ingenuity, Apr. 2020
V. Krakovna, Jonathan Uesato, Vladimir Mikulik, Matthew Rahtz, Tom Everitt, Ramana Kumar, Zac Kenton, Jan Leike, and Shane Legg · 2020
Earlier work this paper cites.
The problem with metrics is a fundamental problem for ai
R. Thomas and D. Uminsky · 2020
Earlier work this paper cites.
Training Verifiers to Solve Math Word Problems, Nov. 2021
K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, C. Hesse, and J. Schulman · 2021
Cited alongside, same era.
“everyone wants to do the model work, not the data work”: Data cascades in high-stakes ai
N. Sambasivan, S. Kapania, H. Highfill, D. Akrong, P. Paritosh, and L. M. Aroyo · 2021
Cited alongside, same era.
TruthfulQA: Measuring How Models Mimic Human Falsehoods, May 2022
S. Lin, J. Hilton, and O. Evans · 2022
Cited alongside, same era.
J. Bai, S. Bai, Y. Chu, Z. Cui, K. Dang, X. Deng, Y. Fan, W. Ge, Y. Han, F. Huang, B. Hui, L. Ji, M. Li, J. Lin, R. Lin, D. Liu, G. Liu, C. Lu, K. Lu, J. Ma, R. Men, X. Ren, X. Ren, C. Tan, S. Tan, J. Tu, P. Wang, S. Wang, W. Wang, S. Wu, B. Xu, J. Xu, A. Yang, H. Yang, J. Yang, S. Yang, Y. Yao, B. Yu, H. Yuan, Z. Yuan, J. Zhang, X. Zhang, Y. Zhang, Z. Zhang, C. Zhou, J. Zhou, X. Zhou, and T. Zhu · 2023
Cited alongside, same era.
A survey of large language models, 2023
W. X. Zhao, K. Zhou, J. Li, T. Tang, X. Wang, Y. Hou, Y. Min, B. Zhang, J. Zhang, Z. Dong, Y. Du, C. Yang, Y. Chen, Z. Chen, J. Jiang, R. Ren, Y. Li, X. Tang, Z. Liu, P. Liu, J.-Y. Nie, and J.-R. Wen · 2023
Later among the works it cites.
N. Alzahrani, H. A. Alyahya, Y. Alnumay, S. Alrashed, S. Alsubaie, Y. Almushaykeh, F. Mirza, N. Alotaibi, N. Altwairesh, A. Alowisheq, M. S. Bari, and H. Khan · 2024
Closest in time.
A survey on evaluation of large language models
Y. Chang, X. Wang, J. Wang, Y. Wu, L. Yang, K. Zhu, H. Chen, X. Yi, C. Wang, Y. Wang, et al · 2024
Closest in time.
Ultrafeedback: Boosting language models with scaled ai feedback, 2024
G. Cui, L. Yuan, N. Ding, G. Yao, B. He, W. Zhu, Y. Ni, G. Xie, R. Xie, Y. Lin, Z. Liu, and M. Sun · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
R. Bommasani, K. Klyman, S. Longpre, S. Kapoor, N. Maslej, B. Xiong, D. Zhang, and P. Liang · 2023
Cited alongside, same era.
What’s going on with the Open LLM Leaderboard?, June 2023
C. Fourrier, N. Habib, J. Launay, and T. Wolf · 2023
Cited alongside, same era.
Time travel in llms: Tracing data contamination in large language models
S. Golchin and M. Surdeanu · 2023
Cited alongside, same era.
An Introduction to Statistical Learning: with Applications in Python
G. James, D. Witten, T. Hastie, R. Tibshirani, and J. Taylor · 2023
Cited alongside, same era.
Goodhart’s law in reinforcement learning
J. Karwowski, O. Hayman, X. Bai, K. Kiendlhofer, C. Griffin, and J. Skalse · 2023
Cited alongside, same era.
Toward comprehensive risk assessments and assurance of ai-based systems
H. Khlaaf · 2023
Cited alongside, same era.
The decontaminated evaluation of gpt-4, 2023
B. Marie · 2023
Cited alongside, same era.
Proving Test Set Contamination in Black Box Language Models, Nov. 2023
Y. Oren, N. Meister, N. Chatterji, F. Ladhak, and T. B. Hashimoto · 2023
Cited alongside, same era.
C. Deng, Y. Zhao, X. Tang, M. Gerstein, and A. Cohan · 2024
Closest in time.
Open llm leaderboard v2
C. Fourrier, N. Habib, A. Lozovskaya, K. Szafer, and T. Wolf · 2024
Closest in time.
On the Term “Randomization Test”
J. Hemerik · 2024
Closest in time.
Investigating Data Contamination for Pre-training Language Models, Jan. 2024
M. Jiang, K. Z. Liu, M. Zhong, R. Schaeffer, S. Ouyang, J. Han, and S. Koyejo · 2024
Closest in time.
A comprehensive overview of large language models, 2024
H. Naveed, A. U. Khan, S. Qiu, M. Saqib, S. Anwar, M. Usman, N. Akhtar, N. Barnes, and A. Mian · 2024
Closest in time.
International Scientific Report on the Safety of Advanced AI - Interim Report
D. Privitera, T. Besiroglu, R. Bommasani, S. Casper, S. Longpre, S. Mindermann, B. Adekanmbi, Y. Choi, D. Goldfarb, H. Heidari, L. Khalatbari, V. Mavroudis, M. Mazeika, K. Yee Ng, C. T. Okolo, D. Raji, T. Skeadas, and F. Tramèr · 2024
Closest in time.
We need a science of evals
A. Research · 2024
Closest in time.
CONDA 2024 | The 1st Workshop on Data Contamination, 2024
O. Sainz, I. García-Ferrero, J. Ander, Y. Elazar, and E. Agirre · 2024
Closest in time.
A Careful Examination of Large Language Model Performance on Grade School Arithmetic, May 2024
H. Zhang, J. Da, D. Lee, V. Robinson, C. Wu, W. Song, T. Zhao, P. Raja, D. Slack, Q. Lyu, S. Hendryx, R. Kaplan, M. Lunati, and S. Yue · 2024
Closest in time.
Large Language Models Are Not Robust Multiple Choice Selectors, Feb. 2024
C. Zheng, H. Zhou, F. Meng, J. Zhou, and M. Huang · 2024
Closest in time.
Fake It Till You Make It: Guidelines for Effective Synthetic Data Generation
F. K. Dankar and M. Ibrahim · 2076
Closest in time.