Fetching the paper…
Reading the bibliography…
The advancement of large language models (LLMs) relies on evaluation using public benchmarks, but data contamination can lead to overestimated performance.
F. Jelinek, R. L. Mercer, L. R. Bahl, and J. K. Baker, “Perplexity—a measure of the difficulty of speech recognition tasks,” The Journal of the Acoustical Society of America
1977
Earlier work this paper cites.
J.-l. Gailly and M. Adler, “Zlib compression library,” 2004
2004
Earlier work this paper cites.
R. Koncel-Kedziorski, S. Roy, A. Amini, N. Kushman, and H. Hajishirzi, “MAWPS: A math word problem repository,” in Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies
2016
Earlier work this paper cites.
N. Reimers and I. Gurevych, “Sentence-BERT: Sentence embeddings using Siamese BERT-networks,” in Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)
2019
Earlier work this paper cites.
S.-y. Miao, C.-C. Liang, and K.-Y. Su, “A diverse corpus for evaluating and developing English math word problem solvers,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics
2020
Earlier work this paper cites.
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei, “Language models are few-shot learners,” in Proceedings of NIPS
2020
Earlier work this paper cites.
2021
Earlier work this paper cites.
D. Hendrycks, C. Burns, S. Kadavath, A. Arora, S. Basart, E. Tang, D. Song, and J. Steinhardt, “Measuring mathematical problem solving with the MATH dataset,” in Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2)
2021
Earlier work this paper cites.
D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt, “Measuring massive multitask language understanding,” in Proceedings of ICLR
2021
Earlier work this paper cites.
A. Patel, S. Bhattamishra, and N. Goyal, “Are NLP models really able to solve simple math word problems?,” in Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies
2021
Earlier work this paper cites.
N. Carlini, F. Tramer, E. Wallace, M. Jagielski, A. Herbert-Voss, K. Lee, A. Roberts, T. Brown, D. Song, U. Erlingsson, et al
2021
Earlier work this paper cites.
A. Elangovan, J. He, and K. Verspoor, “Memorization vs. generalization : Quantifying data leakage in NLP performance evaluation,” in Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume
2021
Earlier work this paper cites.
I. Magar and R. Schwartz, “Data contamination: From memorization to exploitation,” in Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers)
2022
Cited alongside, same era.
Z. Ji, N. Lee, R. Frieske, T. Yu, D. Su, Y. Xu, E. Ishii, Y. J. Bang, A. Madotto, and P. Fung, “Survey of hallucination in natural language generation,” ACM Computing Surveys
2023
Cited alongside, same era.
L. Gao, A. Madaan, S. Zhou, U. Alon, P. Liu, Y. Yang, J. Callan, and G. Neubig, “Pal: Program-aided language models,” in International Conference on Machine Learning
2023
Cited alongside, same era.
P. Lu, L. Qiu, K.-W. Chang, Y. N. Wu, S.-C. Zhu, T. Rajpurohit, P. Clark, and A. Kalyan, “Dynamic prompt learning via policy gradient for semi-structured mathematical reasoning,” in The Eleventh International Conference on Learning Representations
2023
Cited alongside, same era.
2024
Closest in time.
2024
Closest in time.
W. Shi, A. Ajith, M. Xia, Y. Huang, D. Liu, T. Blevins, D. Chen, and L. Zettlemoyer, “Detecting pretraining data from large language models,” in The Twelfth International Conference on Learning Representations
2024
Closest in time.
2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2023
Cited alongside, same era.
M. Javaheripi, S. Bubeck, M. Abdin, J. Aneja, C. C. T. Mendes, W. Chen, A. Del Giorno, R. Eldan, S. Gopi, S. Gunasekar, et al
2023
Cited alongside, same era.
W. Lian, B. Goodson, E. Pentland, A. Cook, C. Vong, and "Teknium", “Openorca: An open dataset of gpt augmented flan reasoning traces.” https://https://huggingface.co/Open-Orca/OpenOrca , 2023
2023
Cited alongside, same era.
Y. Bai, J. Ying, Y. Cao, X. Lv, Y. He, X. Wang, J. Yu, K. Zeng, Y. Xiao, H. Lyu, J. Zhang, J. Li, and L. Hou, “Benchmarking foundation models with language-model-as-an-examiner,” in Thirty-seventh Conference on Neural Information Processing Systems Datasets and Benchmarks Track
2023
Cited alongside, same era.
Y. Li, F. Geurin, and C. Lin, “Latesteval: Addressing data contamination in language model evaluation through dynamic and time-sensitive test construction,” in AAAI Conference on Artificial Intelligence
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
K. Zhu, J. Chen, J. Wang, N. Z. Gong, D. Yang, and X. Xie, “Dyval: Dynamic evaluation of large language models for reasoning tasks,” in The Twelfth International Conference on Learning Representations
2024
Closest in time.
H. Zhang, J. Da, D. Lee, V. Robinson, C. Wu, W. Song, T. Zhao, P. Raja, D. Slack, Q. Lyu, et al
2024
Closest in time.
2024
Closest in time.
Y. Oren, N. Meister, N. S. Chatterji, F. Ladhak, and T. Hashimoto, “Proving test set contamination in black-box language models,” in The Twelfth International Conference on Learning Representations
2024
Closest in time.