Fetching the paper…
Reading the bibliography…
Evaluation of large language models (LLMs) has raised great concerns in the community due to the issue of data contamination.
The proof and measurement of association between two things
Spearman, C · 1904
Earlier work this paper cites.
The history and theory of correlation
Pearson, K · 1920
Earlier work this paper cites.
A new measure of rank correlation
Kendall, M. G · 1938
Earlier work this paper cites.
” general intelligence” objectively determined and measured
Spearman, C · 1961
Earlier work this paper cites.
The matthew effect in science: The reward and communication systems of science are considered
Merton, R. K · 1968
Earlier work this paper cites.
Introduction to psychometric theory
Raykov, T. and Marcoulides, G. A · 2011
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge
Clark, P., Cowhey, I., Etzioni, O., Khot, T., Sabharwal, A., Schoenick, C., and Tafjord, O · 2018
Earlier work this paper cites.
Deeptest: Automated testing of deep-neural-network-driven autonomous cars
Tian, Y., Pei, K., Jana, S., and Ray, B · 2018
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Beyond accuracy: Behavioral testing of NLP models with CheckList
Ribeiro, M. T., Wu, T., Guestrin, C., and Singh, S · 2020
Earlier work this paper cites.
On the dangers of stochastic parrots: Can language models be too big?
Bender, E. M., Gebru, T., McMillan-Major, A., and Shmitchell, S · 2021
Earlier work this paper cites.
Evaluating large language models trained on code
Chen, M., Tworek, J., Jun, H., Yuan, Q., Pinto, H. P. d. O., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., et al · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., et al · 2021
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2021
Earlier work this paper cites.
Dynabench: Rethinking benchmarking in NLP
Kiela, D., Bartolo, M., Nie, Y., Kaushik, D., Geiger, A., Wu, Z., Vidgen, B., Prasad, G., Singh, A., Ringshia, P., Ma, Z., Thrush, T., Riedel, S., Waseem, Z., Stenetorp, P., Jia, R., Bansal, M., Potts, C., and Williams, A · 2021
Earlier work this paper cites.
Dynaboard: An evaluation-as-a-service platform for holistic next-generation benchmarking
Ma, Z., Ethayarajh, K., Thrush, T., Jain, S., Wu, L., Jia, R., Potts, C., Williams, A., and Kiela, D · 2021
Earlier work this paper cites.
Adaptive testing of computer vision models
Gao, I., Ilharco, G., Lundberg, S., and Ribeiro, M. T · 2022
Earlier work this paper cites.
Adaptive testing and debugging of nlp models
Ribeiro, M. T. and Lundberg, S · 2022
Earlier work this paper cites.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Srivastava, A., Rastogi, A., Rao, A., Shoeb, A. A. M., Abid, A., Fisch, A., Brown, A. R., Santoro, A., Gupta, A., Garriga-Alonso, A., et al · 2022
Earlier work this paper cites.
Challenging big-bench tasks and whether chain-of-thought can solve them
Suzgun, M., Scales, N., Schärli, N., Gehrmann, S., Tay, Y., Chung, H. W., Chowdhery, A., Le, Q. V., Chi, E. H., Zhou, D., , and Wei, J · 2022
Cited alongside, same era.
Self-instruct: Aligning language model with self generated instructions
Wang, Y., Kordi, Y., Mishra, S., Liu, A., Smith, N. A., Khashabi, D., and Hajishirzi, H · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al · 2022
Cited alongside, same era.
Yang, L., Zhang, S., Qin, L., Li, Y., Wang, Y., Liu, H., Wang, J., Xie, X., and Zhang, Y · 2022
Cited alongside, same era.
Benchmarking foundation models with language-model-as-an-examiner
An open source data contamination report for llama series models
Li, Y · 2023
Later among the works it cites.
Holistic evaluation of language models
Liang, P., Bommasani, R., Lee, T., Tsipras, D., Soylu, D., Yasunaga, M., Zhang, Y., Narayanan, D., Wu, Y., Kumar, A., Newman, B., Yuan, B., Yan, B., Zhang, C., Cosgrove, C. A., Manning, C. D., Re, C., Acosta-Navas, D., Hudson, D. A., Zelikman, E., Durmus, E., Ladhak, F., Rong, F., Ren, H., Yao, H., WANG, J., Santhanam, K., Orr, L., Zheng, L., Yuksekgonul, M., Suzgun, M., Kim, N., Guha, N., Chatterji, N. S., Khattab, O., Henderson, P., Huang, Q., Chi, R. A., Xie, S. M., Santurkar, S., Ganguli, S., Hashimoto, T., Icard, T., Zhang, T., Chaudhary, V., Wang, W., Li, X., Mai, Y., Zhang, Y., and Koreeda, Y · 2023
Later among the works it cites.
Gpt-4 performs significantly worse on coding problems not in its training data
Lovin, B · 2023
Later among the works it cites.
Mixtral-8x7b-v0.1
MistralAITeam · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bai, Y., Ying, J., Cao, Y., Lv, X., He, Y., Wang, X., Yu, J., Zeng, K., Xiao, Y., Lyu, H., et al · 2023
Cited alongside, same era.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
bench authors, B · 2023
Cited alongside, same era.
The reversal curse: Llms trained on “a is b” fail to learn “b is a”
Berglund, L., Tong, M., Kaufmann, M., Balesni, M., Stickland, A. C., Korbak, T., and Evans, O · 2023
Cited alongside, same era.
Emergent and predictable memorization in large language models
Biderman, S., Prashanth, U. S., Sutawika, L., Schoelkopf, H., Anthony, Q., Purohit, S., and Raf, E · 2023
Cited alongside, same era.
Revealing the structure of language model capabilities
Burnell, R., Hao, H., Conway, A. R., and Orallo, J. H · 2023
Cited alongside, same era.
Weak-to-strong generalization: Eliciting strong capabilities with weak supervision
Burns, C., Izmailov, P., Kirchner, J. H., Baker, B., Gao, L., Aschenbrenner, L., Chen, Y., Ecoffet, A., Joglekar, M., Leike, J., et al · 2023
Cited alongside, same era.
A survey on evaluation of large language models
Chang, Y., Wang, X., Wang, J., Wu, Y., Zhu, K., Chen, H., Yang, L., Yi, X., Wang, C., Wang, Y., et al · 2023
Cited alongside, same era.
Alpacafarm: A simulation framework for methods that learn from human feedback
Dubois, Y., Li, X., Taori, R., Zhang, T., Gulrajani, I., Ba, J., Guestrin, C., Liang, P., and Hashimoto, T. B · 2023
Cited alongside, same era.
Oren, Y., Meister, N., Chatterji, N., Ladhak, F., and Hashimoto, T. B · 2023
Later among the works it cites.
Are emergent abilities of large language models a mirage?
Schaeffer, R., Miranda, B., and Koyejo, S · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Later among the works it cites.
Decodingtrust: A comprehensive assessment of trustworthiness in gpt models
Wang, B., Chen, W., Pei, H., Xie, C., Kang, M., Zhang, C., Xu, C., Xiong, Z., Dutta, R., Schaeffer, R., et al · 2023
Later among the works it cites.
Skywork: A more open bilingual foundation model
Wei, T., Zhao, L., Zhang, L., Zhu, B., Wang, L., Yang, H., Li, B., Cheng, C., Lü, W., Hu, R., et al · 2023
Later among the works it cites.
Rethinking benchmark and contamination for language models with rephrased samples
Yang, S., Chiang, W.-L., Zheng, L., Gonzalez, J. E., and Stoica, I · 2023
Later among the works it cites.
Metamath: Bootstrap your own mathematical questions for large language models
Yu, L., Jiang, W., Shi, H., Yu, J., Liu, Z., Zhang, Y., Kwok, J. T., Li, Z., Weller, A., and Liu, W · 2023
Later among the works it cites.
Agieval: A human-centric benchmark for evaluating foundation models
Zhong, W., Cui, R., Guo, Y., Liang, Y., Lu, S., Wang, Y., Saied, A., Chen, W., and Duan, N · 2023
Later among the works it cites.
Don’t make your llm an evaluation benchmark cheater
Zhou, K., Zhu, Y., Chen, Z., Chen, W., Zhao, W. X., Chen, X., Lin, Y., Wen, J.-R., and Han, J · 2023
Later among the works it cites.
Fool your (vision and) language model with embarrassingly simple permutations
Zong, Y., Yu, T., Zhao, B., Chavhan, R., and Hospedales, T · 2023
Later among the works it cites.
Yi: A series of large language models
01-ai · 2024
Closest in time.
Augmenting math word problems via iterative question composing
Liu, H. and Yao, A. C.-C · 2024
Closest in time.
Pandalm: An automatic evaluation benchmark for llm instruction tuning optimization
Wang, Y., Yu, Z., Zeng, Z., Yang, L., Wang, C., Chen, H., Jiang, C., Xie, R., Wang, J., Xie, X., Ye, W., Zhang, S., and Zhang, Y · 2024
Closest in time.
Self-rewarding language models
Yuan, W., Pang, R. Y., Cho, K., Sukhbaatar, S., Xu, J., and Weston, J · 2024
Closest in time.