Fetching the paper…
Reading the bibliography…
The success of Large Language Models (LLMs) relies heavily on the huge amount of pre-training data learned in the pre-training phase.
Language models are few-shot learners
Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J. D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. 2020 · 1901
Earlier work this paper cites.
Training verifiers to solve math word problems
Cobbe, K.; Kosaraju, V.; Bavarian, M.; Chen, M.; Jun, H.; Kaiser, L.; Plappert, M.; Tworek, J.; Hilton, J.; Nakano, R.; et al. 2021 · 2021
Earlier work this paper cites.
Finetuned language models are zero-shot learners
Wei, J.; Bosma, M.; Zhao, V. Y.; Guu, K.; Yu, A. W.; Lester, B.; Du, N.; Dai, A. M.; and Le, Q. V. 2021 · 2021
Earlier work this paper cites.
Bai, J.; Bai, S.; Chu, Y.; Cui, Z.; Dang, K.; Deng, X.; Fan, Y.; Ge, W.; Han, Y.; Huang, F.; et al. 2023 · 2023
Earlier work this paper cites.
Investigating data contamination in modern benchmarks for large language models
Deng, C.; Zhao, Y.; Tang, X.; Gerstein, M.; and Cohan, A. 2023 · 2023
Earlier work this paper cites.
Cmmlu: Measuring massive multitask language understanding in chinese
Li, H.; Zhang, Y.; Koto, F.; Yang, Y.; Zhao, H.; Gong, Y.; Duan, N.; and Baldwin, T. 2023 · 2023
Earlier work this paper cites.
Membership inference attacks against language models via neighbourhood comparison
Mattern, J.; Mireshghallah, F.; Jin, Z.; Schölkopf, B.; Sachan, M.; and Berg-Kirkpatrick, T. 2023 · 2023
Earlier work this paper cites.
Proving test set contamination in black box language models
Oren, Y.; Meister, N.; Chatterji, N.; Ladhak, F.; and Hashimoto, T. B. 2023 · 2023
Earlier work this paper cites.
Internlm: A multilingual language model with progressively enhanced capabilities
Team, I. 2023 · 2023
Cited alongside, same era.
Cmb: A comprehensive medical benchmark in chinese
Wang, X.; Chen, G. H.; Song, D.; Zhang, Z.; Chen, Z.; Xiao, Q.; Jiang, F.; Li, J.; Wan, X.; Wang, B.; et al. 2023 · 2023
Cited alongside, same era.
Skywork: A more open bilingual foundation model
Wei, T.; Zhao, L.; Zhang, L.; Zhu, B.; Wang, L.; Yang, H.; Li, B.; Cheng, C.; Lü, W.; Hu, R.; et al. 2023 · 2023
Cited alongside, same era.
Baichuan 2: Open large-scale language models
Yang, A.; Xiao, B.; Wang, B.; Zhang, B.; Bian, C.; Yin, C.; Lv, C.; Pan, D.; Wang, D.; Yan, D.; et al. 2023 · 2023
Cited alongside, same era.
Agieval: A human-centric benchmark for evaluating foundation models
Generalization or Memorization: Data Contamination and Trustworthy Evaluation for Large Language Models
Dong, Y.; Jiang, X.; Liu, H.; Jin, Z.; Gu, B.; Yang, M.; and Li, G. 2024 · 2024
Closest in time.
E-EVAL: A Comprehensive Chinese K-12 Education Evaluation Benchmark for Large Language Models
Hou, J.; Ao, C.; Wu, H.; Kong, X.; Zheng, Z.; Tang, D.; Li, C.; Hu, X.; Xu, R.; Ni, S.; et al. 2024 · 2024
Closest in time.
C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models
Huang, Y.; Bai, Y.; Zhu, Z.; Zhang, J.; Zhang, J.; Su, T.; Liu, J.; Lv, C.; Zhang, Y.; Fu, Y.; et al. 2024 · 2024
Closest in time.
Benchmarking benchmark leakage in large language models
Xu, R.; Wang, Z.; Fan, R.-Z.; and Liu, P. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zhong, W.; Cui, R.; Guo, Y.; Liang, Y.; Lu, S.; Wang, Y.; Saied, A.; Chen, W.; and Duan, N. 2023 · 2023
Cited alongside, same era.
Don’t make your llm an evaluation benchmark cheater
Zhou, K.; Zhu, Y.; Chen, Z.; Chen, W.; Zhao, W. X.; Chen, X.; Lin, Y.; Wen, J.-R.; and Han, J. 2023 · 2023
Cited alongside, same era.
Phi-3 technical report: A highly capable language model locally on your phone
Abdin, M.; Jacobs, S. A.; Awan, A. A.; Aneja, J.; Awadallah, A.; Awadalla, H.; Bach, N.; Bahree, A.; Bakhtiari, A.; Behl, H.; et al. 2024 · 2024
Cited alongside, same era.
Deepseek llm: Scaling open-source language models with longtermism
Bi, X.; Chen, D.; Chen, G.; Chen, S.; Dai, D.; Deng, C.; Ding, H.; Dong, K.; Du, Q.; Fu, Z.; et al. 2024 · 2024
Cited alongside, same era.
Measuring Massive Multitask Language Understanding
Hendrycks, D.; Burns, C.; Basart, S.; Zou, A.; Mazeika, M.; Song, D.; and Steinhardt, J. 2021a
Cited in the paper.
Measuring mathematical problem solving with the math dataset
Hendrycks, D.; Burns, C.; Kadavath, S.; Arora, A.; Basart, S.; Tang, E.; Song, D.; and Steinhardt, J. 2021b
Cited in the paper.
Llama: Open and efficient foundation language models
Touvron, H.; Lavril, T.; Izacard, G.; Martinet, X.; Lachaux, M.-A.; Lacroix, T.; Rozière, B.; Goyal, N.; Hambro, E.; Azhar, F.; et al. 2023a
Cited in the paper.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H.; Martin, L.; Stone, K.; Albert, P.; Almahairi, A.; Babaei, Y.; Bashlykov, N.; Batra, S.; Bhargava, P.; Bhosale, S.; et al. 2023b
Cited in the paper.
Young, A.; Chen, B.; Li, C.; Huang, C.; Zhang, G.; Zhang, G.; Li, H.; Zhu, J.; Chen, J.; Chang, J.; et al. 2024 · 2024
Closest in time.
Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
Yue, X.; Ni, Y.; Zhang, K.; Zheng, T.; Liu, R.; Zhang, G.; Stevens, S.; Jiang, D.; Ren, W.; Sun, Y.; et al. 2024 · 2024
Closest in time.
Cmmmu: A chinese massive multi-discipline multimodal understanding benchmark
Zhang, G.; Du, X.; Chen, B.; Liang, Y.; Luo, T.; Zheng, T.; Zhu, K.; Cheng, Y.; Xu, C.; Guo, S.; et al. 2024 · 2024
Closest in time.
Judging llm-as-a-judge with mt-bench and chatbot arena
Zheng, L.; Chiang, W.-L.; Sheng, Y.; Zhuang, S.; Wu, Z.; Zhuang, Y.; Lin, Z.; Li, Z.; Li, D.; Xing, E.; et al. 2024 · 2024
Closest in time.