Fetching the paper…
Reading the bibliography…
New Natural Langauge Process~(NLP) benchmarks are urgently needed to align with the rapid development of large language models (LLMs).
HellaSwag: Can a machine really finish your sentence?
Zellers, R.; Holtzman, A.; Bisk, Y.; Farhadi, A.; and Choi, Y. 2019 · 1905
Earlier work this paper cites.
Cosmos QA: Machine reading comprehension with contextual commonsense reasoning
Huang, L.; Bras, R. L.; Bhagavatula, C.; and Choi, Y. 2019 · 1909
Earlier work this paper cites.
Neural Correction Model for Open-Domain Named Entity Recognition
Zhu, M.; Deng, Z.; Xiong, W.; Yu, M.; Zhang, M.; and Wang, W. Y. 2019 · 1909
Earlier work this paper cites.
Piqa: Reasoning about physical commonsense in natural language
Bisk, Y.; Zellers, R.; Gao, J.; Choi, Y.; et al. 2020 · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C.; Shazeer, N.; Roberts, A.; Lee, K.; Narang, S.; Matena, M.; Zhou, Y.; Li, W.; and Liu, P. J. 2020 · 2020
Earlier work this paper cites.
Ext5: Towards extreme multi-task scaling for transfer learning
Aribandi, V.; Tay, Y.; Schuster, T.; Rao, J.; Zheng, H. S.; Mehta, S. V.; Zhuang, H.; Tran, V. Q.; Bahri, D.; Ni, J.; et al. 2021 · 2021
Earlier work this paper cites.
Measuring coding challenge competence with apps
Hendrycks, D.; Basart, S.; Kadavath, S.; Mazeika, M.; Arora, A.; Guo, E.; Burns, C.; Puranik, S.; He, H.; Song, D.; et al. 2021 · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Hu, E. J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; and Chen, W. 2021 · 2021
Earlier work this paper cites.
Constitutional AI: Harmlessness from AI Feedback
Bai, Y.; Kadavath, S.; Kundu, S.; Askell, A.; Kernion, J.; Jones, A.; Chen, A.; Goldie, A.; Mirhoseini, A.; McKinnon, C.; et al. 2022 · 2022
Earlier work this paper cites.
GPT-NeoX-20B: An Open-Source Autoregressive Language Model
Black, S.; Biderman, S.; Hallahan, E.; Anthony, Q.; Gao, L.; Golding, L.; He, H.; Leahy, C.; McDonell, K.; Phang, J.; Pieler, M.; Prashanth, U. S.; Purohit, S.; Reynolds, L.; Tow, J.; Wang, B.; and Weinbach, S. 2022 · 2022
Earlier work this paper cites.
Scaling instruction-finetuned language models
Chung, H. W.; Hou, L.; Longpre, S.; Zoph, B.; Tay, Y.; Fedus, W.; Li, E.; Wang, X.; Dehghani, M.; Brahma, S.; et al. 2022 · 2022
Earlier work this paper cites.
GLM: General Language Model Pretraining with Autoregressive Blank Infilling
Du, Z.; Qian, Y.; Liu, X.; Ding, M.; Qiu, J.; Yang, Z.; and Tang, J. 2022 · 2022
Earlier work this paper cites.
How does gpt obtain its ability? tracing emergent abilities of language models to their sources
Fu, Y.; Peng, H.; and Khot, T. 2022 · 2022
Earlier work this paper cites.
An empirical analysis of compute-optimal large language model training
Hoffmann, J.; Borgeaud, S.; Mensch, A.; Buchatskaya, E.; Cai, T.; Rutherford, E.; de Las Casas, D.; Hendricks, L. A.; Welbl, J.; Clark, A.; et al. 2022 · 2022
Earlier work this paper cites.
Holistic evaluation of language models
Liang, P.; Bommasani, R.; Lee, T.; Tsipras, D.; Soylu, D.; Yasunaga, M.; Zhang, Y.; Narayanan, D.; Wu, Y.; Kumar, A.; et al. 2022 · 2022
Earlier work this paper cites.
Crosslingual Generalization through Multitask Finetuning
Muennighoff, N.; Wang, T.; Sutawika, L.; Roberts, A.; Biderman, S.; Scao, T. L.; Bari, M. S.; Shen, S.; Yong, Z.-X.; Schoelkopf, H.; Tang, X.; Radev, D.; Aji, A. F.; Almubarak, K.; Albanie, S.; Alyafeai, Z.; Webson, A.; Raff, E.; and Raffel, C. 2022 · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L.; Wu, J.; Jiang, X.; Almeida, D.; Wainwright, C.; Mishkin, P.; Zhang, C.; Agarwal, S.; Slama, K.; Ray, A.; et al. 2022 · 2022
Cited alongside, same era.
Discovering Language Model Behaviors with Model-Written Evaluations
Perez, E.; Ringer, S.; Lukošiūtė, K.; Nguyen, K.; Chen, E.; Heiner, S.; Pettit, C.; Olsson, C.; Kundu, S.; Kadavath, S.; et al. 2022 · 2022
Cited alongside, same era.
A survey of evaluation metrics used for NLG systems
Sai, A. B.; Mohankumar, A. K.; and Khapra, M. M. 2022 · 2022
Cited alongside, same era.
Bloom: A 176b-parameter open-access multilingual language model
C-Eval: A Multi-Level Multi-Discipline Chinese Evaluation Suite for Foundation Models
Huang, Y.; Bai, Y.; Zhu, Z.; Zhang, J.; Zhang, J.; Su, T.; Liu, J.; Lv, C.; Zhang, Y.; Lei, J.; et al. 2023 · 2023
Closest in time.
OpenAssistant Conversations – Democratizing Large Language Model Alignment
Köpf, A.; Kilcher, Y.; von Rütte, D.; Anagnostidis, S.; Tam, Z.-R.; Stevens, K.; Barhoum, A.; Duc, N. M.; Stanley, O.; Nagyfi, R.; ES, S.; Suri, S.; Glushkov, D.; Dantuluri, A.; Maguire, A.; Schuhmann, C.; Nguyen, H.; and Mattick, A. 2023 · 2023
Closest in time.
CMMLU: Measuring massive multitask language understanding in Chinese
Li, H.; Zhang, Y.; Koto, F.; Yang, Y.; Zhao, H.; Gong, Y.; Duan, N.; and Baldwin, T. 2023 · 2023
Closest in time.
Liu, C.; Jin, R.; Ren, Y.; Yu, L.; Dong, T.; Peng, X.; Zhang, S.; Peng, J.; Zhang, P.; Lyu, Q.; et al. 2023 · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Scao, T. L.; Fan, A.; Akiki, C.; Pavlick, E.; Ilić, S.; Hesslow, D.; Castagné, R.; Luccioni, A. S.; Yvon, F.; Gallé, M.; et al. 2022 · 2022
Cited alongside, same era.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Srivastava, A.; Rastogi, A.; Rao, A.; Shoeb, A. A. M.; Abid, A.; Fisch, A.; Brown, A. R.; Santoro, A.; Gupta, A.; Garriga-Alonso, A.; et al. 2022 · 2022
Cited alongside, same era.
Challenging BIG-Bench tasks and whether chain-of-thought can solve them
Suzgun, M.; Scales, N.; Schärli, N.; Gehrmann, S.; Tay, Y.; Chung, H. W.; Chowdhery, A.; Le, Q. V.; Chi, E. H.; Zhou, D.; et al. 2022 · 2022
Cited alongside, same era.
Emergent abilities of large language models
Wei, J.; Tay, Y.; Bommasani, R.; Raffel, C.; Zoph, B.; Borgeaud, S.; Yogatama, D.; Bosma, M.; Zhou, D.; Metzler, D.; et al. 2022 · 2022
Cited alongside, same era.
Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling
Biderman, S.; Schoelkopf, H.; Anthony, Q.; Bradley, H.; O’Brien, K.; Hallahan, E.; Khan, M. A.; Purohit, S.; Prashanth, U. S.; Raff, E.; Skowron, A.; Sutawika, L.; and van der Wal, O. 2023 · 2023
Cited alongside, same era.
Sparks of artificial general intelligence: Early experiments with gpt-4
Bubeck, S.; Chandrasekaran, V.; Eldan, R.; Gehrke, J.; Horvitz, E.; Kamar, E.; Lee, P.; Lee, Y. T.; Li, Y.; Lundberg, S.; et al. 2023 · 2023
Cited alongside, same era.
Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90%* ChatGPT Quality
Chiang, W.-L.; Li, Z.; Lin, Z.; Sheng, Y.; Wu, Z.; Zhang, H.; Zheng, L.; Zhuang, S.; Zhuang, Y.; Gonzalez, J. E.; Stoica, I.; and Xing, E. P. 2023 · 2023
Cited alongside, same era.
Free Dolly: Introducing the World’s First Truly Open Instruction-Tuned LLM
Conover, M.; Hayes, M.; Mathur, A.; Meng, X.; Xie, J.; Wan, J.; Shah, S.; Ghodsi, A.; Wendell, P.; Zaharia, M.; and Xin, R. 2023 · 2023
Cited alongside, same era.
Lextreme: A multi-lingual and multi-task benchmark for the legal domain
Niklaus, J.; Matoshi, V.; Rani, P.; Galassi, A.; Stürmer, M.; and Chalkidis, I. 2023 · 2023
Closest in time.
StableLM: Stability AI Language Models
StabilityAI. 2023 · 2023
Closest in time.
Stanford Alpaca: An Instruction-following LLaMA model
Taori, R.; Gulrajani, I.; Zhang, T.; Dubois, Y.; Li, X.; Guestrin, C.; Liang, P.; and Hashimoto, T. B. 2023 · 2023
Closest in time.
Llama: Open and efficient foundation language models
Touvron, H.; Lavril, T.; Izacard, G.; Martinet, X.; Lachaux, M.-A.; Lacroix, T.; Rozière, B.; Goyal, N.; Hambro, E.; Azhar, F.; et al. 2023 · 2023
Closest in time.
DoctorGLM: Fine-tuning your Chinese Doctor is not a Herculean Task
Xiong, H.; Wang, S.; Zhu, Y.; Zhao, Z.; Liu, Y.; Huang, L.; Wang, Q.; and Shen, D. 2023 · 2023
Closest in time.
Baize: An Open-Source Chat Model with Parameter-Efficient Tuning on Self-Chat Data
Xu, C.; Guo, D.; Duan, N.; and McAuley, J. 2023 · 2023
Closest in time.
GLM-130B: An Open Bilingual Pre-trained Model
Zeng, A.; Liu, X.; Du, Z.; Wang, Z.; Lai, H.; Ding, M.; Yang, Z.; Xu, Y.; Zheng, W.; Xia, X.; Tam, W. L.; Ma, Z.; Xue, Y.; Zhai, J.; Chen, W.; Liu, Z.; Zhang, P.; Dong, Y.; and Tang, J. 2023 · 2023
Closest in time.
Measuring Massive Multitask Chinese Understanding
Zeng, H. 2023 · 2023
Closest in time.
A survey of large language models
Zhao, W. X.; Zhou, K.; Li, J.; Tang, T.; Wang, X.; Hou, Y.; Min, Y.; Zhang, B.; Zhang, J.; Dong, Z.; et al. 2023 · 2023
Closest in time.
AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Zhong, W.; Cui, R.; Guo, Y.; Liang, Y.; Lu, S.; Wang, Y.; Saied, A.; Chen, W.; and Duan, N. 2023 · 2023
Closest in time.
A comprehensive survey on pretrained foundation models: A history from bert to chatgpt
Zhou, C.; Li, Q.; Li, C.; Yu, J.; Liu, Y.; Wang, G.; Zhang, K.; Ji, C.; Yan, Q.; He, L.; et al. 2023 · 2023
Closest in time.