Fetching the paper…
Reading the bibliography…
Large language models (LLMs) can understand human instructions, showing their potential for pragmatic applications beyond traditional NLP tasks.
Language models are few-shot learners
Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J. D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. 2020 · 1901
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D.; Burns, C.; Basart, S.; Zou, A.; Mazeika, M.; Song, D.; and Steinhardt, J. 2020 · 2009
Earlier work this paper cites.
CN-DBpedia: A never-ending Chinese knowledge extraction system
Xu, B.; Xu, Y.; Liang, J.; Xie, C.; Liang, B.; Cui, W.; and Xiao, Y. 2017 · 2017
Earlier work this paper cites.
Evaluating large language models trained on code
Chen, M.; Tworek, J.; Jun, H.; Yuan, Q.; Pinto, H. P. d. O.; Kaplan, J.; Edwards, H.; Burda, Y.; Joseph, N.; Brockman, G.; et al. 2021 · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems
Cobbe, K.; Kosaraju, V.; Bavarian, M.; Chen, M.; Jun, H.; Kaiser, L.; Plappert, M.; Tworek, J.; Hilton, J.; Nakano, R.; et al. 2021 · 2021
Earlier work this paper cites.
Controllable Text Generation with Language Constraints
Chen, H.; Li, H.; Chen, D.; and Narasimhan, K. 2022 · 2022
Earlier work this paper cites.
Palm: Scaling language modeling with pathways
Chowdhery, A.; Narang, S.; Devlin, J.; Bosma, M.; Mishra, G.; Roberts, A.; Barham, P.; Chung, H. W.; Sutton, C.; Gehrmann, S.; et al. 2022 · 2022
Earlier work this paper cites.
Unnatural instructions: Tuning language models with (almost) no human labor
Honovich, O.; Scialom, T.; Levy, O.; and Schick, T. 2022 · 2022
Earlier work this paper cites.
Holistic evaluation of language models
Liang, P.; Bommasani, R.; Lee, T.; Tsipras, D.; Soylu, D.; Yasunaga, M.; Zhang, Y.; Narayanan, D.; Wu, Y.; Kumar, A.; et al. 2022 · 2022
Earlier work this paper cites.
L-Eval: Instituting Standardized Evaluation for Long Context Language Models
An, C.; Gong, S.; Zhong, M.; Li, M.; Zhang, J.; Kong, L.; and Qiu, X. 2023 · 2023
Earlier work this paper cites.
INSTRUCTEVAL: Towards Holistic Evaluation of Instruction-Tuned Large Language Models
Chia, Y. K.; Hong, P.; Bing, L.; and Poria, S. 2023 · 2023
Earlier work this paper cites.
Efficient and Effective Text Encoding for Chinese LLaMA and Alpaca
Cui, Y.; Yang, Z.; and Yao, X. 2023 · 2023
Cited alongside, same era.
Enhancing Chat Language Models by Scaling High-quality Instructional Conversations
Ding, N.; Chen, Y.; Xu, B.; Qin, Y.; Zheng, Z.; Hu, S.; Liu, Z.; Sun, M.; and Zhou, B. 2023 · 2023
Cited alongside, same era.
Alpacafarm: A simulation framework for methods that learn from human feedback
Dubois, Y.; Li, X.; Taori, R.; Zhang, T.; Gulrajani, I.; Ba, J.; Guestrin, C.; Liang, P.; and Hashimoto, T. B. 2023 · 2023
Cited alongside, same era.
Xiezhi: An Ever-Updating Benchmark for Holistic Domain Knowledge Evaluation
Gu, Z.; Zhu, X.; Ye, H.; Zhang, L.; Wang, J.; Jiang, S.; Xiong, Z.; Li, Z.; He, Q.; Xu, R.; et al. 2023 · 2023
Cited alongside, same era.
Auto-GPT: An Autonomous GPT-4 Experiment
Richards, T. B. 2023 · 2023
Closest in time.
Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models
Srivastava, A.; Rastogi, A.; Rao, A.; Shoeb, A. A. M.; Abid, A.; Fisch, A.; Brown, A. R.; Santoro, A.; Gupta, A.; Garriga-Alonso, A.; et al. 2023 · 2023
Closest in time.
Stanford alpaca: An instruction-following llama model
Taori, R.; Gulrajani, I.; Zhang, T.; Dubois, Y.; Li, X.; Guestrin, C.; Liang, P.; and Hashimoto, T. B. 2023 · 2023
Closest in time.
InternLM: A Multilingual Language Model with Progressively Enhanced Capabilities
Team, I. 2023 · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H.; Martin, L.; Stone, K.; Albert, P.; Almahairi, A.; Babaei, Y.; Bashlykov, N.; Batra, S.; Bhargava, P.; Bhosale, S.; et al. 2023 · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Guo, B.; Zhang, X.; Wang, Z.; Jiang, M.; Nie, J.; Ding, Y.; Yue, J.; and Wu, Y. 2023 · 2023
Cited alongside, same era.
C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models
Huang, Y.; Bai, Y.; Zhu, Z.; Zhang, J.; Zhang, J.; Su, T.; Liu, J.; Lv, C.; Zhang, Y.; Lei, J.; et al. 2023 · 2023
Cited alongside, same era.
BELLE: Be Everyone’s Large Language model Engine
Ji, Y.; Deng, Y.; Gong, Y.; Peng, Y.; Niu, Q.; Ma, B.; and Li, X. 2023 · 2023
Cited alongside, same era.
How Long Can Open-Source LLMs Truly Promise on Context Length?
Li*, D.; Shao*, R.; Xie, A.; Sheng, Y.; Zheng, L.; Gonzalez, J. E.; Stoica, I.; Ma, X.; ; and Zhang, H. 2023 · 2023
Cited alongside, same era.
WizardCoder: Empowering Code Large Language Models with Evol-Instruct
Luo, Z.; Xu, C.; Zhao, P.; Sun, Q.; Geng, X.; Hu, W.; Tao, C.; Ma, J.; Lin, Q.; and Jiang, D. 2023 · 2023
Cited alongside, same era.
Orca: Progressive learning from complex explanation traces of gpt-4
Mukherjee, S.; Mitra, A.; Jawahar, G.; Agarwal, S.; Palangi, H.; and Awadallah, A. 2023 · 2023
Cited alongside, same era.
ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs
Qin, Y.; Liang, S.; Ye, Y.; Zhu, K.; Yan, L.; Lu, Y.; Lin, Y.; Cong, X.; Tang, X.; Qian, B.; et al. 2023 · 2023
Cited alongside, same era.
Camel: Communicative agents for” mind” exploration of large scale language model society
Li, G.; Hammoud, H. A. A. K.; Itani, H.; Khizbullin, D.; and Ghanem, B. 2023a
Cited in the paper.
Yu, J.; Wang, X.; Tu, S.; Cao, S.; Zhang-Li, D.; Lv, X.; Peng, H.; Yao, Z.; Zhang, X.; Li, H.; et al. 2023 · 2023
Closest in time.
GLM-130B: An Open Bilingual Pre-trained Model
Zeng, A.; Liu, X.; Du, Z.; Wang, Z.; Lai, H.; Ding, M.; Yang, Z.; Xu, Y.; Zheng, W.; Xia, X.; Tam, W. L.; Ma, Z.; Xue, Y.; Zhai, J.; Chen, W.; Liu, Z.; Zhang, P.; Dong, Y.; and Tang, J. 2023 · 2023
Closest in time.
TableGPT: Towards Unifying Tables, Nature Language and Commands into One GPT
Zha, L.; Zhou, J.; Li, L.; Wang, R.; Huang, Q.; Yang, S.; Yuan, J.; Su, C.; Li, X.; Su, A.; et al. 2023 · 2023
Closest in time.
Judging LLM-as-a-judge with MT-Bench and Chatbot Arena
Zheng, L.; Chiang, W.-L.; Sheng, Y.; Zhuang, S.; Wu, Z.; Zhuang, Y.; Lin, Z.; Li, Z.; Li, D.; Xing, E. P.; Zhang, H.; Gonzalez, J. E.; and Stoica, I. 2023 · 2023
Closest in time.
Agieval: A human-centric benchmark for evaluating foundation models
Zhong, W.; Cui, R.; Guo, Y.; Liang, Y.; Lu, S.; Wang, Y.; Saied, A.; Chen, W.; and Duan, N. 2023 · 2023
Closest in time.