Fetching the paper…
Reading the bibliography…
Instruction Fine-Tuning (IFT) significantly enhances the zero-shot capabilities of pretrained Large Language Models (LLMs).
CLUTRR: A Diagnostic Benchmark for Inductive Reasoning from Text
Sinha, K.; Sodhani, S.; Dong, J.; Pineau, J.; and Hamilton, W. L. 2019 · 1908
Earlier work this paper cites.
Reasoning about a rule
Wason, P. C. 1968 · 1968
Earlier work this paper cites.
Psychology of reasoning: Structure and content , volume 86
Wason, P. C.; and Johnson-Laird, P. N. 1972 · 1972
Earlier work this paper cites.
Language Models are Few-Shot Learners
Brown, T. B.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; Agarwal, S.; Herbert-Voss, A.; Krueger, G.; Henighan, T.; Child, R.; Ramesh, A.; Ziegler, D. M.; Wu, J.; Winter, C.; Hesse, C.; Chen, M.; Sigler, E.; Litwin, M.; Gray, S.; Chess, B.; Clark, J.; Berner, C.; McCandlish, S.; Radford, A.; Sutskever, I.; and Amodei, D. 2020 · 2005
Earlier work this paper cites.
Towards ai-complete question answering: A set of prerequisite toy tasks
Weston, J.; Bordes, A.; Chopra, S.; Rush, A. M.; Van Merriënboer, B.; Joulin, A.; and Mikolov, T. 2015 · 2015
Earlier work this paper cites.
MAWPS: A Math Word Problem Repository
Koncel-Kedziorski, R.; Roy, S.; Amini, A.; Kushman, N.; and Hajishirzi, H. 2016 · 2016
Earlier work this paper cites.
A Diverse Corpus for Evaluating and Developing English Math Word Problem Solvers
Miao, S.-y.; Liang, C.-C.; and Su, K.-Y. 2020 · 2020
Earlier work this paper cites.
The child as hacker: building more human-like models of learning
Rule, J. S. 2020 · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems
Cobbe, K.; Kosaraju, V.; Bavarian, M.; Chen, M.; Jun, H.; Kaiser, L.; Plappert, M.; Tworek, J.; Hilton, J.; Nakano, R.; et al. 2021 · 2021
Earlier work this paper cites.
Are NLP models really able to solve simple math word problems?
Patel, A.; Bhattamishra, S.; and Goyal, N. 2021 · 2021
Earlier work this paper cites.
Scaling Instruction-Finetuned Language Models
Chung, H. W.; Hou, L.; Longpre, S.; Zoph, B.; Tay, Y.; Fedus, W.; Li, Y.; Wang, X.; Dehghani, M.; Brahma, S.; Webson, A.; Gu, S. S.; Dai, Z.; Suzgun, M.; Chen, X.; Chowdhery, A.; Castro-Ros, A.; Pellat, M.; Robinson, K.; Valter, D.; Narang, S.; Mishra, G.; Yu, A.; Zhao, V.; Huang, Y.; Dai, A.; Yu, H.; Petrov, S.; Chi, E. H.; Dean, J.; Devlin, J.; Roberts, A.; Zhou, D.; Le, Q. V.; and Wei, J. 2022 · 2022
Earlier work this paper cites.
How does GPT Obtain its Ability? Tracing Emergent Abilities of Language Models to their Sources
Fu, H., Yao; Peng; and Khot, T. 2022 · 2022
Earlier work this paper cites.
Explanations from Large Language Models Make Small Reasoners Better
Li, S.; Chen, J.; Shen, Y.; Chen, Z.; Zhang, X.; Li, Z.; Wang, H.; Qian, J.; Peng, B.; Mao, Y.; Chen, W.; and Yan, X. 2022 · 2022
Earlier work this paper cites.
Language Models of Code are Few-Shot Commonsense Learners
Madaan, A.; Zhou, S.; Alon, U.; Yang, Y.; and Neubig, G. 2022 · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Ouyang, L.; Wu, J.; Jiang, X.; Almeida, D.; Wainwright, C. L.; Mishkin, P.; Zhang, C.; Agarwal, S.; Slama, K.; Ray, A.; Schulman, J.; Hilton, J.; Kelton, F.; Miller, L.; Simens, M.; Askell, A.; Welinder, P.; Christiano, P.; Leike, J.; and Lowe, R. 2022 · 2022
Earlier work this paper cites.
Claude 2
Anthropic. 2023 · 2023
Cited alongside, same era.
Bai, J.; Bai, S.; Chu, Y.; Cui, Z.; Dang, K.; Deng, X.; Fan, Y.; Ge, W.; Han, Y.; Huang, F.; Hui, B.; Ji, L.; Li, M.; Lin, J.; Lin, R.; Liu, D.; Liu, G.; Lu, C.; Lu, K.; Ma, J.; Men, R.; Ren, X.; Ren, X.; Tan, C.; Tan, S.; Tu, J.; Wang, P.; Wang, S.; Wang, W.; Wu, S.; Xu, B.; Xu, J.; Yang, A.; Yang, H.; Yang, J.; Yang, S.; Yao, Y.; Yu, B.; Yuan, H.; Yuan, Z.; Zhang, J.; Zhang, X.; Zhang, Y.; Zhang, Z.; Zhou, C.; Zhou, J.; Zhou, X.; and Zhu, T. 2023 · 2023
Cited alongside, same era.
AlpaGasus: Training A Better Alpaca with Fewer Data
Chen, L.; Li, S.; Yan, J.; Wang, H.; Gunaratna, K.; Yadav, V.; Tang, Z.; Srinivasan, V.; Zhou, T.; Huang, H.; and Jin, H. 2023 · 2023
Cited alongside, same era.
Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90%* ChatGPT Quality
Chiang, W.-L.; Li, Z.; Lin, Z.; Sheng, Y.; Wu, Z.; Zhang, H.; Zheng, L.; Zhuang, S.; Zhuang, Y.; Gonzalez, J. E.; Stoica, I.; and Xing, E. P. 2023 · 2023
Cited alongside, same era.
Simple Dataset Generation
Stanford Alpaca: An Instruction-following LLaMA model
Taori, R.; Gulrajani, I.; Zhang, T.; Dubois, Y.; Li, X.; Guestrin, C.; Liang, P.; and Hashimoto, T. B. 2023 · 2023
Later among the works it cites.
Code4Struct: Code Generation for Few-Shot Event Structure Prediction
Wang, X.; Li, S.; and Ji, H. 2023 · 2023
Later among the works it cites.
WizardLM: Empowering Large Language Models to Follow Complex Instructions
Xu, C.; Sun, Q.; Zheng, K.; Geng, X.; Zhao, P.; Feng, J.; Tao, C.; and Jiang, D. 2023 · 2023
Later among the works it cites.
Exploring the Limits of ChatGPT for Query or Aspect-based Text Summarization
Yang, X.; Li, Y.; Zhang, X.; Chen, H.; and Cheng, W. 2023 · 2023
Later among the works it cites.
Evaluating instruction-tuned large language models on code comprehension and generation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Fortes, A. 2023 · 2023
Cited alongside, same era.
Jiang, A. Q.; Sablayrolles, A.; Mensch, A.; Bamford, C.; Chaplot, D. S.; Casas, D. d. l.; Bressand, F.; Lengyel, G.; Lample, G.; Saulnier, L.; et al. 2023 · 2023
Cited alongside, same era.
Holistic Evaluation of Language Models
Liang, P.; Bommasani, R.; Lee, T.; Tsipras, D.; Soylu, D.; Yasunaga, M.; Zhang, Y.; Narayanan, D.; Wu, Y.; Kumar, A.; Newman, B.; Yuan, B.; Yan, B.; Zhang, C.; Cosgrove, C.; Manning, C. D.; Ré, C.; Acosta-Navas, D.; Hudson, D. A.; Zelikman, E.; Durmus, E.; Ladhak, F.; Rong, F.; Ren, H.; Yao, H.; Wang, J.; Santhanam, K.; Orr, L.; Zheng, L.; Yuksekgonul, M.; Suzgun, M.; Kim, N.; Guha, N.; Chatterji, N.; Khattab, O.; Henderson, P.; Huang, Q.; Chi, R.; Xie, S. M.; Santurkar, S.; Ganguli, S.; Hashimoto, T.; Icard, T.; Zhang, T.; Chaudhary, V.; Wang, W.; Li, X.; Mai, Y.; Zhang, Y.; and Koreeda, Y. 2023 · 2023
Cited alongside, same era.
The Flan Collection: Designing Data and Methods for Effective Instruction Tuning
Longpre, S.; Hou, L.; Vu, T.; Webson, A.; Chung, H. W.; Tay, Y.; Zhou, D.; Le, Q. V.; Zoph, B.; Wei, J.; and Roberts, A. 2023 · 2023
Cited alongside, same era.
At Which Training Stage Does Code Data Help LLMs Reasoning?
Ma, Y.; Liu, Y.; Yu, Y.; Zhang, Y.; Jiang, Y.; Wang, C.; and Li, S. 2023 · 2023
Cited alongside, same era.
Introducing ChatGPT
OpenAI. 2022 · 2023
Cited alongside, same era.
OpenAI. 2023 · 2023
Cited alongside, same era.
Peng, B.; Li, C.; He, P.; Galley, M.; and Gao, J. 2023 · 2023
Cited alongside, same era.
Yuan, Z.; Liu, J.; Zi, Q.; Liu, M.; Peng, X.; and Lou, Y. 2023 · 2023
Later among the works it cites.
MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning
Yue, X.; Qu, X.; Zhang, G.; Fu, Y.; Huang, W.; Sun, H.; Su, Y.; and Chen, W. 2023 · 2023
Later among the works it cites.
Enhancing Small Medical Learners with Privacy-preserving Contextual Prompting
Zhang, X.; Li, S.; Yang, X.; Tian, C.; Qin, Y.; and Petzold, L. R. 2023 · 2023
Later among the works it cites.
Llama 3 Model Card
AI@Meta. 2024 · 2024
Closest in time.
A Survey on Data Selection for Language Models
Albalak, A.; Elazar, Y.; Xie, S. M.; Longpre, S.; Lambert, N.; Wang, X.; Muennighoff, N.; Hou, B.; Pan, L.; Jeong, H.; Raffel, C.; Chang, S.; Hashimoto, T.; and Wang, W. Y. 2024 · 2024
Closest in time.
A Survey on Large Language Models for Critical Societal Domains: Finance, Healthcare, and Law
Chen, Z. Z.; Ma, J.; Zhang, X.; Hao, N.; Yan, A.; Nourbakhsh, A.; Yang, X.; McAuley, J.; Petzold, L.; and Wang, W. Y. 2024 · 2024
Closest in time.
DeepSeek-Coder: When the Large Language Model Meets Programming – The Rise of Code Intelligence
Guo, D.; Zhu, Q.; Yang, D.; Xie, Z.; Dong, K.; Zhang, W.; Chen, G.; Bi, X.; Wu, Y.; Li, Y. K.; Luo, F.; Xiong, Y.; and Liang, W. 2024 · 2024
Closest in time.
Testing the general deductive reasoning capacity of large language models using ood examples
Saparov, A.; Pang, R. Y.; Padmakumar, V.; Joshi, N.; Kazemi, M.; Kim, N.; and He, H. 2024 · 2024
Closest in time.
Gemma: Open models based on gemini research and technology
Team, G.; Mesnard, T.; Hardin, C.; Dadashi, R.; Bhupatiraju, S.; Pathak, S.; Sifre, L.; Rivière, M.; Kale, M. S.; Love, J.; et al. 2024 · 2024
Closest in time.
AlpaCare:Instruction-tuned Large Language Models for Medical Application
Zhang, X.; Tian, C.; Yang, X.; Chen, L.; Li, Z.; and Petzold, L. R. 2024 · 2024
Closest in time.