Fetching the paper…
Reading the bibliography…
Large Language Models are traditionally finetuned on large instruction datasets.
Mathqa: Towards interpretable math word problem solving with operation-based formalisms
Amini, A., Gabriel, S., Lin, S., Koncel-Kedziorski, R., Choi, Y., and Hajishirzi, H. (2019) · 1905
Earlier work this paper cites.
Hellaswag: Can a machine really finish your sentence?
Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y. (2019) · 1905
Earlier work this paper cites.
WINOGRANDE: an adversarial winograd schema challenge at scale
Sakaguchi, K., Bras, R. L., Bhagavatula, C., and Choi, Y. (2019) · 1907
Earlier work this paper cites.
Pubmedqa: A dataset for biomedical research question answering
Jin, Q., Dhingra, B., Liu, Z., Cohen, W. W., and Lu, X. (2019) · 1909
Earlier work this paper cites.
Unifiedqa: Crossing format boundaries with a single qa system
Khashabi, D., Min, S., Khot, T., Sabharwal, A., Tafjord, O., Clark, P., and Hajishirzi, H. (2020) · 2005
Earlier work this paper cites.
Logiqa: A challenge dataset for machine reading comprehension with logical reasoning
Liu, J., Cui, L., Liu, H., Huang, D., Wang, Y., and Zhang, Y. (2020) · 2007
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J. (2020) · 2009
Earlier work this paper cites.
Choice of plausible alternatives: An evaluation of commonsense causal reasoning
Roemmele, M., Beja, C. A., and Gordon, A. S. (2011) · 2011
Earlier work this paper cites.
The winograd schema challenge
Levesque, H., Davis, E., and Morgenstern, L. (2012) · 2012
Earlier work this paper cites.
The lambada dataset: Word prediction requiring a broad discourse context
Paperno, D., Kruszewski, G., Lazaridou, A., Pham, Q. N., Bernardi, R., Pezzelle, S., Baroni, M., Boleda, G., and Fernández, R. (2016) · 2016
Earlier work this paper cites.
SQuAD: 100,000+ Questions for Machine Comprehension of Text
Rajpurkar, P., Zhang, J., Lopyrev, K., and Liang, P. (2016) · 2016
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F. (2017) · 2017
Earlier work this paper cites.
Universal language model fine-tuning for text classification
Howard, J. and Ruder, S. (2018) · 2018
Earlier work this paper cites.
Can a suit of armor conduct electricity? a new dataset for open book question answering
Mihaylov, T., Clark, P., Khot, T., and Sabharwal, A. (2018) · 2018
Earlier work this paper cites.
Duorc: Towards complex language understanding with paraphrased reading comprehension
Saha, A., Aralikatte, R., Khapra, M. M., and Sankaranarayanan, K. (2018) · 2018
Earlier work this paper cites.
A simple method for commonsense reasoning
Trinh, T. H. and Le, Q. V. (2018) · 2018
Earlier work this paper cites.
Yu, T., Zhang, R., Yang, K., Yasunaga, M., Wang, D., Li, Z., Ma, J., Li, I., Yao, Q., Roman, S., et al. (2018) · 2018
Earlier work this paper cites.
Boolq: Exploring the surprising difficulty of natural yes/no questions
Clark, C., Lee, K., Ming-Wei Chang, T. K., Collins, M., and Toutanova, K. (2019) · 2019
Earlier work this paper cites.
Natural questions: a benchmark for question answering research
Kwiatkowski, T., Palomaki, J., Redfield, O., Collins, M., Parikh, A., Alberti, C., Epstein, D., Polosukhin, I., Kelcey, M., Devlin, J., Lee, K., Toutanova, K. N., Jones, L., Chang, M.-W., Dai, A., Uszkoreit, J., Le, Q., and Petrov, S. (2019) · 2019
Cited alongside, same era.
Superglue: A stickier benchmark for general-purpose language understanding systems
Wang, A., Pruksachatkun, Y., Nangia, N., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. (2019) · 2019
Cited alongside, same era.
Piqa: Reasoning about physical commonsense in natural language
Bisk, Y., Zellers, R., Gao, J., Choi, Y., et al. (2020) · 2020
Cited alongside, same era.
Training verifiers to solve math word problems
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., Hesse, C., and Schulman, J. (2021) · 2021
Cited alongside, same era.
A dataset of information-seeking questions and answers anchored in research papers
Qlora: Efficient finetuning of quantized llms
Dettmers, T., Pagnoni, A., Holtzman, A., and Zettlemoyer, L. (2023) · 2023
Closest in time.
Alpacafarm: A simulation framework for methods that learn from human feedback
Dubois, Y., Li, X., Taori, R., Zhang, T., Gulrajani, I., Ba, J., Guestrin, C., Liang, P., and Hashimoto, T. B. (2023) · 2023
Closest in time.
The false promise of imitating proprietary llms
Gudibande, A., Wallace, E., Snell, C., Geng, X., Liu, H., Abbeel, P., Levine, S., and Song, D. (2023) · 2023
Closest in time.
Human feedback is not gold standard
Hosking, T., Blunsom, P., and Bartolo, M. (2023) · 2023
Closest in time.
Camels in a changing climate: Enhancing lm adaptation with tulu 2
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dasigi, P., Lo, K., Beltagy, I., Cohan, A., Smith, N. A., and Gardner, M. (2021) · 2021
Cited alongside, same era.
Measuring mathematical problem solving with the math dataset
Hendrycks, D., Burns, C., Kadavath, S., Arora, A., Basart, S., Tang, E., Song, D., and Steinhardt, J. (2021) · 2021
Cited alongside, same era.
Train short, test long: Attention with linear biases enables input length extrapolation
Press, O., Smith, N. A., and Lewis, M. (2021) · 2021
Cited alongside, same era.
Multitask prompted training enables zero-shot task generalization
Sanh, V., Webson, A., Raffel, C., Bach, S. H., Sutawika, L., Alyafeai, Z., Chaffin, A., Stiegler, A., Scao, T. L., Raja, A., et al. (2021) · 2021
Cited alongside, same era.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Bai, Y., Jones, A., Ndousse, K., Askell, A., Chen, A., DasSarma, N., Drain, D., Fort, S., Ganguli, D., Henighan, T., et al. (2022) · 2022
Cited alongside, same era.
Gpt-neox-20b: An open-source autoregressive language model
Black, S., Biderman, S., Hallahan, E., Anthony, Q., Gao, L., Golding, L., He, H., Leahy, C., McDonell, K., Phang, J., et al. (2022) · 2022
Cited alongside, same era.
SummScreen: A dataset for abstractive screenplay summarization
Chen, M., Chu, Z., Wiseman, S., and Gimpel, K. (2022) · 2022
Cited alongside, same era.
Scaling instruction-finetuned language models
Chung, H. W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., Li, E., Wang, X., Dehghani, M., Brahma, S., et al. (2022) · 2022
Cited alongside, same era.
Ivison, H., Wang, Y., Pyatkin, V., Lambert, N., Peters, M., Dasigi, P., Jang, J., Wadden, D., Smith, N. A., Beltagy, I., and Hajishirzi, H. (2023) · 2023
Closest in time.
Alpacaeval: An automatic evaluator of instruction-following models
Li, X., Zhang, T., Dubois, Y., Taori, R., Gulrajani, I., Guestrin, C., Liang, P., and Hashimoto, T. B. (2023) · 2023
Closest in time.
The flan collection: Designing data and methods for effective instruction tuning
Longpre, S., Hou, L., Vu, T., Webson, A., Chung, H. W., Tay, Y., Zhou, D., Le, Q. V., Zoph, B., Wei, J., et al. (2023) · 2023
Closest in time.
Blazingly fast llm evaluation for in-context learning
MosaicML NLP Team (2023a) · 2023
Closest in time.
Introducing mpt-30b: Raising the bar for open-source foundation models
MosaicML NLP Team (2023c) · 2023
Closest in time.
Introducing mpt-7b: A new standard for open-source, commercially usable llms
MosaicML NLP Team (2023d) · 2023
Closest in time.
Llm evaluation scores
MosaicML NLP Team (2023e) · 2023
Closest in time.
Stanford alpaca: An instruction-following llama model
Taori, R., Gulrajani, I., Zhang, T., Dubois, Y., Li, X., Guestrin, C., Liang, P., and Hashimoto, T. B. (2023) · 2023
Closest in time.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al. (2023) · 2023
Closest in time.
How far can camels go? exploring the state of instruction tuning on open resources
Wang, Y., Ivison, H., Dasigi, P., Hessel, J., Khot, T., Chandu, K. R., Wadden, D., MacMillan, K., Smith, N. A., Beltagy, I., et al. (2023) · 2023
Closest in time.
Baize: An open-source chat model with parameter-efficient tuning on self-chat data
Xu, C., Guo, D., Duan, N., and McAuley, J. (2023) · 2023
Closest in time.
Judging llm-as-a-judge with mt-bench and chatbot arena
Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E., et al. (2023) · 2023
Closest in time.
Lima: Less is more for alignment
Zhou, C., Liu, P., Xu, P., Iyer, S., Sun, J., Mao, Y., Ma, X., Efrat, A., Yu, P., Yu, L., Zhang, S., Ghosh, G., Lewis, M., Zettlemoyer, L., and Levy, O. (2023) · 2023
Closest in time.