Fetching the paper…
Reading the bibliography…
Instruction tuning is widely recognized as a key technique for building generalist language models, which has attracted the attention of researchers and the public with the release of InstructGPT~\citep{ouyang2022training} and ChatGPT\footnote{\url{https://chat.openai.com/}}.
Curriculum learning
Bengio, Y., Louradour, J., Collobert, R., and Weston, J. (2009) · 2009
Earlier work this paper cites.
Social chemistry 101: Learning to reason about social and moral norms
Forbes, M., Hwang, J. D., Shwartz, V., Sap, M., and Choi, Y. (2020) · 2011
Earlier work this paper cites.
Cn-dbpedia: A never-ending chinese knowledge extraction system
Xu, B., Xu, Y., Liang, J., Xie, C., Liang, B., Cui, W., and Xiao, Y. (2017) · 2017
Earlier work this paper cites.
Learning to reweight examples for robust deep learning
Ren, M., Zeng, W., Yang, B., and Urtasun, R. (2018) · 2018
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J. (2020) · 2020
Earlier work this paper cites.
Gradient surgery for multi-task learning
Yu, T., Kumar, S., Gupta, A., Levine, S., Hausman, K., and Finn, C. (2020) · 2020
Earlier work this paper cites.
Ext5: Towards extreme multi-task scaling for transfer learning
Aribandi, V., Tay, Y., Schuster, T., Rao, J., Zheng, H., Mehta, S. V., Zhuang, H., Tran, V., Bahri, D., Ni, J., Gupta, J., Hui, K., Ruder, S., and Metzler, D. (2021) · 2021
Earlier work this paper cites.
Moral stories: Situated reasoning about norms, intents, actions, and their consequences
Emelin, D., Le Bras, R., Hwang, J. D., Forbes, M., and Choi, Y. (2021) · 2021
Earlier work this paper cites.
Active curriculum learning
Jafarpour, B., Sepehr, D., and Pogrebnyakov, N. (2021) · 2021
Earlier work this paper cites.
Metaicl: Learning to learn in context
Min, S., Lewis, M., Zettlemoyer, L., and Hajishirzi, H. (2021) · 2021
Earlier work this paper cites.
Multitask prompted training enables zero-shot task generalization
Sanh, V., Webson, A., Raffel, C., Bach, S. H., Sutawika, L., Alyafeai, Z., Chaffin, A., Stiegler, A., Scao, T. L., Raja, A., Dey, M., Bari, M. S., Xu, C., Thakker, U., Sharma, S. S., Szczechla, E., Kim, T., Chhablani, G., Nayak, N., Datta, D., Chang, J., Jiang, M. T.-J., Wang, H., Manica, M., Shen, S., Yong, Z. X., Pandey, H., Bawden, R., Wang, T., Neeraj, T., Rozen, J., Sharma, A., Santilli, A., Fevry, T., Fries, J. A., Teehan, R., Biderman, S., Gao, L., Bers, T., Wolf, T., and Rush, A. M. (2021) · 2021
Earlier work this paper cites.
A survey on curriculum learning
Wang, X., Chen, Y., and Zhu, W. (2021) · 2021
Earlier work this paper cites.
Finetuned language models are zero-shot learners
Wei, J., Bosma, M., Zhao, V., Guu, K., Yu, A. W., Lester, B., Du, N., Dai, A. M., and Le, Q. V. (2021) · 2021
Earlier work this paper cites.
Cuge: A chinese language understanding and generation evaluation benchmark
Yao, Y., Dong, Q., Guan, J., Cao, B., Zhang, Z., Xiao, C., Wang, X., Qi, F., Bao, J., Nie, J., Zeng, Z., Gu, Y., Zhou, K., Huang, X., Li, W., Ren, S., Lu, J., Xu, C., Wang, H., Zeng, G., Zhou, Z., Zhang, J., Li, J., Huang, M., Yan, R., He, X., Wan, X., Zhao, X., Sun, X., Liu, Y., Liu, Z., Han, X., Yang, E., Sui, Z., and Sun, M. (2021) · 2021
Earlier work this paper cites.
CrossFit: A few-shot learning challenge for cross-task generalization in NLP
Ye, Q., Lin, B. Y., and Ren, X. (2021) · 2021
Earlier work this paper cites.
Promptsource: An integrated development environment and repository for natural language prompts
Bach, S. H., Sanh, V., Yong, Z. X., Webson, A., Raffel, C., Nayak, N. V., Sharma, A., Kim, T., Bari, M. S., Févry, T., Alyafeai, Z., Dey, M., Santilli, A., Sun, Z., Ben-David, S., Xu, C., Chhablani, G., Wang, H., Fries, J. A., Al-shaibani, M. S., Sharma, S., Thakker, U., Almubarak, K., Tang, X., Jiang, M. T.-J., and Rush, A. M. (2022) · 2022
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Bai, Y., Jones, A., Ndousse, K., Askell, A., Chen, A., DasSarma, N., Drain, D., Fort, S., Ganguli, D., Henighan, T., Joseph, N., Kadavath, S., Kernion, J., Conerly, T., El-Showk, S., Elhage, N., Hatfield-Dodds, Z., Hernandez, D., Hume, T., Johnston, S., Kravec, S., Lovitt, L., Nanda, N., Olsson, C., Amodei, D., Brown, T. B., Clark, J., McCandlish, S., Olah, C., Mann, B., and Kaplan, J. (2022) · 2022
Earlier work this paper cites.
Scaling instruction-finetuned language models
Chung, H. W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., Li, E., Wang, X., Dehghani, M., Brahma, S., Webson, A., Gu, S., Dai, Z., Suzgun, M., Chen, X., Chowdhery, A., Valter, D., Narang, S., Mishra, G., Yu, A., Zhao, V., Huang, Y., Dai, A. M., Yu, H., Petrov, S., Chi, E., Dean, J., Devlin, J., Roberts, A., Zhou, D., Le, Q. V., and Wei, J. (2022) · 2022
Cited alongside, same era.
Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned
Ganguli, D., Lovitt, L., Kernion, J., Askell, A., Bai, Y., Kadavath, S., Mann, B., Perez, E., Schiefer, N., Ndousse, K., Jones, A., Bowman, S., Chen, A., Conerly, T., DasSarma, N., Drain, D., Elhage, N., El-Showk, S., Fort, S., Hatfield-Dodds, Z., Henighan, T., Hernandez, D., Hume, T., Jacobson, J., Johnston, S., Kravec, S., Olsson, C., Ringer, S., Tran-Johnson, E., Amodei, D., Brown, T., Joseph, N., McCandlish, S., Olah, C., Kaplan, J., and Clark, J. (2022) · 2022
Cited alongside, same era.
Unnatural instructions: Tuning language models with (almost) no human labor
Honovich, O., Scialom, T., Levy, O., and Schick, T. (2022) · 2022
Cited alongside, same era.
CSL: A large-scale Chinese scientific literature dataset
Li, Y., Zhang, Y., Zhao, Z., Shen, L., Liu, W., Mao, W., and Zhang, H. (2022) · 2022
Cited alongside, same era.
Koala: A dialogue model for academic research
Geng, X., Gudibande, A., Liu, H., Wallace, E., Abbeel, P., Levine, S., and Song, D. (2023) · 2023
Closest in time.
How close is chatgpt to human experts? comparison corpus, evaluation, and detection
Guo, B., Zhang, X., Wang, Z., Jiang, M., Nie, J., Ding, Y., Yue, J., and Wu, Y. (2023) · 2023
Closest in time.
Camel: Communicative agents for “mind” exploration of large scale language model society
Guohao, L., Hasan Abed Al, K. H., Hani, I., Dmitrii, K., and Ghanem, B. (2023) · 2023
Closest in time.
Guanacodataset
JosephusCheung (2021) · 2023
Closest in time.
Huggingface h4 stack exchange preference dataset
Lambert, N., Tunstall, L., Rajani, N., and Thrush, T. (2023) · 2023
Closest in time.
Chinese alpaca dataset
Liu, B., Huang, K., Jiao, L., He, Y., Zhang, R., Liang, Y., and Wang, Y. (2023) · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cross-task generalization via natural language crowdsourcing instructions
Mishra, S., Khashabi, D., Baral, C., and Hajishirzi, H. (2022) · 2022
Cited alongside, same era.
Crosslingual generalization through multitask finetuning
Muennighoff, N., Wang, T., Sutawika, L., Roberts, A., Biderman, S., Scao, T. L., Bari, M. S., Shen, S., Yong, Z.-X., Schoelkopf, H., Tang, X., Radev, D., Aji, A. F., Almubarak, K., Albanie, S., Alyafeai, Z., Webson, A., Raff, E., and Raffel, C. (2022) · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al. (2022) · 2022
Cited alongside, same era.
Potato: The portable text annotation tool
Pei, J., Ananthasubramaniam, A., Wang, X., Zhou, N., Dedeloudis, A., Sargent, J., and Jurgens, D. (2022) · 2022
Cited alongside, same era.
Unifiedskg: Unifying and multi-tasking structured knowledge grounding with text-to-text language models
Xie, T., Wu, C. H., Shi, P., Zhong, R., Scholak, T., Yasunaga, M., Wu, C., Zhong, M., Yin, P., Wang, S. I., Zhong, V., Wang, B., Li, C., Boyle, C., Ni, A., Yao, Z., Radev, D., Xiong, C., Kong, L., Zhang, R., Smith, N. A., Zettlemoyer, L., and Yu, T. (2022) · 2022
Cited alongside, same era.
ZeroPrompt: Scaling prompt-based pretraining to 1,000 tasks improves zero-shot generalization
Xu, H., Chen, Y., Du, Y., Shao, N., Yanggang, W., Li, H., and Yang, Z. (2022) · 2022
Cited alongside, same era.
Oig dataset
AI, L. (2021) · 2023
Cited alongside, same era.
Huggingface datasets: hh-rlhf
Anthropic (2022) · 2023
Cited alongside, same era.
Closest in time.
The flan collection: Designing data and methods for effective instruction tuning
Longpre, S., Hou, L., Vu, T., Webson, A., Chung, H. W., Tay, Y., Zhou, D., Le, Q. V., Zoph, B., Wei, J., and Roberts, A. (2023) · 2023
Closest in time.
Generative agents: Interactive simulacra of human behavior
Park, J. S., O’Brien, J. C., Cai, C. J., Morris, M. R., Liang, P., and Bernstein, M. S. (2023) · 2023
Closest in time.
Peng, B., Li, C., He, P., Galley, M., and Gao, J. (2023) · 2023
Closest in time.
Sharegpt
ShareGPT (2021) · 2023
Closest in time.
Stanford alpaca: An instruction-following llama model
Taori, R., Gulrajani, I., Zhang, T., Dubois, Y., Li, X., Guestrin, C., Liang, P., and Hashimoto, T. B. (2023) · 2023
Closest in time.
Baize: An open-source chat model with parameter-efficient tuning on self-chat data
Xu, C., Guo, D., Duan, N., and McAuley, J. (2023) · 2023
Closest in time.
Instruction in the wild: A user-based instruction dataset
Xue, F., Zheng, Z., and You, Y. (2023) · 2023
Closest in time.
Chinese-chatllama
YDli-ai (2021) · 2023
Closest in time.
Belle: Be everyone’s large language model engine
Yunjie, J., Yong, D., Yan, G., Yiping, P., Qiang, N., Baochang, M., and Xiangang, L. (2023) · 2023
Closest in time.
GLM-130b: An open bilingual pre-trained model
Zeng, A., Liu, X., Du, Z., Wang, Z., Lai, H., Ding, M., Yang, Z., Xu, Y., Zheng, W., Xia, X., Tam, W. L., Ma, Z., Xue, Y., Zhai, J., Chen, W., Liu, Z., Zhang, P., Dong, Y., and Tang, J. (2023) · 2023
Closest in time.
Luotuo: An instruction-following chinese language model, lora tuning on llama
Ziang Leng, Q. C. and Li, C. (2023) · 2023
Closest in time.