Fetching the paper…
Reading the bibliography…
Instruction tuning is essential for large language models (LLMs) to become interactive.
1902
Earlier work this paper cites.
1907
Earlier work this paper cites.
1909
Earlier work this paper cites.
N. Shimizu, N. Rong, and T. Miyazaki, “Visual Question Answering Dataset for Bilingual Image Understanding: A Study of Cross-Lingual Transfer Using Attention Maps,” in Proceedings of the 27th International Conference on Computational Linguistics . Association for Computational Linguistics, 2018, pp. 1918–1928
1928
Earlier work this paper cites.
F. Jelinek, R. L. Mercer, L. R. Bahl, and J. K. Baker, “Perplexity—a measure of the difficulty of speech recognition tasks,” The Journal of the Acoustical Society of America , vol. 62, no. S1, pp. S63–S63, 1977
1977
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention Is All You Need,” in Advances in Neural Information Processing Systems , vol. 30, 2017, pp. 5999–6009
2017
Earlier work this paper cites.
P. F. Christiano, J. Leike, T. Brown, M. Martic, S. Legg, and D. Amodei, “Deep reinforcement learning from human preferences,” in Advances in neural information processing systems , vol. 30, 2017, pp. 4299–4307
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
A. Radford, K. Narasimhan, T. Salimans, and I. Sutskever, “Improving Language Understanding by Generative Pre-Training,” 2018. [Online]. Available: https://cdn.openai.com/research-covers/language-unsupervised/language_understanding_paper.pdf
2018
Earlier work this paper cites.
H. Tanioka, K. Kimura, K. Takaoka, R. Nakatani, and Y. Uchida, “Automatic Generation of Japanese Question-Answering Pairs,” in Fourth Asia Pacific Corpus Linguistics Conference (APCLC 2018) , 2018
2018
Earlier work this paper cites.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics . Association for Computational Linguistics, 2019, pp. 4171–4186
2019
Earlier work this paper cites.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever, “Language Models are Unsupervised Multitask Learners,” 2019. [Online]. Available: https://cdn.openai.com/better-language-models/language_models_are_unsupervised_multitask_learners.pdf
2019
Earlier work this paper cites.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al. , “Language Models are Few-Shot Learners,” Advances in Neural Information Processing Systems , vol. 33, pp. 1877–1901, 2020
2020
Earlier work this paper cites.
S. Rajbhandari, J. Rasley, O. Ruwase, and Y. He, “ZeRO: Memory Optimizations toward Training Trillion Parameter Models,” in SC20: International Conference for High Performance Computing, Networking, Storage and Analysis , 2020, pp. 1–16
2020
Earlier work this paper cites.
P. Keung, Y. Lu, G. Szarvas, and N. A. Smith, “The Multilingual Amazon Reviews Corpus,” in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing , 2020
2020
Earlier work this paper cites.
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu, “Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer,” Journal of Machine Learning Research , vol. 21, no. 140, pp. 1–67, 2020
2020
Earlier work this paper cites.
Z. Lin, A. Madotto, and P. Fung, “Exploring Versatile Generative Language Model Via Parameter-Efficient Transfer Learning,” in Findings of the Association for Computational Linguistics: EMNLP 2020 . Online: Association for Computational Linguistics, Nov. 2020, pp. 441–459
2020
Earlier work this paper cites.
Y. Tanaka, Y. Murawaki, D. Kawahara, and S. Kurohashi, “Improving a Japanese Typo Dataset and Typo Correction System Based on Wikipedia’s Revision History,” in Proceedings of the Twenty-seventh Annual Meeting of the Association for Natural Language Processing , 2021, pp. 1540–1545, (in Japanese). [Online]. Available: https://www.anlp.jp/proceedings/annual_meeting/2021/pdf_dir/E8-3.pdf
2021
Earlier work this paper cites.
L. Gao, J. Tow, S. Biderman, S. Black, A. DiPofi, C. Foster, L. Golding, J. Hsu, K. McDonell, N. Muennighoff et al. , “A framework for few-shot language model evaluation,” 2021. [Online]. Available: https://doi.org/10.5281/zenodo.5371628
2021
Earlier work this paper cites.
P. He, X. Liu, J. Gao, and W. Chen, “DeBERTa: Decoding-enhanced BERT with Disentangled Attention,” in International Conference on Learning Representations , 2021
2021
Cited alongside, same era.
X. L. Li and P. Liang, “Prefix-tuning: Optimizing continuous prompts for generation,” in Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers) . Association for Computational Linguistics, 2021, pp. 4582–4597
2021
Cited alongside, same era.
B. Lester, R. Al-Rfou, and N. Constant, “The power of scale for parameter-efficient prompt tuning,” in Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing . Association for Computational Linguistics, 2021, pp. 3045–3059
2021
Cited alongside, same era.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2022
Cited alongside, same era.
S. Black, S. Biderman, E. Hallahan, Q. Anthony, L. Gao, L. Golding, H. He, C. Leahy, K. McDonell, J. Phang et al. , “GPT-NeoX-20B: An Open-Source Autoregressive Language Model,” in Proceedings of BigScience Episode #5 – Workshop on Challenges & Perspectives in Creating Large Language Models . Association for Computational Linguistics, 2022, pp. 95–136
2022
Cited alongside, same era.
2022
Cited alongside, same era.
2022
Cited alongside, same era.
J. Wei, Y. Tay, R. Bommasani, C. Raffel, B. Zoph, S. Borgeaud, D. Yogatama, M. Bosma, D. Zhou, D. Metzler et al. , “Emergent abilities of large language models,” Transactions on Machine Learning Research , 2022
2022
Cited alongside, same era.
J. Wei, M. Bosma, V. Zhao, K. Guu, A. W. Yu, B. Lester, N. Du, A. M. Dai, and Q. V. Le, “Finetuned language models are zero-shot learners,” in International Conference on Learning Representations , 2022
2022
Cited alongside, same era.
V. Sanh, A. Webson, C. Raffel, S. Bach, L. Sutawika, Z. Alyafeai, A. Chaffin, A. Stiegler, A. Raja, M. Dey et al. , “Multitask Prompted Training Enables Zero-Shot Task Generalization,” in International Conference on Learning Representations , 2022
2022
Cited alongside, same era.
2022
Cited alongside, same era.
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen, “LoRA: Low-Rank Adaptation of Large Language Models,” in International Conference on Learning Representations , 2022, pp. 1–13
2022
Cited alongside, same era.
2023
Closest in time.
Databricks, “Dolly,” https://github.com/databrickslabs/dolly , 2023
2023
Closest in time.
Vicuna, “Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90% ChatGPT Quality,” https://vicuna.lmsys.org/ , 2023
2023
Closest in time.
R. Taori, I. Gulrajani, T. Zhang, Y. Dubois, X. Li, C. Guestrin, P. Liang, and T. B. Hashimoto, “Stanford Alpaca: An Instruction-following LLaMA model,” https://github.com/tatsu-lab/stanford_alpaca , 2023
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
P. He, J. Gao, and W. Chen, “DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding Sharing,” in International Conference on Learning Representations , 2023
2023
Closest in time.
OpenAI, “GPT-4 Technical Report,” 2023. [Online]. Available: https://arxiv.org/abs/2303.08774
2023
Closest in time.
S. Chaudhary, “Code Alpaca: An Instruction-following LLaMA model for code generation,” https://github.com/sahil280114/codealpaca , 2023
2023
Closest in time.
E. Wang, “Alpaca-LoRA,” https://github.com/tloen/alpaca-lora , 2023
2023
Closest in time.
Q. Zhang, M. Chen, A. Bukharin, P. He, Y. Cheng, W. Chen, and T. Zhao, “Adaptive budget allocation for parameter-efficient fine-tuning,” in International Conference on Learning Representations , 2023
2023
Closest in time.