Fetching the paper…
Reading the bibliography…
Large Language Models(LLMs) have dramatically revolutionized the field of Natural Language Processing(NLP), offering remarkable capabilities that have garnered widespread usage.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 1901
Earlier work this paper cites.
A tree-based statistical language model for natural language speech recognition
Bahl, L. R., Brown, P. F., de Souza, P. V., and Mercer, R. L · 1989
Earlier work this paper cites.
Online algorithms: a survey
Albers, S · 2003
Earlier work this paper cites.
A fast learning algorithm for deep belief nets
Hinton, G. E., Osindero, S., and Teh, Y. W · 2006
Earlier work this paper cites.
Scaling learning algorithms towards AI
Bengio, Y., and LeCun, Y · 2007
Earlier work this paper cites.
Natural language processing (almost) from scratch
Collobert, R., Weston, J., Bottou, L., Karlen, M., Kavukcuoglu, K., and Kuksa, P · 2011
Earlier work this paper cites.
Deep learning
Goodfellow, I., Bengio, Y., Courville, A., and Bengio, Y · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding, 2019
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Earlier work this paper cites.
Knowledge enhanced contextual word representations
Peters, M. E., Neumann, M., Logan IV, R. L., Schwartz, R., Joshi, V., Singh, S., and Smith, N. A · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Earlier work this paper cites.
Megatron-lm: Training multi-billion parameter language models using model parallelism
Shoeybi, M., Patwary, M., Puri, R., LeGresley, P., Casper, J., and Catanzaro, B · 2019
Earlier work this paper cites.
Mass: Masked sequence to sequence pre-training for language generation, 2019
Song, K., Tan, X., Qin, T., Lu, J., and Liu, T.-Y · 2019
Earlier work this paper cites.
Continual lifelong learning in natural language processing: A survey
Biesialska, M., Biesialska, K., and Costa-jussà , M. R · 2020
Earlier work this paper cites.
Retrieval augmented language model pre-training
Guu, K., Lee, K., Tung, Z., Pasupat, P., and Chang, M · 2020
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2020
Earlier work this paper cites.
A survey on contrastive self-supervised learning
Jaiswal, A., Babu, A. R., Zadeh, M. Z., Banerjee, D., and Makedon, F · 2020
Earlier work this paper cites.
What disease does this patient have? a large-scale open domain question answering dataset from medical exams, 2020
Jin, D., Pan, E., Oufattole, N., Weng, W.-H., Fang, H., and Szolovits, P · 2020
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W.-t., Rocktäschel, T., et al · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Earlier work this paper cites.
Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters
Rasley, J., Rajbhandari, S., Ruwase, O., and He, Y · 2020
Earlier work this paper cites.
Generalizing from a few examples: A survey on few-shot learning
Wang, Y., Yao, Q., Kwok, J. T., and Ni, L. M · 2020
Earlier work this paper cites.
Colossal-ai: A unified deep learning system for large-scale parallel training
Bian, Z., Liu, H., Wang, B., Huang, H., Li, Y., Wang, C., Cui, F., and You, Y · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., et al · 2021
Earlier work this paper cites.
Fairscale: A general purpose modular pytorch library for high performance and large scale training
FairScale authors · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2021
Earlier work this paper cites.
Colossal-ai: A unified deep learning system for large-scale parallel training
Li, S., Fang, J., Bian, Z., Liu, H., Liu, Y., Huang, H., Wang, B., and You, Y · 2021
Earlier work this paper cites.
What makes good in-context examples for gpt- 3 3 ?
Liu, J., Shen, D., Zhang, Y., Dolan, B., Carin, L., and Chen, W · 2021
Earlier work this paper cites.
Lu, Y., Lu, H., Fu, G., and Liu, Q · 2021
Cited alongside, same era.
Multitask prompted training enables zero-shot task generalization
Sanh, V., Webson, A., Raffel, C., Bach, S. H., Sutawika, L., Alyafeai, Z., Chaffin, A., Stiegler, A., Scao, T. L., Raja, A., et al · 2021
Cited alongside, same era.
Ernie 3.0: Large-scale knowledge enhanced pre-training for language understanding and generation
Sun, Y., Wang, S., Feng, S., Ding, S., Pang, C., Shang, J., Liu, J., Chen, X., Zhao, Y., Lu, Y., et al · 2021
Cited alongside, same era.
Finetuned language models are zero-shot learners
Wei, J., Bosma, M., Zhao, V. Y., Guu, K., Yu, A. W., Lester, B., Du, N., Dai, A. M., and Le, Q. V · 2021
Cited alongside, same era.
A survey of knowledge enhanced pre-trained language models
Hu, L., Liu, Z., Zhao, Z., Hou, L., Nie, L., and Li, J · 2023
Later among the works it cites.
Active retrieval augmented generation
Jiang, Z., Xu, F. F., Gao, L., Sun, Z., Liu, Q., Dwivedi-Yu, J., Yang, Y., Callan, J., and Neubig, G · 2023
Later among the works it cites.
Openassistant conversations–democratizing large language model alignment
Köpf, A., Kilcher, Y., von Rütte, D., Anagnostidis, S., Tam, Z.-R., Stevens, K., Barhoum, A., Duc, N. M., Stanley, O., Nagyfi, R., et al · 2023
Later among the works it cites.
LangChain
Langchain authors · 2023
Later among the works it cites.
Self-alignment with instruction backtranslation
Li, X., Yu, P., Zhou, C., Schick, T., Zettlemoyer, L., Levy, O., Weston, J., and Lewis, M · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Black, S., Biderman, S., Hallahan, E., Anthony, Q., Gao, L., Golding, L., He, H., Leahy, C., McDonell, K., Phang, J., et al · 2022
Cited alongside, same era.
LangChain, Oct. 2022
Chase, H · 2022
Cited alongside, same era.
Scaling instruction-finetuned language models
Chung, H. W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., Li, E., Wang, X., Dehghani, M., Brahma, S., et al · 2022
Cited alongside, same era.
A survey for in-context learning
Dong, Q., Li, L., Dai, D., Zheng, C., Wu, Z., Chang, B., Sun, X., Xu, J., and Sui, Z · 2022
Cited alongside, same era.
Training compute-optimal large language models
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., Casas, D. d. L., Hendricks, L. A., Welbl, J., Clark, A., et al · 2022
Cited alongside, same era.
Large language models can self-improve
Huang, J., Gu, S. S., Hou, L., Wu, Y., Wang, X., Yu, H., and Han, J · 2022
Cited alongside, same era.
Few-shot learning with retrieval augmented language models
Izacard, G., Lewis, P., Lomeli, M., Hosseini, L., Petroni, F., Schick, T., Dwivedi-Yu, J., Joulin, A., Riedel, S., and Grave, E · 2022
Cited alongside, same era.
Relational memory-augmented language models
Liu, Q., Yogatama, D., and Blunsom, P · 2022
Cited alongside, same era.
Self-alignment with instruction backtranslation, 2023
Li, X., Yu, P., Zhou, C., Schick, T., Zettlemoyer, L., Levy, O., Weston, J., and Lewis, M · 2023
Later among the works it cites.
Lost in the middle: How language models use long contexts, 2023
Liu, N. F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., and Liang, P · 2023
Later among the works it cites.
Megatron-DeepSpeed
Megatron-DeepSpeed authors · 2023
Later among the works it cites.
Fine-tuning or retrieval? comparing knowledge injection in llms, 2023
Ovadia, O., Brief, M., Mishaeli, M., and Elisha, O · 2023
Later among the works it cites.
Generative agents: Interactive simulacra of human behavior
Park, J. S., O’Brien, J., Cai, C. J., Morris, M. R., Liang, P., and Bernstein, M. S · 2023
Later among the works it cites.
Gorilla: Large language model connected with massive apis
Patil, S. G., Zhang, T., Wang, X., and Gonzalez, J. E · 2023
Later among the works it cites.
Peng, B., Galley, M., He, P., Cheng, H., Xie, Y., Hu, Y., Huang, Q., Liden, L., Yu, Z., Chen, W., et al · 2023
Later among the works it cites.
Tool learning with foundation models, 2023
Qin, Y., Hu, S., Lin, Y., Chen, W., Ding, N., Cui, G., Zeng, Z., Huang, Y., Xiao, C., Han, C., Fung, Y. R., Su, Y., Wang, H., Qian, C., Tian, R., Zhu, K., Liang, S., Shen, X., Xu, B., Zhang, Z., Ye, Y., Li, B., Tang, Z., Yi, J., Zhu, Y., Dai, Z., Yan, L., Cong, X., Lu, Y., Zhao, W., Huang, Y., Yan, J., Han, X., Sun, X., Li, D., Phang, J., Yang, C., Wu, T., Ji, H., Liu, Z., and Sun, M · 2023
Later among the works it cites.
Toolllm: Facilitating large language models to master 16000+ real-world apis
Qin, Y., Liang, S., Ye, Y., Zhu, K., Yan, L., Lu, Y., Lin, Y., Cong, X., Tang, X., Qian, B., et al · 2023
Later among the works it cites.
Toolformer: Language models can teach themselves to use tools
Schick, T., Dwivedi-Yu, J., Dessì, R., Raileanu, R., Lomeli, M., Zettlemoyer, L., Cancedda, N., and Scialom, T · 2023
Later among the works it cites.
Reflexion: an autonomous agent with dynamic memory and self-reflection
Shinn, N., Labash, B., and Gopinath, A · 2023
Later among the works it cites.
Text Generation Inference
Text Generation Inference authors · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models, 2023
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., Rodriguez, A., Joulin, A., Grave, E., and Lample, G · 2023
Later among the works it cites.
Shall we pretrain autoregressive language models with retrieval? a comprehensive study
Wang, B., Ping, W., Xu, P., McAfee, L., Liu, Z., Shoeybi, M., Dong, Y., Kuchaiev, O., Li, B., Xiao, C., et al · 2023
Later among the works it cites.
The rise and potential of large language model based agents: A survey
Xi, Z., Chen, W., Guo, X., He, W., Ding, Y., Hong, B., Zhang, M., Wang, J., Jin, S., Zhou, E., et al · 2023
Later among the works it cites.
{ \{ SkyPilot } \} : An intercloud broker for sky computing
Yang, Z., Wu, Z., Luo, M., Chiang, W.-L., Bhardwaj, R., Kwon, W., Zhuang, S., Luan, F. S., Mittal, G., Shenker, S., et al · 2023
Later among the works it cites.
Deepspeed-chat: Easy, fast and affordable rlhf training of chatgpt-like models at all scales
Yao, Z., Aminabadi, R. Y., Ruwase, O., Rajbhandari, S., Wu, X., Awan, A. A., Rasley, J., Zhang, M., Li, C., Holmes, C., et al · 2023
Later among the works it cites.
DeepSpeed-Chat: Easy, Fast and Affordable RLHF Training of ChatGPT-like Models at All Scales
Yao, Z., Aminabadi, R. Y., Ruwase, O., Rajbhandari, S., Wu, X., Awan, A. A., Rasley, J., Zhang, M., Li, C., Holmes, C., Zhou, Z., Wyatt, M., Smith, M., Kurilenko, L., Qin, H., Tanaka, M., Che, S., Song, S. L., and He, Y · 2023
Later among the works it cites.
Openbmb: Big model systems for large-scale representation learning
Zeng, G., Han, X., Zhang, Z., Liu, Z., Lin, Y., and Sun, M · 2023
Later among the works it cites.
Measuring massive multitask chinese understanding
Zeng, H · 2023
Later among the works it cites.
A survey of large language models
Zhao, W. X., Zhou, K., Li, J., Tang, T., Wang, X., Hou, Y., Min, Y., Zhang, B., Zhang, J., Dong, Z., et al · 2023
Later among the works it cites.
Judging llm-as-a-judge with mt-bench and chatbot arena, 2023
Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E. P., Zhang, H., Gonzalez, J. E., and Stoica, I · 2023
Later among the works it cites.
Lima: Less is more for alignment, 2023
Zhou, C., Liu, P., Xu, P., Iyer, S., Sun, J., Mao, Y., Ma, X., Efrat, A., Yu, P., Yu, L., Zhang, S., Ghosh, G., Lewis, M., Zettlemoyer, L., and Levy, O · 2023
Later among the works it cites.