DQ-BART: Efficient Sequence-to-Sequence Model via Joint Distillation and Quantization
Original
Li, Z., Wang, Z., Tan, M., Nallapati, R., Bhatia, P., Arnold, A., Xiang, B., & Roth, D. (2022) · 2022
Later among the works it cites.
Teaching language models to support answers with verified quotes
Original
Menick, J., Trebacz, M., Mikulik, V., Aslanides, J., Song, F., Chadwick, M., Glaese, M., Young, S., Campbell-Gillingham, L., Irving, G., et al. (2022) · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Original
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al. (2022) · 2022
Later among the works it cites.
Hierarchical text-conditional image generation with clip latents
Original
Ramesh, A., Dhariwal, P., Nichol, A., Chu, C., & Chen, M. (2022) · 2022
Later among the works it cites.
A generalist agent
Original
Reed, S., Zolna, K., Parisotto, E., Colmenarejo, S. G., Novikov, A., Barth-Maron, G., Gimenez, M., Sulsky, Y., Kay, J., Springenberg, J. T., et al. (2022) · 2022
Later among the works it cites.
High-resolution image synthesis with latent diffusion models
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., & Ommer, B. (2022) · 2022
Later among the works it cites.
Photorealistic text-to-image diffusion models with deep language understanding
Original
Saharia, C., Chan, W., Saxena, S., Li, L., Whang, J., Denton, E., Ghasemipour, S. K. S., Ayan, B. K., Mahdavi, S. S., Lopes, R. G., et al. (2022) · 2022
Later among the works it cites.
Using deepspeed and megatron to train megatron-turing nlg 530b, a large-scale generative language model
Original
Smith, S., Patwary, M., Norick, B., LeGresley, P., Rajbhandari, S., Casper, J., Liu, Z., Prabhumoye, S., Zerveas, G., Korthikanti, V., et al. (2022) · 2022
Later among the works it cites.
AlexaTM 20B: Few-Shot Learning Using a Large-Scale Multilingual Seq2Seq Model
Original
Soltan, S., Ananthakrishnan, S., FitzGerald, J., Gupta, R., Hamza, W., Khan, H., Peris, C., Rawls, S., Rosenbaum, A., Rumshisky, A., Prakash, C. S., Sridhar, M., Triefenbach, F., Verma, A., Tur, G., & Natarajan, P. (2022) · 2022
Later among the works it cites.
One Embedder, Any Task: Instruction-Finetuned Text Embeddings
Original
Su, H. S., Shi, W. S., Kasai, J., Wang, Y., Hu, Y., Ostendorf, M., Yih, W.-t., Smith, N. A., Zettlemoyer, L., & Yu, T. (2022) · 2022
Later among the works it cites.
Unifying Language Learning Paradigms
Original
Tay, Y., Dehghani, M., Tran, V. Q., Garcia, X., Bahri, D., Schuster, T., Zheng, H. S., Houlsby, N., & Metzler, D. (2022) · 2022
Later among the works it cites.
GALACTICA: A Large Language Model for Science
Original
Taylor, R., Kardas, M., Cucurull, G., Scialom, T., Hartshorn, A., Saravia, E., Poulton, A., Kerkez, V., & Stojnic, R. (2022) · 2022
Later among the works it cites.
Lamda: Language models for dialog applications
Original
Thoppilan, R., De Freitas, D., Hall, J., Shazeer, N., Kulshreshtha, A., Cheng, H.-T., Jin, A., Bos, T., Baker, L., Du, Y., et al. (2022) · 2022
Later among the works it cites.
Text Embeddings by Weakly-Supervised Contrastive Pre-training
Original
Wang, L., Yang, N., Huang, X., Jiao, B., Yang, L., Jiang, D., Majumder, R., & Wei, F. (2022) · 2022
Later among the works it cites.
Diffusion models: A comprehensive survey of methods and applications
Original
Yang, L., Zhang, Z., Song, Y., Hong, S., Xu, R., Zhao, Y., Shao, Y., Zhang, W., Cui, B., & Yang, M.-H. (2022) · 2022
Later among the works it cites.
Opt: Open pre-trained transformer language models
Original
Zhang, S., Roller, S., Goyal, N., Artetxe, M., Chen, M., Chen, S., Dewan, C., Diab, M., Li, X., Lin, X. V., et al. (2022) · 2022
Later among the works it cites.
Education in the era of generative artificial intelligence (AI): Understanding the potential benefits of ChatGPT in promoting teaching and learning
Baidoo-Anu, D., & Owusu Ansah, L. (2023) · 2023
Closest in time.
Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling
Original
Biderman, S., Schoelkopf, H., & and, Q. A. (2023) · 2023
Closest in time.
Deep reinforcement learning from human preferences
Original
Christiano, P., Leike, J., Brown, T. B., Martic, M., Legg, S., & Amodei, D. (2023) · 2023
Closest in time.
Let’s chat about ChatGPT
Dennean, K., Gantori, S., Limas, D. K., & Allen Pu, a. R. G. (2023) · 2023
Closest in time.
OpenAssistant Conversations – Democratizing Large Language Model Alignment
Original
Köpf, A., Kilcher, Y., von Rütte, D., Anagnostidis, S., Tam, Z.-R., Stevens, K., Barhoum, A., Duc, N. M., Stanley, O., Nagyfi, R., ES, S., Suri, S., Glushkov, D., Dantuluri, A., Maguire, A., Schuhmann, C., Nguyen, H., & Mattick, A. (2023) · 2023
Closest in time.
GPT-4 Technical Report
Original
OpenAI (2023) · 2023
Closest in time.
What ChatGPT and generative AI mean for science
Stokel-Walker, C., & Noorden, R. V. (2023) · 2023
Closest in time.
Stanford Alpaca: An Instruction-following LLaMA model
Taori, R., Gulrajani, I., Zhang, T., Dubois, Y., Li, X., Guestrin, C., Liang, P., & Hashimoto, T. B. (2023) · 2023
Closest in time.
LLaMA: Open and Efficient Foundation Language Models
Original
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., Rodriguez, A., Joulin, A., Grave, E., & Lample, G. (2023) · 2023
Closest in time.
Large Language Models: A Survey.
Minaee, S., Mikolov, T., Nikzad, N., Chenaghlu, M., Socher, R., Amatriain, X., & Gao, J. (2024) · 2024
Closest in time.
ERNIE: Enhanced language representation with informative entities
Original
Zhang, Z., Han, X., Liu, Z., Jiang, X., Sun, M., & Liu, Q. (2019b) · 2024
Closest in time.