Fetching the paper…
Reading the bibliography…
This work introduces Weaver, our first family of large language models (LLMs) dedicated to content creation.
The CoNLL-2014 shared task on grammatical error correction
H. T. Ng, S. M. Wu, T. Briscoe, C. Hadiwinoto, R. H. Susanto, and C. Bryant · 2014
Earlier work this paper cites.
Toward controlled generation of text
Z. Hu, Z. Yang, X. Liang, R. Salakhutdinov, and E. P. Xing · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
I. Loshchilov and F. Hutter · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Generating sentences by editing prototypes
K. Guu, T. B. Hashimoto, Y. Oren, and P. Liang · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
A. Radford, K. Narasimhan, T. Salimans, I. Sutskever, et al · 2018
Earlier work this paper cites.
The BEA-2019 shared task on grammatical error correction
C. Bryant, M. Felice, Ø. E. Andersen, and T. Briscoe · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever, et al · 2019
Earlier work this paper cites.
Megatron-lm: Training multi-billion parameter language models using model parallelism
M. Shoeybi, M. Patwary, R. Puri, P. LeGresley, J. Casper, and B. Catanzaro · 2019
Earlier work this paper cites.
Root mean square layer normalization
B. Zhang and R. Sennrich · 2019
Earlier work this paper cites.
Hierarchical summary-to-article generation, 2019
W. Zhou, T. Ge, K. Xu, F. Wei, and M. Zhou · 2019
Earlier work this paper cites.
Language models are few-shot learners, 2020
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei · 2020
Earlier work this paper cites.
Measuring massive multitask language understanding
D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt · 2020
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
P. Lewis, E. Perez, A. Piktus, F. Petroni, V. Karpukhin, N. Goyal, H. Küttler, M. Lewis, W.-t. Yih, T. Rocktäschel, et al · 2020
Earlier work this paper cites.
CommonGen: A constrained text generation challenge for generative commonsense reasoning
B. Y. Lin, W. Zhou, M. Shen, P. Zhou, C. Bhagavatula, Y. Choi, and X. Ren · 2020
Earlier work this paper cites.
Glu variants improve transformer
N. Shazeer · 2020
Earlier work this paper cites.
Constitutional ai: Harmlessness from ai feedback, 2022
Y. Bai, S. Kadavath, S. Kundu, A. Askell, J. Kernion, A. Jones, A. Chen, A. Goldie, A. Mirhoseini, C. McKinnon, C. Chen, C. Olsson, C. Olah, D. Hernandez, D. Drain, D. Ganguli, D. Li, E. Tran-Johnson, E. Perez, J. Kerr, J. Mueller, J. Ladish, J. Landau, K. Ndousse, K. Lukosuite, L. Lovitt, M. Sellitto, N. Elhage, N. Schiefer, N. Mercado, N. DasSarma, R. Lasenby, R. Larson, S. Ringer, S. Johnston, S. Kravec, S. E. Showk, S. Fort, T. Lanham, T. Telleen-Lawton, T. Conerly, T. Henighan, T. Hume, S. R. Bowman, Z. Hatfield-Dodds, B. Mann, D. Amodei, N. Joseph, S. McCandlish, T. Brown, and J. Kaplan · 2022
Earlier work this paper cites.
Scaling instruction-finetuned language models, 2022
H. W. Chung, L. Hou, S. Longpre, B. Zoph, Y. Tay, W. Fedus, Y. Li, X. Wang, M. Dehghani, S. Brahma, A. Webson, S. S. Gu, Z. Dai, M. Suzgun, X. Chen, A. Chowdhery, A. Castro-Ros, M. Pellat, K. Robinson, D. Valter, S. Narang, G. Mishra, A. Yu, V. Zhao, Y. Huang, A. Dai, H. Yu, S. Petrov, E. H. Chi, J. Dean, J. Devlin, A. Roberts, D. Zhou, Q. V. Le, and J. Wei · 2022
Earlier work this paper cites.
FlashAttention: Fast and memory-efficient exact attention with IO-awareness
T. Dao, D. Y. Fu, S. Ermon, A. Rudra, and C. Ré · 2022
Cited alongside, same era.
Introducing ChatGPT, 2022
OpenAI · 2022
Cited alongside, same era.
Multitask prompted training enables zero-shot task generalization
V. Sanh, A. Webson, C. Raffel, S. Bach, L. Sutawika, Z. Alyafeai, A. Chaffin, A. Stiegler, A. Raja, M. Dey, M. S. Bari, C. Xu, U. Thakker, S. S. Sharma, E. Szczechla, T. Kim, G. Chhablani, N. Nayak, D. Datta, J. Chang, M. T.-J. Jiang, H. Wang, M. Manica, S. Shen, Z. X. Yong, H. Pandey, R. Bawden, T. Wang, T. Neeraj, J. Rozen, A. Sharma, A. Santilli, T. Fevry, J. A. Fries, R. Teehan, T. L. Scao, S. Biderman, L. Gao, T. Wolf, and A. M. Rush · 2022
Cited alongside, same era.
Summarize, outline, and elaborate: Long-text generation via hierarchical supervision from extractive summaries
X. Sun, Z. Sun, Y. Meng, J. Li, and C. Fan · 2022
Cited alongside, same era.
Chain of thought prompting elicits reasoning in large language models
J. Wei, X. Wang, D. Schuurmans, M. Bosma, brian ichter, F. Xia, E. H. Chi, Q. V. Le, and D. Zhou · 2022
Cited alongside, same era.
What makes good data for alignment? a comprehensive study of automatic data selection in instruction tuning, 2023
W. Liu, W. Zeng, K. He, Y. Jiang, and J. He · 2023
Later among the works it cites.
OpenAI · 2023
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model
R. Rafailov, A. Sharma, E. Mitchell, S. Ermon, C. D. Manning, and C. Finn · 2023
Later among the works it cites.
Code llama: Open foundation models for code, 2023
B. Rozière, J. Gehring, F. Gloeckle, S. Sootla, I. Gat, X. E. Tan, Y. Adi, J. Liu, T. Remez, J. Rapin, A. Kozhevnikov, I. Evtimov, J. Bitton, M. Bhatt, C. C. Ferrer, A. Grattafiori, W. Xiong, A. Défossez, J. Copet, F. Azhar, H. Touvron, L. Martin, N. Usunier, T. Scialom, and G. Synnaeve · 2023
Later among the works it cites.
Toolformer: Language models can teach themselves to use tools
T. Schick, J. Dwivedi-Yu, R. Dessi, R. Raileanu, M. Lomeli, E. Hambro, L. Zettlemoyer, N. Cancedda, and T. Scialom · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Re3: Generating longer stories with recursive reprompting and revision
K. Yang, Y. Tian, N. Peng, and D. Klein · 2022
Cited alongside, same era.
Gqa: Training generalized multi-query transformer models from multi-head checkpoints
J. Ainslie, J. Lee-Thorp, M. de Jong, Y. Zemlyanskiy, F. Lebrón, and S. Sanghai · 2023
Cited alongside, same era.
Introducing Claude, 2023
Anthropic · 2023
Cited alongside, same era.
Chateval: Towards better llm-based evaluators through multi-agent debate, 2023
C.-M. Chan, W. Chen, Y. Su, J. Yu, W. Xue, S. Zhang, J. Fu, and Z. Liu · 2023
Cited alongside, same era.
Chatlaw: Open-source legal large language model with integrated external knowledge bases, 2023
J. Cui, Z. Li, Y. Yan, B. Chen, and L. Yuan · 2023
Cited alongside, same era.
FlashAttention-2: Faster attention with better parallelism and work partitioning
T. Dao · 2023
Cited alongside, same era.
Doolittle: Benchmarks and corpora for academic writing formalization
S. Diao, Y. Lei, L. Pan, T. Fang, W. Zhou, S. S. Keh, M.-Y. Kan, and T. Zhang · 2023
Cited alongside, same era.
Later among the works it cites.
Large language model alignment: A survey
T. Shen, R. Jin, Y. Huang, C. Liu, W. Dong, Z. Guo, X. Wu, Y. Liu, and D. Xiong · 2023
Later among the works it cites.
Salmon: Self-alignment with principle-following reward models, 2023
Z. Sun, Y. Shen, H. Zhang, Q. Zhou, Z. Chen, D. Cox, Y. Yang, and C. Gan · 2023
Later among the works it cites.
Gemini: A family of highly capable multimodal models, 2023
Gemini Team · 2023
Later among the works it cites.
Zephyr: Direct distillation of lm alignment, 2023
L. Tunstall, E. Beeching, N. Lambert, N. Rajani, K. Rasul, Y. Belkada, S. Huang, L. von Werra, C. Fourrier, N. Habib, N. Sarrazin, O. Sanseviero, A. M. Rush, and T. Wolf · 2023
Later among the works it cites.
Self-instruct: Aligning language models with self-generated instructions
Y. Wang, Y. Kordi, S. Mishra, A. Liu, N. A. Smith, D. Khashabi, and H. Hajishirzi · 2023
Later among the works it cites.
Bloomberggpt: A large language model for finance, 2023
S. Wu, O. Irsoy, S. Lu, V. Dabravolski, M. Dredze, S. Gehrmann, P. Kambadur, D. Rosenberg, and G. Mann · 2023
Later among the works it cites.
DOC: Improving long story coherence with detailed outline control
K. Yang, D. Klein, N. Peng, and Y. Tian · 2023
Later among the works it cites.
A survey on multimodal large language models
S. Yin, C. Fu, S. Zhao, K. Li, X. Sun, T. Xu, and E. Chen · 2023
Later among the works it cites.
Socratic models: Composing zero-shot multimodal reasoning with language
A. Zeng, M. Attarian, brian ichter, K. M. Choromanski, A. Wong, S. Welker, F. Tombari, A. Purohit, M. S. Ryoo, V. Sindhwani, J. Lee, V. Vanhoucke, and P. Florence · 2023
Later among the works it cites.
A survey of large language models
W. X. Zhao, K. Zhou, J. Li, T. Tang, X. Wang, Y. Hou, Y. Min, B. Zhang, J. Zhang, Z. Dong, et al · 2023
Later among the works it cites.
Roformer: Enhanced transformer with rotary position embedding
J. Su, M. Ahmed, Y. Lu, S. Pan, W. Bo, and Y. Liu · 2024
Closest in time.
Secrets of rlhf in large language models part ii: Reward modeling
B. Wang, R. Zheng, L. Chen, Y. Liu, S. Dou, C. Huang, W. Shen, S. Jin, E. Zhou, C. Shi, et al · 2024
Closest in time.