Fetching the paper…
Reading the bibliography…
This is a book about large language models.
Prediction and entropy of printed english
[Shannon, 1951] Claude E Shannon · 1951
Earlier work this paper cites.
Rank analysis of incomplete block designs: I. the method of paired comparisons
[Bradley and Terry, 1952] Ralph Allan Bradley and Milton E. Terry · 1952
Earlier work this paper cites.
Some moral and technical consequences of automation: As machines learn they may develop unforeseen strategies at rates that baffle their programmers
[Wiener, 1960] Norbert Wiener · 1960
Earlier work this paper cites.
Sequential Decoding
[Wozengraft and Reiffen, 1961] John M. Wozengraft and Barney Reiffen · 1961
Earlier work this paper cites.
Error bounds for convolutional codes and an asymptotically optimum decoding algorithm
[Viterbi, 1967] Andrew J Viterbi · 1967
Earlier work this paper cites.
Maximum-likelihood sequence estimation of digital sequences in the presence of intersymbol interference
[Forney, 1972] GDJR Forney · 1972
Earlier work this paper cites.
The analysis of permutations
[Plackett, 1975] Robin L Plackett · 1975
Earlier work this paper cites.
Problems of monetary management: the UK experience
[Goodhart, 1984] Charles AE Goodhart · 1984
Earlier work this paper cites.
A simple rule-based part of speech tagger
[Brill, 1992] Eric Brill · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
[Williams, 1992] Ronald J Williams · 1992
Earlier work this paper cites.
The mathematics of statistical machine translation: Parameter estimation
[Brown et al., 1993] Peter F. Brown, Stephen A. Della Pietra, Vincent J. Della Pietra, and Robert L. Mercer · 1993
Earlier work this paper cites.
Negative evidence in language acquisition
[Marcus, 1993] Gary F Marcus · 1993
Earlier work this paper cites.
Unsupervised word sense disambiguation rivaling supervised methods
[Yarowsky, 1995] David Yarowsky · 1995
Earlier work this paper cites.
Is it an agent, or just a program?: A taxonomy for autonomous agents
[Franklin and Graesser, 1996] Stan Franklin and Art Graesser · 1996
Earlier work this paper cites.
Statistical parsing with a context-free grammar and word statistics
[Charniak, 1997] Eugene Charniak · 1997
Earlier work this paper cites.
Long short-term memory
[Hochreiter and Schmidhuber, 1997] Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Combining labeled and unlabeled data with co-training
[Blum and Mitchell, 1998] Avrim Blum and Tom Mitchell · 1998
Earlier work this paper cites.
Statistical methods for speech recognition
[Jelinek, 1998] Frederick Jelinek · 1998
Earlier work this paper cites.
Nonlinear multiobjective optimization , volume 12
[Miettinen, 1999] Kaisa Miettinen · 1999
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
[Ng et al., 1999] Andrew Y Ng, Daishi Harada, and Stuart J Russell · 1999
Earlier work this paper cites.
A neural probabilistic language model
[Bengio et al., 2003] Yoshua Bengio, Réjean Ducharme, Pascal Vincent, and Christian Jauvin · 2003
Earlier work this paper cites.
Minimum bayes-risk decoding for statistical machine translation
[Kumar and Byrne, 2004] Shankar Kumar and William Byrne · 2004
Earlier work this paper cites.
Learning to rank using gradient descent
[Burges et al., 2005] Chris Burges, Tal Shaked, Erin Renshaw, Ari Lazier, Matt Deeds, Nicole Hamilton, and Greg Hullender · 2005
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases
[Dolan and Brockett, 2005] Bill Dolan and Chris Brockett · 2005
Earlier work this paper cites.
Greedy layer-wise training of deep networks
[Bengio et al., 2006] Yoshua Bengio, Pascal Lamblin, Dan Popovici, and Hugo Larochelle · 2006
Earlier work this paper cites.
Pattern Recognition and Machine Learning
[Bishop, 2006] Christopher M. Bishop · 2006
Earlier work this paper cites.
Learning to rank: from pairwise approach to listwise approach
[Cao et al., 2007] Zhe Cao, Tao Qin, Tie-Yan Liu, Ming-Feng Tsai, and Hang Li · 2007
Earlier work this paper cites.
Speech and Language Processing (2nd ed.)
[Jurafsky and Martin, 2008] Dan Jurafsky and James H. Martin · 2008
Earlier work this paper cites.
Dynamic programming-based search algorithms in NLP
[Huang, 2009] Liang Huang · 2009
Earlier work this paper cites.
Learning to rank for information retrieval
[Liu, 2009] Tie-Yan Liu · 2009
Earlier work this paper cites.
Why does unsupervised pre-training help deep learning?
[Erhan et al., 2010] Dumitru Erhan, Aaron Courville, Yoshua Bengio, and Pascal Vincent · 2010
Earlier work this paper cites.
Statistical Machine Translation
[Koehn, 2010] Philipp Koehn · 2010
Earlier work this paper cites.
Algorithms for reinforcement learning
[Szepesvári, 2010] Csaba Szepesvári · 2010
Earlier work this paper cites.
Pascal recognizing textual entailment challenge (rte-7) at tac 2011
[Bentivogli and Giampiccolo, 2011] Luisa Bentivogli and Danilo Giampiccolo · 2011
Earlier work this paper cites.
Thinking, fast and slow
[Kahneman, 2011] Daniel Kahneman · 2011
Earlier work this paper cites.
Learning to Rank for Information Retrieval and Natural Language Processing
[Li, 2011] Hang Li · 2011
Earlier work this paper cites.
Efficient estimation of word representations in vector space
[Mikolov et al., 2013] Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean · 2013
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
[Mikolov et al., 2013] Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg Corrado, and Jeffrey Dean · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
[Socher et al., 2013] Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Y Ng, and Christopher Potts · 2013
Earlier work this paper cites.
Bagging and boosting statistical machine translation systems
[Xiao et al., 2013] Tong Xiao, Jingbo Zhu, and Tongran Liu · 2013
Earlier work this paper cites.
Complex problem solving: The European perspective
[Frensch and Funke, 2014] Peter A Frensch and Joachim Funke · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
[Pennington et al., 2014] Jeffrey Pennington, Richard Socher, and Christopher D. Manning · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
[Sutskever et al., 2014] Ilya Sutskever, Oriol Vinyals, and Quoc V Le · 2014
Earlier work this paper cites.
Deep learning in neural networks: An overview
[Schmidhuber, 2015] Jürgen Schmidhuber · 2015
Earlier work this paper cites.
Trust region policy optimization
[Schulman et al., 2015] John Schulman, Sergey Levine, Philipp Moritz, Michael Jordan, and Pieter Abbeel · 2015
Earlier work this paper cites.
Neural module networks
[Andreas et al., 2016] Jacob Andreas, Marcus Rohrbach, Trevor Darrell, and Dan Klein · 2016
Earlier work this paper cites.
Unitary evolution recurrent neural networks
[Arjovsky et al., 2016] Martin Arjovsky, Amar Shah, and Yoshua Bengio · 2016
Earlier work this paper cites.
Gaussian error linear units (gelus)
[Hendrycks and Gimpel, 2016] Dan Hendrycks and Kevin Gimpel · 2016
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
[Mnih et al., 2016] Volodymyr Mnih, Adrià Puigdomènech Badia, Mehdi Mirza, Alex Graves, Tim Harley, Timothy P Lillicrap, David Silver, and Koray Kavukcuoglu · 2016
Earlier work this paper cites.
Improving neural machine translation models with monolingual data
[Sennrich et al., 2016] Rico Sennrich, Barry Haddow, and Alexandra Birch · 2016
Earlier work this paper cites.
Neural architecture search with reinforcement learning
[Zoph and Le, 2016] Barret Zoph and Quoc Le · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
[Christiano et al., 2017] Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Earlier work this paper cites.
Deep learning scaling is predictable, empirically
[Hestness et al., 2017] Joel Hestness, Sharan Narang, Newsha Ardalani, Gregory Diamos, Heewoo Jun, Hassan Kianinejad, Md Mostofa Ali Patwary, Yang Yang, and Yanqi Zhou · 2017
Earlier work this paper cites.
Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension
[Joshi et al., 2017] Mandar Joshi, Eunsol Choi, Daniel S Weld, and Luke Zettlemoyer · 2017
Earlier work this paper cites.
Overcoming catastrophic forgetting in neural networks
[Kirkpatrick et al., 2017] James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A. Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, Demis Hassabis, Claudia Clopath, Dharshan Kumaran, and Raia Hadsell · 2017
Earlier work this paper cites.
Searching for activation functions
[Ramachandran et al., 2017] Prajit Ramachandran, Barret Zoph, and Quoc V Le · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
[Schulman et al., 2017] John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Bidirectional attention flow for machine comprehension
[Seo et al., 2017] Minjoon Seo, Aniruddha Kembhavi, Ali Farhadi, and Hannaneh Hajishirzi · 2017
Earlier work this paper cites.
Attention is all you need
[Vaswani et al., 2017] Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
[Dehghani et al., 2018] Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, and Łukasz Kaiser · 2018
Earlier work this paper cites.
Hierarchical neural story generation
[Fan et al., 2018] Angela Fan, Mike Lewis, and Yann Dauphin · 2018
Earlier work this paper cites.
Pipedream: Fast and efficient pipeline parallel dnn training
[Harlap et al., 2018] Aaron Harlap, Deepak Narayanan, Amar Phanishayee, Vivek Seshadri, Nikhil Devanur, Greg Ganger, and Phil Gibbons · 2018
Earlier work this paper cites.
Generalization without systematicity: On the compositional skills of sequence-to-sequence recurrent networks
[Lake and Baroni, 2018] Brenden Lake and Marco Baroni · 2018
Earlier work this paper cites.
Mixed precision training
[Micikevicius et al., 2018] Paulius Micikevicius, Sharan Narang, Jonah Alben, Gregory Diamos, Erich Elsen, David Garcia, Boris Ginsburg, Michael Houston, Oleksii Kuchaiev, Ganesh Venkatesh, and Hao Wu · 2018
Earlier work this paper cites.
Image transformer
[Parmar et al., 2018] Niki Parmar, Ashish Vaswani, Jakob Uszkoreit, Lukasz Kaiser, Noam Shazeer, Alexander Ku, and Dustin Tran · 2018
Earlier work this paper cites.
Deep contextualized word representations
[Peters et al., 2018] Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
[Radford et al., 2018] Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever · 2018
Earlier work this paper cites.
Self-attention with relative position representations
[Shaw et al., 2018] Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani · 2018
Earlier work this paper cites.
Reinforcement Learning: An Introduction (2nd ed.)
[Sutton and Barto, 2018] Richard S. Sutton and Andrew G. Barto · 2018
Earlier work this paper cites.
The web as a knowledge-base for answering complex questions
[Talmor and Berant, 2018] Alon Talmor and Jonathan Berant · 2018
Earlier work this paper cites.
A broad-coverage challenge corpus for sentence understanding through inference
[Williams et al., 2018] Adina Williams, Nikita Nangia, and Samuel Bowman · 2018
Earlier work this paper cites.
Swag: A large-scale adversarial dataset for grounded commonsense inference
[Zellers et al., 2018] Rowan Zellers, Yonatan Bisk, Roy Schwartz, and Yejin Choi · 2018
Earlier work this paper cites.
Electra: Pre-training text encoders as discriminators rather than generators
[Clark et al., 2019] Kevin Clark, Minh-Thang Luong, Quoc V Le, and Christopher D Manning · 2019
Earlier work this paper cites.
Transformer-xl: Attentive language models beyond a fixed-length context
[Dai et al., 2019] Zihang Dai, Zhilin Yang, Yiming Yang, Jaime G Carbonell, Quoc Le, and Ruslan Salakhutdinov · 2019
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
[Devlin et al., 2019] Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
Unified language model pre-training for natural language understanding and generation
[Dong et al., 2019] Li Dong, Nan Yang, Wenhui Wang, Furu Wei, Xiaodong Liu, Yu Wang, Jianfeng Gao, Ming Zhou, and Hsiao-Wuen Hon · 2019
Earlier work this paper cites.
Neural architecture search: A survey
[Elsken et al., 2019] Thomas Elsken, Jan Hendrik Metzen, and Frank Hutter · 2019
Earlier work this paper cites.
Reducing transformer depth on demand with structured dropout
[Fan et al., 2019] Angela Fan, Edouard Grave, and Armand Joulin · 2019
Earlier work this paper cites.
The state of sparsity in deep neural networks
[Gale et al., 2019] Trevor Gale, Erich Elsen, and Sara Hooker · 2019
Earlier work this paper cites.
Rethinking imagenet pre-training
[He et al., 2019] Kaiming He, Ross Girshick, and Piotr Dollár · 2019
Earlier work this paper cites.
Parameter-efficient transfer learning for NLP
[Houlsby et al., 2019] Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly · 2019
Earlier work this paper cites.
Gpipe: Efficient training of giant neural networks using pipeline parallelism
[Huang et al., 2019] Yanping Huang, Youlong Cheng, Ankur Bapna, Orhan Firat, Mia Xu Chen, Dehao Chen, HyoukJoong Lee, Jiquan Ngiam, Quoc V Le, Yonghui Wu, and Zhifeng Chen · 2019
Earlier work this paper cites.
Cross-lingual language model pretraining
[Lample and Conneau, 2019] Guillaume Lample and Alexis Conneau · 2019
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
[Liu et al., 2019] Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Earlier work this paper cites.
Are sixteen heads really better than one?
[Michel et al., 2019] Paul Michel, Omer Levy, and Graham Neubig · 2019
Earlier work this paper cites.
Multi-hop reading comprehension through question decomposition and rescoring
[Min et al., 2019] Sewon Min, Victor Zhong, Luke Zettlemoyer, and Hannaneh Hajishirzi · 2019
Earlier work this paper cites.
Continual lifelong learning with neural networks: A review
[Parisi et al., 2019] German I Parisi, Ronald Kemker, Jose L Part, Christopher Kanan, and Stefan Wermter · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
[Radford et al., 2019] Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Earlier work this paper cites.
Compressive transformers for long-range sequence modelling
[Rae et al., 2019] Jack W Rae, Anna Potapenko, Siddhant M Jayakumar, Chloe Hillier, and Timothy P Lillicrap · 2019
Earlier work this paper cites.
Experience replay for continual learning
[Rolnick et al., 2019] David Rolnick, Arun Ahuja, Jonathan Schwarz, Timothy Lillicrap, and Gregory Wayne · 2019
Earlier work this paper cites.
Human Compatible: Artificial Intelligence and the Problem of Controls
[Russell, 2019] Stuart Russell · 2019
Earlier work this paper cites.
Fast transformer decoding: One write-head is all you need
[Shazeer, 2019] Noam Shazeer · 2019
Earlier work this paper cites.
Megatron-lm: Training multi-billion parameter language models using model parallelism
[Shoeybi et al., 2019] Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro · 2019
Earlier work this paper cites.
Mass: Masked sequence to sequence pre-training for language generation
[Song et al., 2019] Kaitao Song, Xu Tan, Tao Qin, Jianfeng Lu, and Tie-Yan Liu · 2019
Earlier work this paper cites.
Learning deep transformer models for machine translation
[Wang et al., 2019] Qiang Wang, Bei Li, Tong Xiao, Jingbo Zhu, Changliang Li, Derek F Wong, and Lidia S Chao · 2019
Earlier work this paper cites.
Neural network acceptability judgments
[Warstadt et al., 2019] Alex Warstadt, Amanpreet Singh, and Samuel R Bowman · 2019
Earlier work this paper cites.
Sharing attention weights for fast transformer
[Xiao et al., 2019] Tong Xiao, Yinqiao Li, Jingbo Zhu, Zhengtao Yu, and Tongran Liu · 2019
Earlier work this paper cites.
Xlnet: Generalized autoregressive pretraining for language understanding
[Yang et al., 2019] Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R Salakhutdinov, and Quoc V Le · 2019
Earlier work this paper cites.
Root mean square layer normalization
[Zhang and Sennrich, 2019] Biao Zhang and Rico Sennrich · 2019
Earlier work this paper cites.
Etc: Encoding long and structured inputs in transformers
[Ainslie et al., 2020] Joshua Ainslie, Santiago Ontanon, Chris Alberti, Vaclav Cvicek, Zachary Fisher, Philip Pham, Anirudh Ravula, Sumit Sanghai, Qifan Wang, and Li Yang · 2020
Earlier work this paper cites.
Language models are few-shot learners
[Brown et al., 2020] Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Earlier work this paper cites.
The lottery ticket hypothesis for pre-trained bert networks
[Chen et al., 2020] Tianlong Chen, Jonathan Frankle, Shiyu Chang, Sijia Liu, Yang Zhang, Zhangyang Wang, and Michael Carbin · 2020
Earlier work this paper cites.
Unsupervised cross-lingual representation learning at scale
[Conneau et al., 2020] Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Édouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov · 2020
Earlier work this paper cites.
Gmat: Global memory augmentation for transformers
[Gupta and Berant, 2020] Ankit Gupta and Jonathan Berant · 2020
Earlier work this paper cites.
Pretrained transformers improve out-of-distribution robustness
[Hendrycks et al., 2020] Dan Hendrycks, Xiaoyuan Liu, Eric Wallace, Adam Dziedzic, Rishabh Krishnan, and Dawn Song · 2020
Earlier work this paper cites.
The curious case of neural text degeneration
[Holtzman et al., 2020] Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi · 2020
Earlier work this paper cites.
How can we know what language models know?
[Jiang et al., 2020] Zhengbao Jiang, Frank F Xu, Jun Araki, and Graham Neubig · 2020
Earlier work this paper cites.
Tinybert: Distilling bert for natural language understanding
[Jiao et al., 2020] Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu · 2020
Cited alongside, same era.
Spanbert: Improving pre-training by representing and predicting spans
[Joshi et al., 2020] Mandar Joshi, Danqi Chen, Yinhan Liu, Daniel S Weld, Luke Zettlemoyer, and Omer Levy · 2020
Cited alongside, same era.
Scaling laws for neural language models
[Kaplan et al., 2020] Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Cited alongside, same era.
Transformers are rnns: Fast autoregressive transformers with linear attention
[Katharopoulos et al., 2020] Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and François Fleuret · 2020
Cited alongside, same era.
Generalization through memorization: Nearest neighbor language models
[Khandelwal et al., 2020] Urvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer, and Mike Lewis · 2020
Full stack optimization of transformer inference: a survey
[Kim et al., 2023] Sehoon Kim, Coleman Hooper, Thanakul Wattanawong, Minwoo Kang, Ruohan Yan, Hasan Genc, Grace Dinh, Qijing Huang, Kurt Keutzer, Michael W. Mahoney, Yakun Sophia Shao, and Amir Gholami · 2023
Later among the works it cites.
Reducing activation recomputation in large transformer models
[Korthikanti et al., 2023] Vijay Anand Korthikanti, Jared Casper, Sangkug Lym, Lawrence McAfee, Michael Andersch, Mohammad Shoeybi, and Bryan Catanzaro · 2023
Later among the works it cites.
Do models really learn to follow instructions? an empirical study of instruction tuning
[Kung and Peng, 2023] Po-Nien Kung and Nanyun Peng · 2023
Later among the works it cites.
Efficient memory management for large language model serving with pagedattention
[Kwon et al., 2023] Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E Gonzalez, Hao Zhang, and Ion Stoica · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Specification gaming: the flip side of ai ingenuity
[Krakovna et al., 2020] Victoria Krakovna, Jonathan Uesato, Vladimir Mikulik, Matthew Rahtz, Tom Everitt, Ramana Kumar, Zac Kenton, Jan Leike, and Shane Legg · 2020
Cited alongside, same era.
Albert: A lite bert for self-supervised learning of language representations
[Lan et al., 2020] Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut · 2020
Cited alongside, same era.
Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension
[Lewis et al., 2020] Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer · 2020
Cited alongside, same era.
Pre-trained models for natural language processing: A survey
[Qiu et al., 2020] Xipeng Qiu, Tianxiang Sun, Yige Xu, Yunfan Shao, Ning Dai, and Xuanjing Huang · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
[Raffel et al., 2020] Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu · 2020
Cited alongside, same era.
A constructive prediction of the generalization error across scales
[Rosenfeld et al., 2020] Jonathan S Rosenfeld, Amir Rosenfeld, Yonatan Belinkov, and Nir Shavit · 2020
Cited alongside, same era.
Movement pruning: Adaptive sparsity by fine-tuning
[Sanh et al., 2020] Victor Sanh, Thomas Wolf, and Alexander Rush · 2020
Cited alongside, same era.
[Lee et al., 2023] Harrison Lee, Samrat Phatale, Hassan Mansoor, Kellie Ren Lu, Thomas Mesnard, Johan Ferret, Colton Bishop, Ethan Hall, Victor Carbune, and Abhinav Rastogi · 2023
Later among the works it cites.
Fast inference from transformers via speculative decoding
[Leviathan et al., 2023] Yaniv Leviathan, Matan Kalman, and Yossi Matias · 2023
Later among the works it cites.
Deliberate then generate: Enhanced prompting framework for text generation
[Li et al., 2023] Bei Li, Rui Wang, Junliang Guo, Kaitao Song, Xu Tan, Hany Hassan, Arul Menezes, Tong Xiao, Jiang Bian, and JingBo Zhu · 2023
Later among the works it cites.
Sequence parallelism: Long sequence training from system perspective
[Li et al., 2023] Shenggui Li, Fuzhao Xue, Chaitanya Baranwal, Yongbin Li, and Yang You · 2023
Later among the works it cites.
A practical survey on zero-shot prompt design for in-context learning
[Li, 2023] Yinheng Li · 2023
Later among the works it cites.
Compressing context to enhance inference efficiency of large language models
[Li et al., 2023] Yucheng Li, Bo Dong, Frank Guerin, and Chenghua Lin · 2023
Later among the works it cites.
Scaling down to scale up: A guide to parameter-efficient fine-tuning
[Lialin et al., 2023] Vladislav Lialin, Vijeta Deshpande, and Anna Rumshisky · 2023
Later among the works it cites.
Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing
[Liu et al., 2023] Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig · 2023
Later among the works it cites.
Gpt understands, too
[Liu et al., 2023] Xiao Liu, Yanan Zheng, Zhengxiao Du, Ming Ding, Yujie Qian, Zhilin Yang, and Jie Tang · 2023
Later among the works it cites.
Prompting frameworks for large language models: A survey
[Liu et al., 2023] Xiaoxia Liu, Jingyi Wang, Jun Sun, Xiaohan Yuan, Guoliang Dong, Peng Di, Wenhai Wang, and Dongxia Wang · 2023
Later among the works it cites.
The flan collection: Designing data and methods for effective instruction tuning
[Longpre et al., 2023] Shayne Longpre, Le Hou, Tu Vu, Albert Webson, Hyung Won Chung, Yi Tay, Denny Zhou, Quoc V. Le, Barret Zoph, Jason Wei, and Adam Roberts · 2023
Later among the works it cites.
Mega: Moving average equipped gated attention
[Ma et al., 2023] Xuezhe Ma, Chunting Zhou, Xiang Kong, Junxian He, Liangke Gui, Graham Neubig, Jonathan May, and Luke Zettlemoyer · 2023
Later among the works it cites.
Future lens: Anticipating subsequent tokens from a single hidden state
[Pal et al., 2023] Koyena Pal, Jiuding Sun, Andrew Yuan, Byron C Wallace, and David Bau · 2023
Later among the works it cites.
[Penedo et al., 2023] Guilherme Penedo, Quentin Malartic, Daniel Hesslow, Ruxandra Cojocaru, Alessandro Cappelli, Hamza Alobeidli, Baptiste Pannier, Ebtesam Almazrouei, and Julien Launay · 2023
Later among the works it cites.
Efficiently scaling transformer inference
[Pope et al., 2023] Reiner Pope, Sholto Douglas, Aakanksha Chowdhery, Jacob Devlin, James Bradbury, Jonathan Heek, Kefan Xiao, Shivani Agrawal, and Jeff Dean · 2023
Later among the works it cites.
Grips: Gradient-free, edit-based instruction search for prompting large language models
[Prasad et al., 2023] Archiki Prasad, Peter Hase, Xiang Zhou, and Mohit Bansal · 2023
Later among the works it cites.
Measuring and narrowing the compositionality gap in language models
[Press et al., 2023] Ofir Press, Muru Zhang, Sewon Min, Ludwig Schmidt, Noah A Smith, and Mike Lewis · 2023
Later among the works it cites.
Automatic prompt optimization with "gradient descent" and beam search
[Pryzant et al., 2023] Reid Pryzant, Dan Iter, Jerry Li, Yin Tat Lee, Chenguang Zhu, and Michael Zeng · 2023
Later among the works it cites.
PEER: A collaborative language model
[Schick et al., 2023] Timo Schick, Jane A. Yu, Zhengbao Jiang, Fabio Petroni, Patrick Lewis, Gautier Izacard, Qingfei You, Christoforos Nalmpantis, Edouard Grave, and Sebastian Riedel · 2023
Later among the works it cites.
Reflexion: Language agents with verbal reinforcement learning
[Shinn et al., 2023] Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao · 2023
Later among the works it cites.
Stanford alpaca: An instruction-following llama model
[Taori et al., 2023] Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto · 2023
Later among the works it cites.
Openhermes 2.5: An open dataset of synthetic data for generalist llm assistants, 2023
[Teknium, 2023] Teknium · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models
[Touvron et al., 2023] Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
[Touvron et al., 2023] Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom · 2023
Later among the works it cites.
Transformers learn in-context by gradient descent
[Von Oswald et al., 2023] Johannes Von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento, Alexander Mordvintsev, Andrey Zhmoginov, and Max Vladymyrov · 2023
Later among the works it cites.
A comprehensive survey of continual learning: Theory, method and application
[Wang et al., 2023] Liyuan Wang, Xingxing Zhang, Hang Su, and Jun Zhu · 2023
Later among the works it cites.
Self-consistency improves chain of thought reasoning in language models
[Wang et al., 2023] Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V Le, Ed H Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou · 2023
Later among the works it cites.
How far can camels go? exploring the state of instruction tuning on open resources
[Wang et al., 2023] Yizhong Wang, Hamish Ivison, Pradeep Dasigi, Jack Hessel, Tushar Khot, Khyathi Raghavi Chandu, David Wadden, Kelsey MacMillan, Noah A. Smith, Iz Beltagy, and Hannaneh Hajishirzi · 2023
Later among the works it cites.
Self-instruct: Aligning language models with self-generated instructions
[Wang et al., 2023] Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A Smith, Daniel Khashabi, and Hannaneh Hajishirzi · 2023
Later among the works it cites.
A comprehensive survey of forgetting in deep learning beyond continual learning
[Wang et al., 2023] Zhenyi Wang, Enneng Yang, Li Shen, and Heng Huang · 2023
Later among the works it cites.
Generating sequences by learning to self-correct
[Welleck et al., 2023] Sean Welleck, Ximing Lu, Peter West, Faeze Brahman, Tianxiao Shen, Daniel Khashabi, and Yejin Choi · 2023
Later among the works it cites.
Fast distributed inference serving for large language models
[Wu et al., 2023] Bingyang Wu, Yinmin Zhong, Zili Zhang, Shengyu Liu, Fangyue Liu, Yuanhang Sun, Gang Huang, Xuanzhe Liu, and Xin Jin · 2023
Later among the works it cites.
Fine-grained human feedback gives better rewards for language model training
[Wu et al., 2023] Zeqiu Wu, Yushi Hu, Weijia Shi, Nouha Dziri, Alane Suhr, Prithviraj Ammanabrolu, Noah A. Smith, Mari Ostendorf, and Hannaneh Hajishirzi · 2023
Later among the works it cites.
Introduction to transformers: an nlp perspective
[Xiao and Zhu, 2023] Tong Xiao and Jingbo Zhu · 2023
Later among the works it cites.
Towards better chain-of-thought prompting strategies: A survey
[Yu et al., 2023] Zihan Yu, Liang He, Zhen Wu, Xinyu Dai, and Jiajun Chen · 2023
Later among the works it cites.
[Zhang et al., 2023] Zhuosheng Zhang, Yao Yao, Aston Zhang, Xiangru Tang, Xinbei Ma, Zhiwei He, Yiming Wang, Mark Gerstein, Rui Wang, Gongshen Liu, and Hai Zhao · 2023
Later among the works it cites.
Automatic chain of thought prompting in large language models
[Zhang et al., 2023] Zhuosheng Zhang, Aston Zhang, Mu Li, and Alex Smola · 2023
Later among the works it cites.
A survey of large language models
[Zhao et al., 2023] Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, Yifan Du, Chen Yang, Yushuo Chen, Z. Chen, Jinhao Jiang, Ruiyang Ren, Yifan Li, Xinyu Tang, Zikang Liu, Peiyu Liu, Jianyun Nie, and Ji rong Wen · 2023
Later among the works it cites.
Lima: Less is more for alignment
[Zhou et al., 2023] Chunting Zhou, Pengfei Liu, Puxin Xu, Srini Iyer, Jiao Sun, Yuning Mao, Xuezhe Ma, Avia Efrat, Ping Yu, Lili Yu, Susan Zhang, Gargi Ghosh, Mike Lewis, Luke Zettlemoyer, and Omer Levy · 2023
Later among the works it cites.
Least-to-most prompting enables complex reasoning in large language models
[Zhou et al., 2023] Denny Zhou, Nathanael Schärli, Le Hou, Jason Wei, Nathan Scales, Xuezhi Wang, Dale Schuurmans, Claire Cui, Olivier Bousquet, Quoc V. Le, and Ed H. Chi · 2023
Later among the works it cites.
Large language models are human-level prompt engineers
[Zhou et al., 2023] Yongchao Zhou, Andrei Ioan Muresanu, Ziwen Han, Keiran Paster, Silviu Pitis, Harris Chan, and Jimmy Ba · 2023
Later among the works it cites.
Taming { \{ Throughput-Latency } \} tradeoff in { \{ LLM } \} inference with { \{ Sarathi-Serve } \}
[Agrawal et al., 2024] Amey Agrawal, Nitin Kedia, Ashish Panwar, Jayashree Mohan, Nipun Kwatra, Bhargav Gulavani, Alexey Tumanov, and Ramachandran Ramjee · 2024
Later among the works it cites.
cosmopedia: how to create large-scale synthetic data for pre-training
[Allal et al., 2024] Loubna Ben Allal, Anton Lozhkov, and Daniel van Strien · 2024
Later among the works it cites.
Situational awareness: The decade ahead, 2024
[Aschenbrenner, 2024] Leopold Aschenbrenner · 2024
Later among the works it cites.
Managing extreme ai risks amid rapid progress
[Bengio et al., 2024] Yoshua Bengio, Geoffrey Hinton, Andrew Yao, Dawn Song, Pieter Abbeel, Trevor Darrell, Yuval Noah Harari, Ya-Qin Zhang, Lan Xue, Shai Shalev-Shwartz, Gillian K. Hadfield, Jeff Clune, Tegan Maharaj, Frank Hutter, Atilim Gunes Baydin, Sheila A. McIlraith, Qiqi Gao, Ashwin Acharya, David Krueger, Anca Dragan, Philip Torr, Stuart Russell, Daniel Kahneman, Jan Markus Brauner, and Sören Mindermann · 2024
Later among the works it cites.
Graph of thoughts: Solving elaborate problems with large language models
[Besta et al., 2024] Maciej Besta, Nils Blach, Ales Kubicek, Robert Gerstenberger, Michal Podstawski, Lukas Gianinazzi, Joanna Gajda, Tomasz Lehmann, Hubert Niewiadomski, Piotr Nyczyk, and Torsten Hoefler · 2024
Later among the works it cites.
Reducing transformer key-value cache size with cross-layer attention
[Brandon et al., 2024] William Brandon, Mayank Mishra, Aniruddha Nrusimha, Rameswar Panda, and Jonathan Ragan Kelly · 2024
Later among the works it cites.
Large language monkeys: Scaling inference compute with repeated sampling
[Brown et al., 2024] Bradley Brown, Jordan Juravsky, Ryan Ehrlich, Ronald Clark, Quoc V Le, Christopher Ré, and Azalia Mirhoseini · 2024
Later among the works it cites.
Efficient prompting methods for large language models: A survey
[Chang et al., 2024] Kaiyan Chang, Songcheng Xu, Chenglong Wang, Yingfeng Luo, Tong Xiao, and Jingbo Zhu · 2024
Later among the works it cites.
Alpagasus: Training a better alpaca with fewer data
[Chen et al., 2024] Lichang Chen, Shiyang Li, Jun Yan, Hai Wang, Kalpa Gunaratna, Vikas Yadav, Zheng Tang, Vijay Srinivasan, Tianyi Zhou, Heng Huang, and Hongxia Jin · 2024
Later among the works it cites.
Self-play fine-tuning converts weak language models to strong language models
[Chen et al., 2024] Zixiang Chen, Yihe Deng, Huizhuo Yuan, Kaixuan Ji, and Quanquan Gu · 2024
Later among the works it cites.
Reward model ensembles help mitigate overoptimization
[Coste et al., 2024] Thomas Coste, Usman Anwar, Robert Kirk, and David Krueger · 2024
Later among the works it cites.
ULTRAFEEDBACK: Boosting language models with scaled AI feedback
[Cui et al., 2024] Ganqu Cui, Lifan Yuan, Ning Ding, Guanming Yao, Bingxiang He, Wei Zhu, Yuan Ni, Guotong Xie, Ruobing Xie, Yankai Lin, Zhiyuan Liu, and Maosong Sun · 2024
Later among the works it cites.
Language modeling is compression
[Deletang et al., 2024] Gregoire Deletang, Anian Ruoss, Paul-Ambroise Duquenne, Elliot Catt, Tim Genewein, Christopher Mattern, Jordi Grau-Moya, Li Kevin Wenliang, Matthew Aitchison, Laurent Orseau, Marcus Hutter, and Joel Veness · 2024
Later among the works it cites.
Longrope: Extending llm context window beyond 2 million tokens
[Ding et al., 2024] Yiran Ding, Li Lyna Zhang, Chengruidong Zhang, Yuanyuan Xu, Ning Shang, Jiahang Xu, Fan Yang, and Mao Yang · 2024
Later among the works it cites.
[Dubey et al., 2024] Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al · 2024
Later among the works it cites.
Alpacafarm: A simulation framework for methods that learn from human feedback
[Dubois et al., 2024] Yann Dubois, Chen Xuechen Li, Rohan Taori, Tianyi Zhang, Ishaan Gulrajani, Jimmy Ba, Carlos Guestrin, Percy S Liang, and Tatsunori B Hashimoto · 2024
Later among the works it cites.
[Ge et al., 2024] Yuan Ge, Yilun Liu, Chi Hu, Weibin Meng, Shimin Tao, Xiaofeng Zhao, Hongxia Ma, Li Zhang, Boxing Chen, Hao Yang, Bei Li, Tong Xiao, and Jingbo Zhu · 2024
Later among the works it cites.
Gemma: Open Models Based on Gemini Research and Technology, 2024
[Gemma Team, 2024] Google DeepMind Gemma Team · 2024
Later among the works it cites.
Critic: Large language models can self-correct with tool-interactive critiquing
[Gou et al., 2024] Zhibin Gou, Zhihong Shao, Yeyun Gong, Yujiu Yang, Nan Duan, Weizhu Chen, et al · 2024
Later among the works it cites.
Connecting large language models with evolutionary algorithms yields powerful prompt optimizers
[Guo et al., 2024] Qingyan Guo, Rui Wang, Junliang Guo, Bei Li, Kaitao Song, Xu Tan, Guoqing Liu, Jiang Bian, and Yujiu Yang · 2024
Later among the works it cites.
Parameter-efficient fine-tuning for large models: A comprehensive survey
[Han et al., 2024] Zeyu Han, Chao Gao, Jinyang Liu, Jeff Zhang, and Sai Qian Zhang · 2024
Later among the works it cites.
Instruction following without instruction tuning, 2024
[Hewitt, 2024] John Hewitt · 2024
Later among the works it cites.
Instruction following without instruction tuning
[Hewitt et al., 2024] John Hewitt, Nelson F Liu, Percy Liang, and Christopher D Manning · 2024
Later among the works it cites.
Rewardbench: Evaluating reward models for language modeling
[Lambert et al., 2024] Nathan Lambert, Valentina Pyatkin, Jacob Morrison, LJ Miranda, Bill Yuchen Lin, Khyathi Chandu, Nouha Dziri, Sachin Kumar, Tom Zick, Yejin Choi, Noah A. Smith, and Hannaneh Hajishirzi · 2024
Later among the works it cites.
Llm inference serving: Survey of recent advances and opportunities
[Li et al., 2024] Baolin Li, Yankai Jiang, Vijay Gadepally, and Devesh Tiwari · 2024
Later among the works it cites.
Functional interpolation for relative positions improves long context transformers
[Li et al., 2024] Shanda Li, Chong You, Guru Guruganesh, Joshua Ainslie, Santiago Ontanon, Manzil Zaheer, Sumit Sanghai, Yiming Yang, Sanjiv Kumar, and Srinadh Bhojanapalli · 2024
Later among the works it cites.
Let’s verify step by step
[Lightman et al., 2024] Hunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe · 2024
Later among the works it cites.
[Liu et al., 2024] Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, et al · 2024
Later among the works it cites.
Statistical rejection sampling improves preference optimization
[Liu et al., 2024] Tianqi Liu, Yao Zhao, Rishabh Joshi, Misha Khalman, Mohammad Saleh, Peter J Liu, and Jialu Liu · 2024
Later among the works it cites.
Forgetting curve: A reliable method for evaluating memorization capability for long-context models
[Liu et al., 2024] Xinyu Liu, Runsong Zhao, Pengcheng Huang, Chunyang Xiao, Bei Li, Jingang Wang, Tong Xiao, and Jingbo Zhu · 2024
Later among the works it cites.
Megalodon: Efficient llm pretraining and inference with unlimited context length
[Ma et al., 2024] Xuezhe Ma, Xiaomeng Yang, Wenhan Xiong, Beidi Chen, Lili Yu, Hao Zhang, Jonathan May, Luke Zettlemoyer, Omer Levy, and Chunting Zhou · 2024
Later among the works it cites.
Self-refine: Iterative refinement with self-feedback
[Madaan et al., 2024] Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, Shashank Gupta, Bodhisattwa Prasad Majumder, Katherine Hermann, Sean Welleck, Amir Yazdanbakhsh, and Peter Clark · 2024
Later among the works it cites.
Multi-hop question answering
[Mavi et al., 2024] Vaibhav Mavi, Anubhav Jangra, and Adam Jatowt · 2024
Later among the works it cites.
Large language models: A survey
[Minaee et al., 2024] Shervin Minaee, Tomas Mikolov, Narjes Nikzad, Meysam Chenaghlu, Richard Socher, Xavier Amatriain, and Jianfeng Gao · 2024
Later among the works it cites.
Random-access infinite context length for transformers
[Mohtashami and Jaggi, 2024] Amirkeivan Mohtashami and Martin Jaggi · 2024
Later among the works it cites.
Learning to compress prompts with gist tokens
[Mu et al., 2024] Jesse Mu, Xiang Li, and Noah Goodman · 2024
Later among the works it cites.
Leave no context behind: Efficient infinite context transformers with infini-attention
[Munkhdalai et al., 2024] Tsendsuren Munkhdalai, Manaal Faruqui, and Siddharth Gopal · 2024
Later among the works it cites.
Learning to reason with llms, September 2024
[OpenAI, 2024] OpenAI · 2024
Later among the works it cites.
Automatically correcting large language models: Surveying the landscape of diverse automated correction strategies
[Pan et al., 2024] Liangming Pan, Michael Saxon, Wenda Xu, Deepak Nathani, Xinyi Wang, and William Yang Wang · 2024
Later among the works it cites.
Splitwise: Efficient generative llm inference using phase splitting
[Patel et al., 2024] Pratyush Patel, Esha Choukse, Chaojie Zhang, Aashaka Shah, Íñigo Goiri, Saeed Maleki, and Ricardo Bianchini · 2024
Later among the works it cites.
YaRN: Efficient context window extension of large language models
[Peng et al., 2024] Bowen Peng, Jeffrey Quesnelle, Honglu Fan, and Enrico Shippole · 2024
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model
[Rafailov et al., 2024] Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn · 2024
Later among the works it cites.
A survey of llm surveys
[Ruan et al., 2024] Junhao Ruan, Long Meng, Weiqiao Shan, Tong Xiao, and Jingbo Zhu · 2024
Later among the works it cites.
Toolformer: Language models can teach themselves to use tools
[Schick et al., 2024] Timo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu, Maria Lomeli, Eric Hambro, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom · 2024
Later among the works it cites.
Scaling llm test-time compute optimally can be more effective than scaling model parameters
[Snell et al., 2024] Charlie Snell, Jaehoon Lee, Kelvin Xu, and Aviral Kumar · 2024
Later among the works it cites.
Roformer: Enhanced transformer with rotary position embedding
[Su et al., 2024] Jianlin Su, Murtadha Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu · 2024
Later among the works it cites.
Gemma 2: Improving open language models at a practical size
[Team et al., 2024] Gemma Team, Morgane Riviere, Shreya Pathak, Pier Giuseppe Sessa, Cassidy Hardin, Surya Bhupatiraju, Léonard Hussenot, Thomas Mesnard, Bobak Shahriari, Alexandre Ramé, et al · 2024
Later among the works it cites.
Esrl: Efficient sampling-based reinforcement learning for sequence generation
[Wang et al., 2024] Chenglong Wang, Hang Zhou, Yimin Hu, Yifu Huo, Bei Li, Tongran Liu, Tong Xiao, and Jingbo Zhu · 2024
Later among the works it cites.
Do language models plan for future tokens?
[Wu et al., 2024] Wilson Wu, John X Morris, and Lionel Levine · 2024
Later among the works it cites.
Less: Selecting influential data for targeted instruction tuning
[Xia et al., 2024] Mengzhou Xia, Sadhika Malladi, Suchin Gururangan, Sanjeev Arora, and Danqi Chen · 2024
Later among the works it cites.
Efficient streaming language models with attention sinks
[Xiao et al., 2024] Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, and Mike Lewis · 2024
Later among the works it cites.
Wizardlm: Empowering large pre-trained language models to follow complex instructions
[Xu et al., 2024] Can Xu, Qingfeng Sun, Kai Zheng, Xiubo Geng, Pu Zhao, Jiazhan Feng, Chongyang Tao, Qingwei Lin, and Daxin Jiang · 2024
Later among the works it cites.
[Yang et al., 2024] An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, et al · 2024
Later among the works it cites.
Tree of thoughts: Deliberate problem solving with large language models
[Yao et al., 2024] Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Tom Griffiths, Yuan Cao, and Karthik Narasimhan · 2024
Later among the works it cites.
Small language models need strong verifiers to self-correct reasoning
[Zhang et al., 2024] Yunxiang Zhang, Muhammad Khalifa, Lajanugen Logeswaran, Jaekyeom Kim, Moontae Lee, Honglak Lee, and Lu Wang · 2024
Later among the works it cites.
Long is more for alignment: A simple but tough-to-beat baseline for instruction fine-tuning
[Zhao et al., 2024] Hao Zhao, Maksym Andriushchenko, Francesco Croce, and Nicolas Flammarion · 2024
Later among the works it cites.
{ \{ DistServe } \} : Disaggregating prefill and decoding for goodput-optimized large language model serving
[Zhong et al., 2024] Yinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu, Yibo Zhu, Xuanzhe Liu, Xin Jin, and Hao Zhang · 2024
Later among the works it cites.
How scaling laws drive smarter, more powerful ai, 2025
[Briski, 2025] Kari Briski · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
[Deepseek, 2025] Deepseek · 2025
Closest in time.
Nvidia nim llms benchmarking
[Nvidia, 2025] Nvidia · 2025
Closest in time.
Scaling LLM test-time compute optimally can be more effective than scaling parameters for reasoning
[Snell et al., 2025] Charlie Victor Snell, Jaehoon Lee, Kelvin Xu, and Aviral Kumar · 2025
Closest in time.