Fetching the paper…
Reading the bibliography…
The rapid growth of Large Language Models (LLMs) has been a driving force in transforming various domains, reshaping the artificial general intelligence landscape.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 1901
Earlier work this paper cites.
A tree-based statistical language model for natural language speech recognition
Lalit R Bahl, Peter F Brown, Peter V de Souza, and Robert L Mercer. 1989 · 1989
Earlier work this paper cites.
Statistical theory of learning a rule
Géza Györgyi and Naftali Tishby. 1990 · 1990
Earlier work this paper cites.
Improving retrieval performance by relevance feedback
Gerard Salton and Chris Buckley. 1990 · 1990
Earlier work this paper cites.
Four types of learning curves
Shun-ichi Amari, Naotake Fujita, and Shigeru Shinomoto. 1992 · 1992
Earlier work this paper cites.
A method for disambiguating word senses in a large corpus
William A Gale, Kenneth W Church, and David Yarowsky. 1992 · 1992
Earlier work this paper cites.
Statistical mechanics of learning from examples
Hyunjune Sebastian Seung, Haim Sompolinsky, and Naftali Tishby. 1992 · 1992
Earlier work this paper cites.
Scaling laws in learning of classification tasks
N Barkai, Hyunjune Sebastian Seung, and Haim Sompolinsky. 1993 · 1993
Earlier work this paper cites.
Learning and development in neural networks: The importance of starting small
Jeffrey L Elman. 1993 · 1993
Earlier work this paper cites.
A sequential algorithm for training text classifiers: Corrigendum and additional data. In Acm Sigir Forum , Vol. 29. ACM New York, NY, USA, 13–19
David D Lewis. 1995 · 1995
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Statistical methods for speech recognition
Frederick Jelinek. 1998 · 1998
Earlier work this paper cites.
Two decades of statistical language modeling: Where do we go from here?
Ronald Rosenfeld. 2000 · 2000
Earlier work this paper cites.
Active learning for statistical natural language parsing. In ACL . 120–127
Min Tang, Xiaoqiang Luo, and Salim Roukos. 2002 · 2002
Earlier work this paper cites.
Representative sampling for text classification using support vector machines. In ECIR . 393–407
Zhao Xu, Kai Yu, Volker Tresp, Xiaowei Xu, and Jizhi Wang. 2003 · 2003
Earlier work this paper cites.
Decompositions of a higher-order tensor in block terms—Part II: Definitions and uniqueness
Lieven De Lathauwer. 2008 · 2008
Earlier work this paper cites.
Statistical language models for information retrieval a critical review
ChengXiang Zhai et al · 2008
Earlier work this paper cites.
Curriculum learning. In ICML . 41–48
Yoshua Bengio, Jérôme Louradour, Ronan Collobert, and Jason Weston. 2009 · 2009
Earlier work this paper cites.
Semi-supervised learning (chapelle, o. et al., eds.; 2006)[book reviews]
Olivier Chapelle, Bernhard Scholkopf, and Alexander Zien. 2009 · 2009
Earlier work this paper cites.
Active learning literature survey. University of Wisconsin-Madison Department of Computer Sciences
Burr Settles. 2009 · 2009
Earlier work this paper cites.
Active learning with clustering. In Active Learning and Experimental Design workshop In conjunction with AISTATS 2010 . JMLR Workshop and Conference Proceedings, 127–139
Zalán Bodó, Zsolt Minier, and Lehel Csató. 2011 · 2010
Earlier work this paper cites.
Off to a good start: Using clustering to select the initial training set in active learning
Rong Hu, Brian Mac Namee, and Sarah Jane Delany. 2010 · 2010
Earlier work this paper cites.
Self-paced learning for latent variable models
M Kumar, Benjamin Packer, and Daphne Koller. 2010 · 2010
Earlier work this paper cites.
Recurrent neural network based language model.. In Interspeech , Vol. 2. Makuhari, 1045–1048
Tomas Mikolov, Martin Karafiát, Lukas Burget, Jan Cernockỳ, and Sanjeev Khudanpur. 2010 · 2010
Earlier work this paper cites.
From theories to queries: Active learning in practice. In Active learning and experimental design workshop in conjunction with AISTATS 2010 . JMLR Workshop and Conference Proceedings, 1–18
Burr Settles. 2011 · 2010
Earlier work this paper cites.
Extensions of recurrent neural network language model. In 2011 IEEE international conference on acoustics, speech and signal processing (ICASSP) . IEEE, 5528–5531
Tomáš Mikolov, Stefan Kombrink, Lukáš Burget, Jan Černockỳ, and Sanjeev Khudanpur. 2011 · 2011
Earlier work this paper cites.
Active Learning. In Synthesis Lectures on Artificial Intelligence and Machine Learning . 1–114
Burr Settles. 2012 · 2012
Earlier work this paper cites.
One billion word benchmark for measuring progress in statistical language modeling
Ciprian Chelba, Tomas Mikolov, Mike Schuster, Qi Ge, Thorsten Brants, Phillipp Koehn, and Tony Robinson. 2013 · 2013
Earlier work this paper cites.
Low-rank matrix factorization for deep neural network training with high-dimensional output targets. In 2013 IEEE international conference on acoustics, speech and signal processing . IEEE, 6655–6659
Tara N Sainath, Brian Kingsbury, Vikas Sindhwani, Ebru Arisoy, and Bhuvana Ramabhadran. 2013 · 2013
Earlier work this paper cites.
Do deep nets really need to be deep?
Jimmy Ba and Rich Caruana. 2014 · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba. 2014 · 2014
Earlier work this paper cites.
Mining of Massive Datasets, Cambridge University Press, Cambridge
J Leskovec, A Rajaraman, and JD Ullman. 2014 · 2014
Earlier work this paper cites.
Scaling distributed machine learning with the parameter server. In 11th USENIX Symposium on operating systems design and implementation (OSDI 14) . 583–598
Mu Li, David G Andersen, Jun Woo Park, Alexander J Smola, Amr Ahmed, Vanja Josifovski, James Long, Eugene J Shekita, and Bor-Yiing Su. 2014 · 2014
Earlier work this paper cites.
Fitnets: Hints for thin deep nets
Adriana Romero, Nicolas Ballas, Samira Ebrahimi Kahou, Antoine Chassang, Carlo Gatta, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Samuel R Bowman, Gabor Angeli, Christopher Potts, and Christopher D Manning. 2015 · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. 2015 · 2015
Earlier work this paper cites.
Using random undersampling to alleviate class imbalance on tweet sentiment data. In 2015 IEEE international conference on information reuse and integration . IEEE, 197–202
Joseph Prusa, Taghi M Khoshgoftaar, David J Dittman, and Amri Napolitano. 2015 · 2015
Earlier work this paper cites.
Petuum: A new platform for distributed machine learning on big data. In Proceedings of the 21th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining . 1335–1344
Eric P Xing, Qirong Ho, Wei Dai, Jin-Kyu Kim, Jinliang Wei, Seunghak Lee, Xun Zheng, Pengtao Xie, Abhimanu Kumar, and Yaoliang Yu. 2015 · 2015
Earlier work this paper cites.
Training deep nets with sublinear memory cost
Tianqi Chen, Bing Xu, Chiyuan Zhang, and Carlos Guestrin. 2016 · 2016
Earlier work this paper cites.
Sequence-level knowledge distillation
Yoon Kim and Alexander M Rush. 2016 · 2016
Earlier work this paper cites.
Paying More Attention to Attention: Improving the Performance of Convolutional Neural Networks via Attention Transfer. In ICLR
Sergey Zagoruyko and Nikos Komodakis. 2016 · 2016
Earlier work this paper cites.
Dorefa-net: Training low bitwidth convolutional neural networks with low bitwidth gradients
Shuchang Zhou, Yuxin Wu, Zekun Ni, Xinyu Zhou, He Wen, and Yuheng Zou. 2016 · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei. 2017 · 2017
Earlier work this paper cites.
Deep bayesian active learning with image data. In ICML . PMLR, 1183–1192
Yarin Gal, Riashat Islam, and Zoubin Ghahramani. 2017 · 2017
Earlier work this paper cites.
Deep learning scaling is predictable, empirically
Joel Hestness, Sharan Narang, Newsha Ardalani, Gregory Diamos, Heewoo Jun, Hassan Kianinejad, Md Mostofa Ali Patwary, Yang Yang, and Yanqi Zhou. 2017 · 2017
Earlier work this paper cites.
Quantized neural networks: Training neural networks with low precision weights and activations
Itay Hubara, Matthieu Courbariaux, Daniel Soudry, Ran El-Yaniv, and Yoshua Bengio. 2017 · 2017
Earlier work this paper cites.
Biased importance sampling for deep neural network training
Angelos Katharopoulos and François Fleuret. 2017 · 2017
Earlier work this paper cites.
Curriculum learning and minibatch bucketing in neural machine translation
Tom Kocmi and Ondrej Bojar. 2017 · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter. 2017 · 2017
Earlier work this paper cites.
Paulius Micikevicius, Sharan Narang, Jonah Alben, Gregory Diamos, Erich Elsen, David Garcia, Boris Ginsburg, Michael Houston, Oleksii Kuchaiev, Ganesh Venkatesh, et al · 2017
Earlier work this paper cites.
Active learning for convolutional neural networks: A core-set approach
Ozan Sener and Silvio Savarese. 2017 · 2017
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
A gift from knowledge distillation: Fast optimization, network minimization and transfer learning. In CVPR . 4133–4141
Junho Yim, Donggyu Joo, Jihoon Bae, and Junmo Kim. 2017 · 2017
Earlier work this paper cites.
Pipedream: Fast and efficient pipeline parallel dnn training
Aaron Harlap, Deepak Narayanan, Amar Phanishayee, Vivek Seshadri, Nikhil Devanur, Greg Ganger, and Phil Gibbons. 2018 · 2018
Earlier work this paper cites.
Not all samples are created equal: Deep learning with importance sampling. In International conference on machine learning . PMLR, 2525–2534
Angelos Katharopoulos and François Fleuret. 2018 · 2018
Earlier work this paper cites.
Deep contextualized word representations. In NAACL-HLT . 2227–2237
Matthew E Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018 · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al · 2018
Earlier work this paper cites.
Mesh-tensorflow: Deep learning for supercomputers
Noam Shazeer, Youlong Cheng, Niki Parmar, Dustin Tran, Ashish Vaswani, Penporn Koanantakool, Peter Hawkins, HyoukJoong Lee, Mingsheng Hong, Cliff Young, et al · 2018
Earlier work this paper cites.
Adafactor: Adaptive learning rates with sublinear memory cost. In International Conference on Machine Learning . PMLR, 4596–4604
Noam Shazeer and Mitchell Stern. 2018 · 2018
Earlier work this paper cites.
Aditya Siddhant and Zachary C Lipton. 2018 · 2018
Earlier work this paper cites.
Supporting Very Large Models using Automatic Dataflow Graph Partitioning
Minjie Wang, Chien-chin Huang, and Jinyang Li. 2018 · 2018
Earlier work this paper cites.
Swag: A large-scale adversarial dataset for grounded commonsense inference
Rowan Zellers, Yonatan Bisk, Roy Schwartz, and Yejin Choi. 2018 · 2018
Earlier work this paper cites.
Deep batch active learning by diverse, uncertain gradient lower bounds
Jordan T Ash, Chicheng Zhang, Akshay Krishnamurthy, John Langford, and Alekh Agarwal. 2019 · 2019
Earlier work this paper cites.
Generating long sequences with sparse transformers
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever. 2019 · 2019
Earlier work this paper cites.
Transformer-xl: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc V Le, and Ruslan Salakhutdinov. 2019 · 2019
Earlier work this paper cites.
Learning fast algorithms for linear transforms using butterfly factorizations. In ICML . PMLR, 1517–1527
Tri Dao, Albert Gu, Matthew Eichhorn, Atri Rudra, and Christopher Ré. 2019 · 2019
Earlier work this paper cites.
Discriminative active learning
Daniel Gissin and Shai Shalev-Shwartz. 2019 · 2019
Earlier work this paper cites.
Parameter-efficient transfer learning for NLP. In ICML . PMLR, 2790–2799
Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019 · 2019
Earlier work this paper cites.
Gpipe: Efficient training of giant neural networks using pipeline parallelism
Yanping Huang, Youlong Cheng, Ankur Bapna, Orhan Firat, Dehao Chen, Mia Chen, HyoukJoong Lee, Jiquan Ngiam, Quoc V Le, Yonghui Wu, et al · 2019
Earlier work this paper cites.
Beyond Data and Model Parallelism for Deep Neural Networks
Zhihao Jia, Matei Zaharia, and Alex Aiken. 2019 · 2019
Earlier work this paper cites.
Accelerating deep learning by focusing on the biggest losers
Angela H Jiang, Daniel L-K Wong, Giulio Zhou, David G Andersen, Jeffrey Dean, Gregory R Ganger, Gauri Joshi, Michael Kaminksy, Michael Kozuch, Zachary C Lipton, et al · 2019
Earlier work this paper cites.
Tinybert: Distilling bert for natural language understanding
Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu. 2019 · 2019
Earlier work this paper cites.
A study of BFLOAT16 for deep learning training
Dhiraj Kalamkar, Dheevatsa Mudigere, Naveen Mellempudi, Dipankar Das, Kunal Banerjee, Sasikanth Avancha, Dharma Teja Vooturi, Nataraj Jammalamadaka, Jianyu Huang, Hector Yuen, et al · 2019
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In NAACL-HLT . 4171–4186
Jacob Devlin Ming-Wei Chang Kenton and Lee Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Batchbald: Efficient and diverse batch acquisition for deep bayesian active learning
Andreas Kirsch, Joost Van Amersfoort, and Yarin Gal. 2019 · 2019
Earlier work this paper cites.
Aran Komatsuzaki. 2019 · 2019
Earlier work this paper cites.
Revealing the Dark Secrets of BERT. In EMNLP-IJCNLP . 4365–4374
Olga Kovaleva, Alexey Romanov, Anna Rogers, and Anna Rumshisky. 2019 · 2019
Earlier work this paper cites.
Albert: A lite bert for self-supervised learning of language representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2019 · 2019
Earlier work this paper cites.
What would elsa do? freezing layers during transformer fine-tuning
Jaejun Lee, Raphael Tang, and Jimmy Lin. 2019 · 2019
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 2019
Earlier work this paper cites.
A tensorized transformer for language modeling
Xindian Ma, Peng Zhang, Shuai Zhang, Nan Duan, Yuexian Hou, Ming Zhou, and Dawei Song. 2019 · 2019
Earlier work this paper cites.
PipeDream: Generalized pipeline parallelism for DNN training. In Proceedings of the 27th ACM Symposium on Operating Systems Principles . 1–15
Deepak Narayanan, Aaron Harlap, Amar Phanishayee, Vivek Seshadri, Nikhil R Devanur, Gregory R Ganger, Phillip B Gibbons, and Matei Zaharia. 2019 · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
A constructive prediction of the generalization error across scales
Jonathan S Rosenfeld, Amir Rosenfeld, Yonatan Belinkov, and Nir Shavit. 2019 · 2019
Earlier work this paper cites.
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019 · 2019
Earlier work this paper cites.
Megatron-lm: Training multi-billion parameter language models using model parallelism
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. 2019 · 2019
Earlier work this paper cites.
CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge. In ACL . ACL, Minneapolis, Minnesota, 4149–4158
Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. 2019 · 2019
Earlier work this paper cites.
Autoassist: A framework to accelerate training of deep neural networks
Jiong Zhang, Hsiang-Fu Yu, and Inderjit S Dhillon. 2019 · 2019
Earlier work this paper cites.
Diverse mini-batch active learning
Fedor Zhdanov. 2019 · 2019
Earlier work this paper cites.
Fine-tuning language models from human preferences
Daniel M Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving. 2019 · 2019
Earlier work this paper cites.
Nur Ahmed and Muntasir Wahed. 2020 · 2020
Earlier work this paper cites.
ETC: Encoding long and structured inputs in transformers
Joshua Ainslie, Santiago Ontanon, Chris Alberti, Vaclav Cvicek, Zachary Fisher, Philip Pham, Anirudh Ravula, Sumit Sanghai, Qifan Wang, and Li Yang. 2020 · 2020
Earlier work this paper cites.
Binarybert: Pushing the limit of bert quantization
Haoli Bai, Wei Zhang, Lu Hou, Lifeng Shang, Jing Jin, Xin Jiang, Qun Liu, Michael Lyu, and Irwin King. 2020 · 2020
Earlier work this paper cites.
Zeroq: A novel zero shot quantization framework. In CVPR . 13169–13178
Yaohui Cai, Zhewei Yao, Zhen Dong, Amir Gholami, Michael W Mahoney, and Kurt Keutzer. 2020 · 2020
Earlier work this paper cites.
Neural network compression via sparse optimization
Tianyi Chen, Bo Ji, Yixin Shi, Tianyu Ding, Biyi Fang, Sheng Yi, and Xiao Tu. 2020b · 2020
Earlier work this paper cites.
Half-space proximal stochastic gradient method for group-sparsity regularized problem
Tianyi Chen, Guanyi Wang, Tianyu Ding, Bo Ji, Sheng Yi, and Zhihui Zhu. 2020c · 2020
Earlier work this paper cites.
Making pre-trained language models better few-shot learners
Tianyu Gao, Adam Fisch, and Danqi Chen. 2020 · 2020
Earlier work this paper cites.
Aˆ 3: Accelerating attention mechanisms in neural networks with approximation. In HPCA . 328–341
Tae Jun Ham, Sung Jun Jung, Seonghak Kim, Young H Oh, Yeonhong Park, Yoonho Song, Jung-Hun Park, Sanghee Lee, Kyoung Park, Jae W Lee, et al · 2020
Earlier work this paper cites.
Scaling laws for autoregressive generative modeling
Tom Henighan, Jared Kaplan, Mor Katz, Mark Chen, Christopher Hesse, Jacob Jackson, Heewoo Jun, Tom B Brown, Prafulla Dhariwal, Scott Gray, et al · 2020
Earlier work this paper cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020 · 2020
Earlier work this paper cites.
ALBERT: A Lite BERT for Self-supervised Learning of Language Representations. In ICLR
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2020 · 2020
Earlier work this paper cites.
BioBERT: a pre-trained biomedical language representation model for biomedical text mining
Jinhyuk Lee, Wonjin Yoon, Sungdong Kim, Donghyeon Kim, Sunkyu Kim, Chan Ho So, and Jaewoo Kang. 2020 · 2020
Earlier work this paper cites.
Gshard: Scaling giant models with conditional computation and automatic sharding
Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer, and Zhifeng Chen. 2020 · 2020
Earlier work this paper cites.
BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension. In ACL . 7871–7880
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020 · 2020
Earlier work this paper cites.
PyTorch distributed: experiences on accelerating data parallel training
Shen Li, Yanli Zhao, Rohan Varma, Omkar Salpekar, Pieter Noordhuis, Teng Li, Adam Paszke, Jeff Smith, Brian Vaughan, Pritam Damania, et al · 2020
Earlier work this paper cites.
Fastbert: a self-distilling bert with adaptive inference time
Weijie Liu, Peng Zhou, Zhe Zhao, Zhiruo Wang, Haotang Deng, and Qi Ju. 2020b · 2020
Earlier work this paper cites.
MESA: boost ensemble imbalanced learning with meta-sampler
Zhining Liu, Pengfei Wei, Jing Jiang, Wei Cao, Jiang Bian, and Yi Chang. 2020a · 2020
Earlier work this paper cites.
Do we need to create big datasets to learn a task?. In SustaiNLP Workshop . 169–173
Swaroop Mishra and Bhavdeep Singh Sachdeva. 2020 · 2020
Earlier work this paper cites.
Compressing pre-trained language models by matrix decomposition. In Proceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th International Joint Conference on Natural Language Processing . 884–889
Matan Ben Noach and Yoav Goldberg. 2020 · 2020
Earlier work this paper cites.
Adapterhub: A framework for adapting transformers
Jonas Pfeiffer, Andreas Rücklé, Clifton Poth, Aishwarya Kamath, Ivan Vulić, Sebastian Ruder, Kyunghyun Cho, and Iryna Gurevych. 2020a · 2020
Earlier work this paper cites.
Mad-x: An adapter-based framework for multi-task cross-lingual transfer
Jonas Pfeiffer, Ivan Vulić, Iryna Gurevych, and Sebastian Ruder. 2020b · 2020
Earlier work this paper cites.
Shortformer: Better language modeling using shorter inputs
Ofir Press, Noah A Smith, and Mike Lewis. 2020 · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020a · 2020
Cited alongside, same era.
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020b · 2020
Cited alongside, same era.
Zero: Memory optimizations toward training trillion parameter models. In SC . 1–16
Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He. 2020 · 2020
Cited alongside, same era.
Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters. In SIGKDD . 3505–3506
Jeff Rasley, Samyam Rajbhandari, Olatunji Ruwase, and Yuxiong He. 2020 · 2020
Cited alongside, same era.
Q-bert: Hessian based ultra low precision quantization of bert. In AAAI , Vol. 34. 8815–8821
Sheng Shen, Zhen Dong, Jiayu Ye, Linjian Ma, Zhewei Yao, Amir Gholami, Michael W Mahoney, and Kurt Keutzer. 2020 · 2020
SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
Amey Agrawal, Ashish Panwar, Jayashree Mohan, Nipun Kwatra, Bhargav S Gulavani, and Ramachandran Ramjee. 2023 · 2023
Closest in time.
GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints
Joshua Ainslie, James Lee-Thorp, Michiel de Jong, Yury Zemlyanskiy, Federico Lebrón, and Sumit Sanghai. 2023 · 2023
Closest in time.
Gpt4all: Training an assistant-style chatbot with large scale data distillation from gpt-3.5-turbo
Yuvanesh Anand, Zach Nussbaum, Brandon Duderstadt, Benjamin Schmidt, and Andriy Mulyar. 2023 · 2023
Closest in time.
Rohan Anil, Andrew M Dai, Orhan Firat, Melvin Johnson, Dmitry Lepikhin, Alexandre Passos, Siamak Shakeri, Emanuel Taropa, Paige Bailey, Zhifeng Chen, et al · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Autoprompt: Eliciting knowledge from language models with automatically generated prompts
Taylor Shin, Yasaman Razeghi, Robert L Logan IV, Eric Wallace, and Sameer Singh. 2020 · 2020
Cited alongside, same era.
The role of artificial intelligence in achieving the Sustainable Development Goals
Ricardo Vinuesa, Hossein Azizpour, Iolanda Leite, Madeline Balaam, Virginia Dignum, Sami Domisch, Anna Felländer, Simone Daniela Langhans, Max Tegmark, and Francesco Fuso Nerini. 2020 · 2020
Cited alongside, same era.
Self-paced learning for neural machine translation
Yu Wan, Baosong Yang, Derek F Wong, Yikai Zhou, Lidia S Chao, Haibo Zhang, and Boxing Chen. 2020 · 2020
Cited alongside, same era.
CCNet: Extracting High Quality Monolingual Datasets from Web Crawl Data. In Proceedings of The 12th Language Resources and Evaluation Conference . 4003–4012
Guillaume Wenzek, Marie-Anne Lachaux, Alexis Conneau, Vishrav Chaudhary, Francisco Guzmán, Armand Joulin, and Édouard Grave. 2020 · 2020
Cited alongside, same era.
Automatic neural network compression by sparsity-quantization joint learning: A constrained optimization-based approach. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 2178–2188
Haichuan Yang, Shupeng Gui, Yuhao Zhu, and Ji Liu. 2020 · 2020
Cited alongside, same era.
Cold-start active learning through self-supervised language modeling
Michelle Yuan, Hsuan-Tien Lin, and Jordan Boyd-Graber. 2020 · 2020
Cited alongside, same era.
Big bird: Transformers for longer sequences
Manzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontanon, Philip Pham, Anirudh Ravula, Qifan Wang, Li Yang, et al · 2020
Cited alongside, same era.
Maciej Besta, Nils Blach, Ales Kubicek, Robert Gerstenberger, Lukas Gianinazzi, Joanna Gajda, Tomasz Lehmann, Michal Podstawski, Hubert Niewiadomski, Piotr Nyczyk, et al · 2023
Closest in time.
Sparks of artificial general intelligence: Early experiments with gpt-4
Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, et al · 2023
Closest in time.
A survey on evaluation of large language models
Yupeng Chang, Xu Wang, Jindong Wang, Yuan Wu, Kaijie Zhu, Hao Chen, Linyi Yang, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang, et al · 2023
Closest in time.
MCC-KD: Multi-CoT Consistent Knowledge Distillation
Hongzhan Chen, Siyue Wu, Xiaojun Quan, Rui Wang, Ming Yan, and Ji Zhang. 2023f · 2023
Closest in time.
Maybe Only 0.5% Data is Needed: A Preliminary Exploration of Low Training Data Instruction Tuning
Hao Chen, Yiming Zhang, Qi Zhang, Hantao Yang, Xiaomeng Hu, Xuetao Ma, Yifan Yanggong, and Junbo Zhao. 2023i · 2023
Closest in time.
Extending context window of large language models via positional interpolation
Shouyuan Chen, Sherman Wong, Liangjian Chen, and Yuandong Tian. 2023e · 2023
Closest in time.
LoRAShear: Efficient Large Language Model Structured Pruning and Knowledge Recovery
Tianyi Chen, Tianyu Ding, Badal Yadav, Ilya Zharkov, and Luming Liang. 2023a · 2023
Closest in time.
Towards Automatic Neural Architecture Search within General Super-Networks
Tianyi Chen, Luming Liang, Tianyu Ding, and Ilya Zharkov. 2023c · 2023
Closest in time.
Sparse MoE as the New Dropout: Scaling Dense and Self-Slimmable Transformers
Tianlong Chen, Zhenyu Zhang, Ajay Jaiswal, Shiwei Liu, and Zhangyang Wang. 2023h · 2023
Closest in time.
Pipeline MoE: A Flexible MoE Implementation with Pipeline Parallelism
Xin Chen, Hengheng Zhang, Xiaotao Gu, Kaifeng Bi, Lingxi Xie, and Qi Tian. 2023g · 2023
Closest in time.
Dissecting transformer length extrapolation via the lens of receptive field analysis. In ACL . 13522–13537
Ta-Chung Chi, Ting-Han Fan, Alexander Rudnicky, and Peter Ramadge. 2023 · 2023
Closest in time.
Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90%* ChatGPT Quality
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. 2023 · 2023
Closest in time.
John Joon Young Chung, Ece Kamar, and Saleema Amershi. 2023 · 2023
Closest in time.
InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale Fung, and Steven C. H. Hoi. 2023 · 2023
Closest in time.
Flashattention-2: Faster attention with better parallelism and work partitioning
Tri Dao. 2023 · 2023
Closest in time.
Flash-Decoding for Long-Context Inference
Tri Dao, Daniel Haziza, Francisco Massa, and Grigory Sizov. 2023 · 2023
Closest in time.
Qlora: Efficient finetuning of quantized llms
Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer. 2023a · 2023
Closest in time.
SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression
Tim Dettmers, Ruslan Svirschevski, Vage Egiazarian, Denis Kuznedelev, Elias Frantar, Saleh Ashkboos, Alexander Borzunov, Torsten Hoefler, and Dan Alistarh. 2023b · 2023
Closest in time.
Longnet: Scaling transformers to 1,000,000,000 tokens
Jiayu Ding, Shuming Ma, Li Dong, Xingxing Zhang, Shaohan Huang, Wenhui Wang, and Furu Wei. 2023b · 2023
Closest in time.
Sparse Low-rank Adaptation of Pre-trained Language Models
Ning Ding, Xingtai Lv, Qiaosen Wang, Yulin Chen, Bowen Zhou, Zhiyuan Liu, and Maosong Sun. 2023a · 2023
Closest in time.
TinyStories: How Small Can Language Models Be and Still Speak Coherent English?
Ronen Eldan and Yuanzhi Li. 2023 · 2023
Closest in time.
Depgraph: Towards any structural pruning
Gongfan Fang, Xinyin Ma, Mingli Song, Michael Bi Mi, and Xinchao Wang. 2023 · 2023
Closest in time.
Monarch Mixer: A Simple Sub-Quadratic GEMM-Based Architecture. In NeurIPS
Daniel Y Fu, Simran Arora, Jessica Grogan, Isys Johnson, Sabri Eyuboglu, Armin W Thomas, Benjamin Spector, Michael Poli, Atri Rudra, and Christopher Ré. 2023 · 2023
Closest in time.
Cramming: Training a Language Model on a single GPU in one day.. In ICML . PMLR, 11117–11143
Jonas Geiping and Tom Goldstein. 2023 · 2023
Closest in time.
Mamba: Linear-Time Sequence Modeling with Selective State Spaces
Albert Gu and Tri Dao. 2023 · 2023
Closest in time.
Knowledge Distillation of Large Language Models
Yuxian Gu, Li Dong, Furu Wei, and Minlie Huang. 2023 · 2023
Closest in time.
Suriya Gunasekar, Yi Zhang, Jyoti Aneja, Caio César Teodoro Mendes, Allie Del Giorno, Sivakanth Gopi, Mojan Javaheripi, Piero Kauffmann, Gustavo de Rosa, Olli Saarikivi, et al · 2023
Closest in time.
How good are gpt models at machine translation? a comprehensive evaluation
Amr Hendy, Mohamed Abdelrehim, Amr Sharaf, Vikas Raunak, Mohamed Gabr, Hitokazu Matsushita, Young Jin Kim, Mohamed Afify, and Hany Hassan Awadalla. 2023 · 2023
Closest in time.
FlashDecoding++: Faster Large Language Model Inference on GPUs
Ke Hong, Guohao Dai, Jiaming Xu, Qiuli Mao, Xiuhong Li, Jun Liu, Kangdi Chen, Hanyu Dong, and Yu Wang. 2023 · 2023
Closest in time.
Cheng-Yu Hsieh, Chun-Liang Li, Chih-Kuan Yeh, Hootan Nakhost, Yasuhisa Fujii, Alexander Ratner, Ranjay Krishna, Chen-Yu Lee, and Tomas Pfister. 2023 · 2023
Closest in time.
Output Sensitivity-Aware DETR Quantization
Yafeng Huang, Huanrui Yang, Zhen Dong, Denis Gudovskiy, Tomoyuki Okuno, Yohei Nakata, Yuan Du, Shanghang Zhang, and Kurt Keutzer. 2023 · 2023
Closest in time.
The Emergence of Essential Sparsity in Large Pre-trained Models: The Weights that Matter
Ajay Jaiswal, Shiwei Liu, Tianlong Chen, and Zhangyang Wang. 2023 · 2023
Closest in time.
Lion: Adversarial Distillation of Closed-Source Large Language Model
Yuxin Jiang, Chunkit Chan, Mingyang Chen, and Wei Wang. 2023 · 2023
Closest in time.
The Impact of Positional Encoding on Length Generalization in Transformers
Amirhossein Kazemnejad, Inkit Padhi, Karthikeyan Natesan Ramamurthy, Payel Das, and Siva Reddy. 2023 · 2023
Closest in time.
SqueezeLLM: Dense-and-Sparse Quantization
Sehoon Kim, Coleman Hooper, Amir Gholami, Zhen Dong, Xiuyu Li, Sheng Shen, Michael W Mahoney, and Kurt Keutzer. 2023a · 2023
Closest in time.
Reducing activation recomputation in large transformer models
Vijay Anand Korthikanti, Jared Casper, Sangkug Lym, Lawrence McAfee, Michael Andersch, Mohammad Shoeybi, and Bryan Catanzaro. 2023 · 2023
Closest in time.
Sparse Finetuning for Inference Acceleration of Large Language Models
Eldar Kurtic, Denis Kuznedelev, Elias Frantar, Michael Goin, and Dan Alistarh. 2023 · 2023
Closest in time.
Efficient memory management for large language model serving with pagedattention. In Proceedings of the 29th Symposium on Operating Systems Principles . 611–626
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica. 2023 · 2023
Closest in time.
Symbolic Chain-of-Thought Distillation: Small Models Can Also" Think" Step-by-Step
Liunian Harold Li, Jack Hessel, Youngjae Yu, Xiang Ren, Kai-Wei Chang, and Yejin Choi. 2023b · 2023
Closest in time.
Functional Interpolation for Relative Positions Improves Long Context Transformers
Shanda Li, Chong You, Guru Guruganesh, Joshua Ainslie, Santiago Ontanon, Manzil Zaheer, Sumit Sanghai, Yiming Yang, Sanjiv Kumar, and Srinadh Bhojanapalli. 2023d · 2023
Closest in time.
Q-diffusion: Quantizing diffusion models
Xiuyu Li, Long Lian, Yijiang Liu, Huanrui Yang, Zhen Dong, Daniel Kang, Shanghang Zhang, and Kurt Keutzer. 2023c · 2023
Closest in time.
Textbooks Are All You Need II: phi-1.5 technical report
Yuanzhi Li, Sébastien Bubeck, Ronen Eldan, Allie Del Giorno, Suriya Gunasekar, and Yin Tat Lee. 2023a · 2023
Closest in time.
AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration
Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Xingyu Dang, and Song Han. 2023 · 2023
Closest in time.
Blockwise Parallel Transformer for Long Context Large Models
Hao Liu and Pieter Abbeel. 2023 · 2023
Closest in time.
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023a · 2023
Closest in time.
Ring Attention with Blockwise Transformers for Near-Infinite Context
Hao Liu, Matei Zaharia, and Pieter Abbeel. 2023d · 2023
Closest in time.
Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing
Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. 2023c · 2023
Closest in time.
GPT understands, too
Xiao Liu, Yanan Zheng, Zhengxiao Du, Ming Ding, Yujie Qian, Zhilin Yang, and Jie Tang. 2023e · 2023
Closest in time.
LLM-QAT: Data-Free Quantization Aware Training for Large Language Models
Zechun Liu, Barlas Oguz, Changsheng Zhao, Ernie Chang, Pierre Stock, Yashar Mehdad, Yangyang Shi, Raghuraman Krishnamoorthi, and Vikas Chandra. 2023b · 2023
Closest in time.
Large Language Model Guided Tree-of-Thought
Jieyi Long. 2023 · 2023
Closest in time.
LightFormer: Light-weight Transformer Using SVD-based Weight Transfer and Parameter Sharing. In Findings of the Association for Computational Linguistics: ACL 2023 . 10323–10335
Xiuqing Lv, Peng Zhang, Sunzhu Li, Guobing Gan, and Yueheng Sun. 2023 · 2023
Closest in time.
LLM-Pruner: On the Structural Pruning of Large Language Models
Xinyin Ma, Gongfan Fang, and Xinchao Wang. 2023a · 2023
Closest in time.
Mega: moving average equipped gated attention
Xuezhe Ma, Chunting Zhou, Xiang Kong, Junxian He, Liangke Gui, Graham Neubig, Jonathan May, and Luke Zettlemoyer. 2023b · 2023
Closest in time.
Long range language modeling via gated state spaces
Harsh Mehta, Ankit Gupta, Ashok Cutkosky, and Behnam Neyshabur. 2023 · 2023
Closest in time.
Scaling Data-Constrained Language Models
Niklas Muennighoff, Alexander M Rush, Boaz Barak, Teven Le Scao, Aleksandra Piktus, Nouamane Tazi, Sampo Pyysalo, Thomas Wolf, and Colin Raffel. 2023 · 2023
Closest in time.
INT-FP-QSim: Mixed Precision and Formats For Large Language Models and Vision Transformers
Lakshmi Nair, Mikhail Bernadskiy, Arulselvan Madhavan, Craig Chan, Ayon Basumallik, and Darius Bunandar. 2023 · 2023
Closest in time.
OpenAI. 2023a · 2023
Closest in time.
Introducing ChatGPT
OpenAI. 2023b · 2023
Closest in time.
RWKV: Reinventing RNNs for the Transformer Era
Bo Peng, Eric Alcaide, Quentin Anthony, Alon Albalak, Samuel Arcadinho, Huanqi Cao, Xin Cheng, Michael Chung, Matteo Grella, Kranthi Kiran GV, et al · 2023
Closest in time.
Towards Efficient and Effective Adaptation of Large Language Models for Sequential Recommendation
Bo Peng, Ben Burns, Ziqi Chen, Srinivasan Parthasarathy, and Xia Ning. 2023b · 2023
Closest in time.
Yarn: Efficient context window extension of large language models
Bowen Peng, Jeffrey Quesnelle, Honglu Fan, and Enrico Shippole. 2023c · 2023
Closest in time.
Hyena hierarchy: Towards larger convolutional language models
Michael Poli, Stefano Massaroli, Eric Nguyen, Daniel Y Fu, Tri Dao, Stephen Baccus, Yoshua Bengio, Stefano Ermon, and Christopher Ré. 2023 · 2023
Closest in time.
Train short, test long: Attention with linear biases enables input length extrapolation
Ofir Press, Noah A Smith, and Mike Lewis. 2023 · 2023
Closest in time.
Randomized Positional Encodings Boost Length Generalization of Transformers
Anian Ruoss, Grégoire Delétang, Tim Genewein, Jordi Grau-Moya, Róbert Csordás, Mehdi Bennani, Shane Legg, and Joel Veness. 2023 · 2023
Closest in time.
Promptmix: A class boundary augmentation method for large language model distillation
Gaurav Sahu, Olga Vechtomova, Dzmitry Bahdanau, and Issam H Laradji. 2023 · 2023
Closest in time.
OUTLIER WEIGHED LAYERWISE SPARSITY (OWL): AMissing SECRET SAUCE FOR PRUNING LLMS TO HIGH SPARSITY
AMissing SECRET SAUCE. 2023 · 2023
Closest in time.
One-Shot Sensitivity-Aware Mixed Sparsity Pruning for Large Language Models
Hang Shao, Bei Liu, and Yanmin Qian. 2023 · 2023
Closest in time.
On Efficient Training of Large-Scale Deep Learning Models: A Literature Review
Li Shen, Yan Sun, Zhiyuan Yu, Liang Ding, Xinmei Tian, and Dacheng Tao. 2023b · 2023
Closest in time.
Flan-MoE: Scaling Instruction-Finetuned Language Models with Sparse Mixture of Experts
Sheng Shen, Le Hou, Yanqi Zhou, Nan Du, Shayne Longpre, Jason Wei, Hyung Won Chung, Barret Zoph, William Fedus, Xinyun Chen, et al · 2023
Closest in time.
Simplified state space layers for sequence modeling
Jimmy TH Smith, Andrew Warrington, and Scott W Linderman. 2023 · 2023
Closest in time.
Pandagpt: One model to instruction-follow them all
Yixuan Su, Tian Lan, Huayang Li, Jialu Xu, Yan Wang, and Deng Cai. 2023 · 2023
Closest in time.
A Simple and Effective Pruning Approach for Large Language Models
Mingjie Sun, Zhuang Liu, Anna Bair, and J Zico Kolter. 2023b · 2023
Closest in time.
Retentive network: A successor to transformer for large language models
Yutao Sun, Li Dong, Shaohan Huang, Shuming Ma, Yuqing Xia, Jilong Xue, Jianyong Wang, and Furu Wei. 2023a · 2023
Closest in time.
Stanford Alpaca: An Instruction-following LLaMA model
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023 · 2023
Closest in time.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Closest in time.
Efficient methods for natural language processing: A survey
Marcos Treviso, Ji-Ung Lee, Tianchu Ji, Betty van Aken, Qingqing Cao, Manuel R Ciosici, Michael Hassid, Kenneth Heafield, Sara Hooker, Colin Raffel, et al · 2023
Closest in time.
SCOTT: Self-consistent chain-of-thought distillation
Peifeng Wang, Zhengyang Wang, Zheng Li, Yifan Gao, Bing Yin, and Xiang Ren. 2023a · 2023
Closest in time.
A prompt pattern catalog to enhance prompt engineering with chatgpt
Jules White, Quchen Fu, Sam Hays, Michael Sandborn, Carlos Olea, Henry Gilbert, Ashraf Elnashar, Jesse Spencer-Smith, and Douglas C Schmidt. 2023 · 2023
Closest in time.
Lamini-lm: A diverse herd of distilled models from large-scale instructions
Minghao Wu, Abdul Waheed, Chiyu Zhang, Muhammad Abdul-Mageed, and Alham Fikri Aji. 2023b · 2023
Closest in time.
Bloomberggpt: A large language model for finance
Shijie Wu, Ozan Irsoy, Steven Lu, Vadim Dabravolski, Mark Dredze, Sebastian Gehrmann, Prabhanjan Kambadur, David Rosenberg, and Gideon Mann. 2023a · 2023
Closest in time.
Haojun Xia, Zhen Zheng, Yuchao Li, Donglin Zhuang, Zhongzhu Zhou, Xiafei Qiu, Yong Li, Wei Lin, and Shuaiwen Leon Song. 2023b · 2023
Closest in time.
Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning
Mengzhou Xia, Tianyu Gao, Zhiyuan Zeng, and Danqi Chen. 2023a · 2023
Closest in time.
Smoothquant: Accurate and efficient post-training quantization for large language models. In ICML . PMLR, 38087–38099
Guangxuan Xiao, Ji Lin, Mickael Seznec, Hao Wu, Julien Demouth, and Song Han. 2023 · 2023
Closest in time.
Data selection for language models via importance resampling
Sang Michael Xie, Shibani Santurkar, Tengyu Ma, and Percy Liang. 2023b · 2023
Closest in time.
Decomposition enhances reasoning via self-evaluation guided decoding
Yuxi Xie, Kenji Kawaguchi, Yiran Zhao, Xu Zhao, Min-Yen Kan, Junxian He, and Qizhe Xie. 2023a · 2023
Closest in time.
A survey on model compression and acceleration for pretrained language models. In AAAI , Vol. 37. 10566–10575
Canwen Xu and Julian McAuley. 2023 · 2023
Closest in time.
An efficient 2d method for training super-large deep learning models. In 2023 IEEE International Parallel and Distributed Processing Symposium (IPDPS) . IEEE, 222–232
Qifan Xu and Yang You. 2023 · 2023
Closest in time.
Making pre-trained language models end-to-end few-shot learners with contrastive prompt tuning. In Proceedings of the Sixteenth ACM International Conference on Web Search and Data Mining . 438–446
Ziyun Xu, Chengyu Wang, Minghui Qiu, Fuli Luo, Runxin Xu, Songfang Huang, and Jun Huang. 2023 · 2023
Closest in time.
Harnessing the power of llms in practice: A survey on chatgpt and beyond
Jingfeng Yang, Hongye Jin, Ruixiang Tang, Xiaotian Han, Qizhang Feng, Haoming Jiang, Bing Yin, and Xia Hu. 2023a · 2023
Closest in time.
Exploring the limits of chatgpt for query or aspect-based text summarization
Xianjun Yang, Yan Li, Xinlu Zhang, Haifeng Chen, and Wei Cheng. 2023b · 2023
Closest in time.
Tree of thoughts: Deliberate problem solving with large language models
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L Griffiths, Yuan Cao, and Karthik Narasimhan. 2023 · 2023
Closest in time.
EdgeMoE: Fast On-Device Inference of MoE-based Large Language Models
Rongjie Yi, Liwei Guo, Shiyun Wei, Ao Zhou, Shangguang Wang, and Mengwei Xu. 2023 · 2023
Closest in time.
RPTQ: Reorder-based Post-training Quantization for Large Language Models
Zhihang Yuan, Lin Niu, Jiawei Liu, Wenyu Liu, Xinggang Wang, Yuzhang Shang, Guangyu Sun, Qiang Wu, Jiaxiang Wu, and Bingzhe Wu. 2023 · 2023
Closest in time.
Data-centric artificial intelligence: A survey
Daochen Zha, Zaid Pervaiz Bhat, Kwei-Herng Lai, Fan Yang, Zhimeng Jiang, Shaochen Zhong, and Xia Hu. 2023 · 2023
Closest in time.
Lifting the Curse of Capacity Gap in Distilling Language Models
Chen Zhang, Yang Yang, Jiahao Liu, Jingang Wang, Yunsen Xian, Benyou Wang, and Dawei Song. 2023d · 2023
Closest in time.
Pruning Meets Low-Rank Parameter-Efficient Fine-Tuning
Mingyang Zhang, Chunhua Shen, Zhen Yang, Linlin Ou, Xinyi Yu, Bohan Zhuang, et al · 2023
Closest in time.
Adaptive budget allocation for parameter-efficient fine-tuning
Qingru Zhang, Minshuo Chen, Alexander Bukharin, Pengcheng He, Yu Cheng, Weizhu Chen, and Tuo Zhao. 2023a · 2023
Closest in time.
Instruction Tuning for Large Language Models: A Survey
Shengyu Zhang, Linfeng Dong, Xiaoya Li, Sen Zhang, Xiaofei Sun, Shuhe Wang, Jiwei Li, Runyi Hu, Tianwei Zhang, Fei Wu, et al · 2023
Closest in time.
Dynamic Sparse No Training: Training-Free Fine-tuning for Sparse LLMs
Yuxin Zhang, Lirui Zhao, Mingbao Lin, Yunyun Sun, Yiwu Yao, Xingjia Han, Jared Tanner, Shiwei Liu, and Rongrong Ji. 2023e · 2023
Closest in time.
A survey of large language models
Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, et al · 2023
Closest in time.
MAT: mixed-strategy game of adversarial training in fine-tuning
Zhehua Zhong, Tianyi Chen, and Zhen Wang. 2023 · 2023
Closest in time.
A comprehensive survey on pretrained foundation models: A history from bert to chatgpt
Ce Zhou, Qian Li, Chen Li, Jun Yu, Yixin Liu, Guangjing Wang, Kai Zhang, Cheng Ji, Qiben Yan, Lifang He, et al · 2023
Closest in time.
Minigpt-4: Enhancing vision-language understanding with advanced large language models
Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny. 2023a · 2023
Closest in time.
A Survey on Model Compression for Large Language Models
Xunyu Zhu, Jian Li, Yong Liu, Can Ma, and Weiping Wang. 2023b · 2023
Closest in time.
Physics of Language Models: Part 3.1, Knowledge Storage and Extraction
Zeyuan Allen Zhu and Yuanzhi Li. 2023 · 2023
Closest in time.
A survey on efficient training of transformers
Bohan Zhuang, Jing Liu, Zizheng Pan, Haoyu He, Yuetian Weng, and Chunhua Shen. 2023 · 2023
Closest in time.
DistiLLM: Towards Streamlined Distillation for Large Language Models
Jongwoo Ko, Sungnyun Kim, Tianyi Chen, and Se-Young Yun. 2024 · 2024
Closest in time.
Are NLP Models really able to Solve Simple Math Word Problems?. In ACL . ACL, 2080–2094
Arkil Patel, Satwik Bhattamishra, and Navin Goyal. 2021 · 2094
Closest in time.