Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs), despite their recent impressive accomplishments, are notably cost-prohibitive to deploy, particularly for applications involving long-content generation, such as dialogue systems and story writing.
A study of replacement algorithms for a virtual-storage computer
Laszlo A. Belady · 1966
Earlier work this paper cites.
An anomaly in space-time characteristics of certain programs running in a paging machine
Laszlo A Belady, Robert A Nelson, and Gerald S Shedler · 1969
Earlier work this paper cites.
An analysis of approximations for maximizing submodular set functions—i
George L Nemhauser, Laurence A Wolsey, and Marshall L Fisher · 1978
Earlier work this paper cites.
The lru-k page replacement algorithm for database disk buffering
Elizabeth J O’neil, Patrick E O’neil, and Gerhard Weikum · 1993
Earlier work this paper cites.
Lrfu: A spectrum of policies that subsumes the least recently used and least frequently used policies
Donghee Lee, Jongmoo Choi, Jong-Hun Kim, Sam H Noh, Sang Lyul Min, Yookun Cho, and Chong Sang Kim · 2001
Earlier work this paper cites.
Combinatorial optimization: polyhedra and efficiency
Alexander Schrijver · 2003
Earlier work this paper cites.
Beyond convexity: Submodularity in machine learning
Andreas Krause and Carlos Guestrin · 2008
Earlier work this paper cites.
Choice of plausible alternatives: An evaluation of commonsense causal reasoning
Melissa Roemmele, Cosmin Adrian Bejan, and Andrew S Gordon · 2011
Earlier work this paper cites.
Song Han, Huizi Mao, and William J Dally · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, Jeff Dean, et al · 2015
Earlier work this paper cites.
Submodularity in machine learning applications
Jeff Bilmes · 2015
Earlier work this paper cites.
Abstractive text summarization using sequence-to-sequence rnns and beyond
Ramesh Nallapati, Bowen Zhou, Caglar Gulcehre, Bing Xiang, et al · 2016
Earlier work this paper cites.
Pruning convolutional neural networks for resource efficient inference
Pavlo Molchanov, Stephen Tyree, Tero Karras, Timo Aila, and Jan Kautz · 2016
Earlier work this paper cites.
Heavy hitters via cluster-preserving clustering
Kasper Green Larsen, Jelani Nelson, Huy L Nguyen, and Mikkel Thorup · 2016
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam M. Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2017
Earlier work this paper cites.
Can a suit of armor conduct electricity? a new dataset for open book question answering
Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal · 2018
Earlier work this paper cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman · 2018
Earlier work this paper cites.
Shashi Narayan, Shay B Cohen, and Mirella Lapata · 2018
Earlier work this paper cites.
Quantization and training of neural networks for efficient integer-arithmetic-only inference
Benoit Jacob, Skirmantas Kligys, Bo Chen, Menglong Zhu, Matthew Tang, Andrew Howard, Hartwig Adam, and Dmitry Kalenichenko · 2018
Earlier work this paper cites.
Rethinking the value of network pruning
Zhuang Liu, Mingjie Sun, Tinghui Zhou, Gao Huang, and Trevor Darrell · 2018
Earlier work this paper cites.
Commonsenseqa: A question answering challenge targeting commonsense knowledge
Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Texygen: A benchmarking platform for text generation models
Yaoming Zhu, Sidi Lu, Lei Zheng, Jiaxian Guo, Weinan Zhang, Jun Wang, and Yong Yu · 2018
Earlier work this paper cites.
Generating long sequences with sparse transformers
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever · 2019
Earlier work this paper cites.
Fast transformer decoding: One write-head is all you need
Noam Shazeer · 2019
Earlier work this paper cites.
Resurrecting submodularity for neural text generation
Simeng Han, Xiang Lin, and Shafiq Joty · 2019
Earlier work this paper cites.
MathQA: Towards interpretable math word problem solving with operation-based formalisms
Aida Amini, Saadia Gabriel, Shanchuan Lin, Rik Koncel-Kedziorski, Yejin Choi, and Hannaneh Hajishirzi · 2019
Earlier work this paper cites.
Data-free quantization through weight equalization and bias correction
Markus Nagel, Mart van Baalen, Tijmen Blankevoort, and Max Welling · 2019
Earlier work this paper cites.
Improving neural network quantization without retraining using outlier channel splitting
Ritchie Zhao, Yuwei Hu, Jordan Dotzel, Chris De Sa, and Zhiru Zhang · 2019
Earlier work this paper cites.
Filter pruning via geometric median for deep convolutional neural networks acceleration
Yang He, Ping Liu, Ziwei Wang, Zhilan Hu, and Yi Yang · 2019
Earlier work this paper cites.
On the efficacy of knowledge distillation
Jang Hyun Cho and Bharath Hariharan · 2019
Earlier work this paper cites.
Distilling task-specific knowledge from bert into simple neural networks
Raphael Tang, Yao Lu, Linqing Liu, Lili Mou, Olga Vechtomova, and Jimmy Lin · 2019
Earlier work this paper cites.
Xlnet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R Salakhutdinov, and Quoc V Le · 2019
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Earlier work this paper cites.
End-to-end open-domain question answering with bertserini
Wei Yang, Yuqing Xie, Aileen Lin, Xingyu Li, Luchen Tan, Kun Xiong, Ming Li, and Jimmy Lin · 2019
Earlier work this paper cites.
Cognitive graph for multi-hop reading comprehension at scale
Ming Ding, Chang Zhou, Qibin Chen, Hongxia Yang, and Jie Tang · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
On the variance of the adaptive learning rate and beyond
Liyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen, Xiaodong Liu, Jianfeng Gao, and Jiawei Han · 2019
Earlier work this paper cites.
Stronger l2/l2 compressed sensing; without iterating
Vasileios Nakos and Zhao Song · 2019
Earlier work this paper cites.
(nearly) sample-optimal sparse fourier transform in any dimension; ripless and filterless
Vasileios Nakos, Zhao Song, and Zhengyu Wang · 2019
Earlier work this paper cites.
Resurrecting submodularity for neural text generation
Simeng Han, Xiang Lin, and Shafiq Joty · 2019
Earlier work this paper cites.
Solving linear programs in the current matrix multiplication time
Michael B Cohen, Yin Tat Lee, and Zhao Song · 2019
Earlier work this paper cites.
Solving empirical risk minimization in the current matrix multiplication time
Yin Tat Lee, Zhao Song, and Qiuyi Zhang · 2019
Earlier work this paper cites.
Reformer: The efficient transformer
Nikita Kitaev, Łukasz Kaiser, and Anselm Levskaya · 2020
Earlier work this paper cites.
Rethinking attention with performers
Krzysztof Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlos, Peter Hawkins, Jared Davis, Afroz Mohiuddin, Lukasz Kaiser, et al · 2020
Earlier work this paper cites.
Transformers are rnns: Fast autoregressive transformers with linear attention
Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and François Fleuret · 2020
Earlier work this paper cites.
Efficient transformers: A survey
Yi Tay, Mostafa Dehghani, Dara Bahri, and Donald Metzler · 2020
Earlier work this paper cites.
Piqa: Reasoning about physical commonsense in natural language
Yonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao, and Yejin Choi · 2020
Cited alongside, same era.
Compressive transformers for long-range sequence modelling
Jack W Rae, Anna Potapenko, Siddhant M Jayakumar, and Timothy P Lillicrap · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu · 2020
Cited alongside, same era.
Language models are few-shot learners
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Cited alongside, same era.
Why {adam} beats {sgd} for attention models, 2020
Jingzhao Zhang, Sai Praneeth Karimireddy, Andreas Veit, Seungyeon Kim, Sashank J Reddi, Sanjiv Kumar, and Suvrit Sra · 2020
Cited alongside, same era.
Learning to compress prompts with gist tokens
Jesse Mu, Xiang Lisa Li, and Noah Goodman · 2023
Closest in time.
High-throughput generative inference of large language models with a single gpu
Ying Sheng, Lianmin Zheng, Binhang Yuan, Zhuohan Li, Max Ryabinin, Daniel Y Fu, Zhiqiang Xie, Beidi Chen, Clark Barrett, Joseph E Gonzalez, et al · 2023
Closest in time.
Massive language models can be accurately pruned in one-shot
Elias Frantar and Dan Alistarh · 2023
Closest in time.
A simple and effective pruning approach for large language models
Mingjie Sun, Zhuang Liu, Anna Bair, and J Zico Kolter · 2023
Closest in time.
Outlier weighed layerwise sparsity (owl): A missing secret sauce for pruning llms to high sparsity
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Understanding the difficulty of training transformers
Liyuan Liu, Xiaodong Liu, Jianfeng Gao, Weizhu Chen, and Jiawei Han · 2020
Cited alongside, same era.
Very deep transformers for neural machine translation
Xiaodong Liu, Kevin Duh, Liyuan Liu, and Jianfeng Gao · 2020
Cited alongside, same era.
Optimizing deeper transformers on small datasets
Peng Xu, Dhruv Kumar, Wei Yang, Wenjie Zi, Keyi Tang, Chenyang Huang, Jackie Chi Kit Cheung, Simon JD Prince, and Yanshuai Cao · 2020
Cited alongside, same era.
A faster interior point method for semidefinite programming
Haotian Jiang, Tarun Kathuria, Yin Tat Lee, Swati Padmanabhan, and Zhao Song · 2020
Cited alongside, same era.
S Cliff Liu, Zhao Song, Hengjie Zhang, Lichen Zhang, and Tianyi Zhou · 2020
Cited alongside, same era.
A framework for few-shot language model evaluation
Leo Gao, Jonathan Tow, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Kyle McDonell, Niklas Muennighoff, Jason Phang, Laria Reynolds, Eric Tang, Anish Thite, Ben Wang, Kevin Wang, and Andy Zou · 2021
Cited alongside, same era.
Spatten: Efficient sparse attention architecture with cascade token and head pruning
Hanrui Wang, Zhekai Zhang, and Song Han · 2021
Cited alongside, same era.
Lu Yin, You Wu, Zhenyu Zhang, Cheng-Yu Hsieh, Yaqing Wang, Yiling Jia, Mykola Pechenizkiy, Yi Liang, Zhangyang Wang, and Shiwei Liu · 2023
Closest in time.
Awq: Activation-aware weight quantization for llm compression and acceleration
Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Xingyu Dang, and Song Han · 2023
Closest in time.
Colt5: Faster long-range transformers with conditional computation
Joshua Ainslie, Tao Lei, Michiel de Jong, Santiago Ontañón, Siddhartha Brahma, Yury Zemlyanskiy, David Uthus, Mandy Guo, James Lee-Thorp, Yi Tay, et al · 2023
Closest in time.
Dynamic context pruning for efficient and interpretable autoregressive transformers
Sotiris Anagnostidis, Dario Pavllo, Luca Biggio, Lorenzo Noci, Aurelien Lucchi, and Thomas Hoffmann · 2023
Closest in time.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Closest in time.
Alpacaeval: An automatic evaluator of instruction-following models
Xuechen Li, Tianyi Zhang, Yann Dubois, Rohan Taori, Ishaan Gulrajani, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto · 2023
Closest in time.
Judging llm-as-a-judge with mt-bench and chatbot arena, 2023
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric. P Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica · 2023
Closest in time.
Efficient streaming language models with attention sinks
Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, and Mike Lewis · 2023
Closest in time.
Lm-infinite: Simple on-the-fly length generalization for large language models
Chi Han, Qifan Wang, Wenhan Xiong, Yu Chen, Heng Ji, and Sinong Wang · 2023
Closest in time.
Harnessing the power of llms in practice: A survey on chatgpt and beyond
Jingfeng Yang, Hongye Jin, Ruixiang Tang, Xiaotian Han, Qizhang Feng, Haoming Jiang, Bing Yin, and Xia Hu · 2023
Closest in time.
Lost in the middle: How language models use long contexts
Nelson F Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang · 2023
Closest in time.
Improving length-generalization in transformers via task hinting
Pranjal Awasthi and Anupam Gupta · 2023
Closest in time.
Kdeformer: Accelerating transformers via kernel density estimation
Amir Zandieh, Insu Han, Majid Daliri, and Amin Karbasi · 2023
Closest in time.
Fast attention requires bounded entries
Josh Alman and Zhao Song · 2023
Closest in time.
Superiority of softmax: Unveiling the performance edge over linear attention
Yichuan Deng, Zhao Song, and Tianyi Zhou · 2023
Closest in time.
Algorithm and hardness for dynamic attention maintenance in large language models
Jan van den Brand, Zhao Song, and Tianyi Zhou · 2023
Closest in time.
Differentially private attention computation
Yeqi Gao, Zhao Song, and Xin Yang · 2023
Closest in time.
Representational strengths and limitations of transformers
Clayton Sanford, Daniel Hsu, and Matus Telgarsky · 2023
Closest in time.
Randomized and deterministic attention sparsification algorithms for over-parameterized feature dimension
Yichuan Deng, Sridhar Mahadevan, and Zhao Song · 2023
Closest in time.
Solving regularized exp, cosh and sinh regression problems
Zhihang Li, Zhao Song, and Tianyi Zhou · 2023
Closest in time.
Attention scheme inspired softmax regression
Yichuan Deng, Zhihang Li, and Zhao Song · 2023
Closest in time.
Polysketchformer: Fast transformers via sketches for polynomial kernels
Praneeth Kacham, Vahab Mirrokni, and Peilin Zhong · 2023
Closest in time.
An iterative algorithm for rescaled hyperbolic functions regression
Yeqi Gao, Zhao Song, and Junze Yin · 2023
Closest in time.
Hyperattention: Long-context attention in near-linear time
Insu Han, Rajesh Jarayam, Amin Karbasi, Vahab Mirrokni, David P Woodruff, and Amir Zandieh · 2023
Closest in time.
How to protect copyright data in optimization of large language models?
Timothy Chu, Zhao Song, and Chiwun Yang · 2023
Closest in time.
Ritwik Sinha, Zhao Song, and Tianyi Zhou · 2023
Closest in time.
Yeqi Gao, Zhao Song, Weixin Wang, and Junze Yin · 2023
Closest in time.
Fast quantum algorithm for attention computation
Yeqi Gao, Zhao Song, Xin Yang, and Ruizhe Zhang · 2023
Closest in time.
On the optimization and generalization of multi-head attention
Puneesh Deora, Rouzbeh Ghaderi, Hossein Taheri, and Christos Thrampoulidis · 2023
Closest in time.
Gradientcoin: A peer-to-peer decentralized large language models
Yeqi Gao, Zhao Song, and Junze Yin · 2023
Closest in time.
Unmasking transformers: A theoretical approach to data recovery via attention weights
Yichuan Deng, Zhao Song, Shenghao Xie, and Chiwun Yang · 2023
Closest in time.
Zero-th order algorithm for softmax attention optimization
Yichuan Deng, Zhihang Li, Sridhar Mahadevan, and Zhao Song · 2023
Closest in time.
Josh Alman and Zhao Song · 2023
Closest in time.
Convergence of two-layer regression with nonlinear units
Yichuan Deng, Zhao Song, and Shenghao Xie · 2023
Closest in time.
Fine-tune language models to approximate unbiased in-context learning
Timothy Chu, Zhao Song, and Chiwun Yang · 2023
Closest in time.
Trainable transformer in transformer
Abhishek Panigrahi, Sadhika Malladi, Mengzhou Xia, and Sanjeev Arora · 2023
Closest in time.
Do transformers parse while predicting the masked word?
Haoyu Zhao, Abhishek Panigrahi, Rong Ge, and Sanjeev Arora · 2023
Closest in time.
Fast submodular function maximization
Lianke Qin, Zhao Song, and Yitan Wang · 2023
Closest in time.
Infoprompt: Information-theoretic soft prompt tuning for natural language understanding
Junda Wu, Tong Yu, Rui Wang, Zhao Song, Ruiyi Zhang, Handong Zhao, Chaochao Lu, Shuai Li, and Ricardo Henao · 2023
Closest in time.
The closeness of in-context learning and weight shifting for softmax regression
Shuai Li, Zhao Song, Yu Xia, Tong Yu, and Tianyi Zhou · 2023
Closest in time.
A nearly-linear time algorithm for structured support vector machines
Yuzhou Gu, Zhao Song, and Lichen Zhang · 2023
Closest in time.
An online and unified algorithm for projection matrix vector multiplication with application to empirical risk minimization
Lianke Qin, Zhao Song, Lichen Zhang, and Danyang Zhuo · 2023
Closest in time.
Streaming semidefinite programs: O ( n ) {O}(\sqrt{n}) passes, small space and fast runtime
Zhao Song, Mingquan Ye, and Lichen Zhang · 2023
Closest in time.
Convex minimization with integer minima in O ~ ( n 4 ) \widetilde{O}(n^{4}) time
Haotian Jiang, Yin Tat Lee, Zhao Song, and Lichen Zhang · 2024
Closest in time.