Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have shown immense potential in enhancing various aspects of our daily lives, from conversational AI to search and AI assistants.
A topological property of real analytic subsets
Stanislaw Lojasiewicz · 1963
Earlier work this paper cites.
Gradient methods for the minimisation of functionals
Boris T Polyak · 1963
Earlier work this paper cites.
Linear convergence of gradient and proximal-gradient methods under the polyak-łojasiewicz condition
Hamed Karimi, Julie Nutini, and Mark Schmidt · 2016
Earlier work this paper cites.
Recovery guarantee of weighted low-rank approximation via alternating minimization
Yuanzhi Li, Yingyu Liang, and Andrej Risteski · 2016
Earlier work this paper cites.
Weighted low rank approximations with provable guarantees
Ilya Razenshteyn, Zhao Song, and David P Woodruff · 2016
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Yuanzhi Li and Yingyu Liang · 2018
Earlier work this paper cites.
Blockwise parallel decoding for deep autoregressive models
Mitchell Stern, Noam Shazeer, and Jakob Uszkoreit · 2018
Earlier work this paper cites.
Fine-grained analysis of optimization and generalization for overparameterized two-layer neural networks
Sanjeev Arora, Simon Du, Wei Hu, Zhiyuan Li, and Ruosong Wang · 2019
Earlier work this paper cites.
Learning and generalization in overparameterized neural networks, going beyond two layers
Zeyuan Allen-Zhu, Yuanzhi Li, and Yingyu Liang · 2019
Earlier work this paper cites.
A convergence theory for deep learning via over-parameterization
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2019
Earlier work this paper cites.
On the convergence rate of training recurrent neural networks
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2019
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
Gradient descent provably optimizes over-parameterized neural networks
Simon S Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh · 2019
Earlier work this paper cites.
Algorithm-dependent generalization bounds for overparameterized deep residual networks
Spencer Frei, Yuan Cao, and Quanquan Gu · 2019
Earlier work this paper cites.
Ziwei Ji and Matus Telgarsky · 2019
Earlier work this paper cites.
Are sixteen heads really better than one?
Paul Michel, Omer Levy, and Graham Neubig · 2019
Earlier work this paper cites.
Generalization error bounds of gradient descent for learning over-parameterized deep relu networks
Yuan Cao and Quanquan Gu · 2020
Earlier work this paper cites.
Kernel density estimation through density constrained near neighbor search
Moses Charikar, Michael Kapralov, Navid Nouri, and Paris Siminelakis · 2020
Earlier work this paper cites.
Finding trainable sparse networks through neural tangent transfer
Tianlin Liu and Friedemann Zenke · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu · 2020
Earlier work this paper cites.
On the opportunities and risks of foundation models
Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al · 2021
Earlier work this paper cites.
Only train once: A one-shot neural network training and pruning framework
Tianyi Chen, Bo Ji, Tianyu Ding, Biyi Fang, Guanyi Wang, Zhihui Zhu, Luming Liang, Yixin Shi, Sheng Yi, and Xiao Tu · 2021
Earlier work this paper cites.
Provable generalization of sgd-trained neural networks of any width in the presence of adversarial label noise
Spencer Frei, Yuan Cao, and Quanquan Gu · 2021
Earlier work this paper cites.
Proxy convexity: A unified framework for the analysis of neural networks trained by gradient descent
Spencer Frei and Quanquan Gu · 2021
Earlier work this paper cites.
Making pre-trained language models better few-shot learners
Tianyu Gao, Adam Fisch, and Danqi Chen · 2021
Earlier work this paper cites.
Sparsity in deep learning: Pruning and growth for efficient inference and training in neural networks
Torsten Hoefler, Dan Alistarh, Tal Ben-Nun, Nikoli Dryden, and Alexandra Peste · 2021
Earlier work this paper cites.
Accelerated sparse neural training: A provable and efficient method to find n: m transposable masks
Itay Hubara, Brian Chmiel, Moshe Island, Ron Banner, Joseph Naor, and Daniel Soudry · 2021
Earlier work this paper cites.
The power of scale for parameter-efficient prompt tuning
Brian Lester, Rami Al-Rfou, and Noah Constant · 2021
Earlier work this paper cites.
Prefix-tuning: Optimizing continuous prompts for generation
Xiang Lisa Li and Percy Liang · 2021
Earlier work this paper cites.
Show your work: Scratchpads for intermediate computation with language models
Maxwell Nye, Anders Johan Andreassen, Gur AriGuy, Henryk Michalewski, Jacob Austin, David Bieber, David Dohan, Aitor Lewkowycz, Maarten Bosma, David Luan, et al · 2021
Earlier work this paper cites.
A theoretical analysis on feature learning in neural networks: Emergence from inputs and advantage over fixed features
Zhenmei Shi, Junyi Wei, and Yingyu Liang · 2021
Earlier work this paper cites.
Calibrate before use: Improving few-shot performance of language models
Zihao Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh · 2021
Earlier work this paper cites.
Optimal-degree polynomial approximations for exponentials and gaussian kernel density estimation
Amol Aggarwal and Josh Alman · 2022
Earlier work this paper cites.
Scaling instruction-finetuned language models
Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, et al · 2022
Earlier work this paper cites.
Palm: Scaling language modeling with pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al · 2022
Earlier work this paper cites.
Flashattention: Fast and memory-efficient exact attention with io-awareness
Tri Dao, Dan Fu, Stefano Ermon, Atri Rudra, and Christopher Ré · 2022
Earlier work this paper cites.
Attentive walk-aggregating graph neural networks
Mehmet F Demirel, Shengchao Liu, Siddhant Garg, Zhenmei Shi, and Yingyu Liang · 2022
Earlier work this paper cites.
A nearly optimal size coreset algorithm with nearly linear time
Yichuan Deng, Zhao Song, Yitan Wang, and Yuanyuan Yang · 2022
Earlier work this paper cites.
Optimal brain compression: A framework for accurate post-training quantization and pruning
Elias Frantar and Dan Alistarh · 2022
Earlier work this paper cites.
LoRA: Low-rank adaptation of large language models
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2022
Earlier work this paper cites.
Pruning’s effect on generalization through the lens of training and regularization
Tian Jin, Michael Carbin, Dan Roy, Jonathan Frankle, and Gintare Karolina Dziugaite · 2022
Earlier work this paper cites.
Demonstrate-search-predict: Composing retrieval and language models for knowledge-intensive nlp
Omar Khattab, Keshav Santhanam, Xiang Lisa Li, David Hall, Percy Liang, Christopher Potts, and Matei Zaharia · 2022
Earlier work this paper cites.
Cross-task generalization via natural language crowdsourcing instructions
Swaroop Mishra, Daniel Khashabi, Chitta Baral, and Hannaneh Hajishirzi · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Earlier work this paper cites.
Black-box tuning for language-model-as-a-service
Tianxiang Sun, Yunfan Shao, Hong Qian, Xuanjing Huang, and Xipeng Qiu · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, and Denny Zhou · 2022
Earlier work this paper cites.
Denoising time cycle modeling for recommendation
Sicong Xie, Qunwei Li, Weidi Xu, Kaiming Shen, Shaohu Chen, and Wenliang Zhong · 2022
Earlier work this paper cites.
Extracting trigger-sharing events via an event matrix
Jun Xu, Weidi Xu, Mengshu Sun, Taifeng Wang, and Wei Chu · 2022
Earlier work this paper cites.
Fast attention requires bounded entries
Josh Alman and Zhao Song · 2023
Earlier work this paper cites.
Sparks of artificial general intelligence: Early experiments with gpt-4
Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, Harsha Nori, Hamid Palangi, Marco Tulio Ribeiro, and Yi Zhang · 2023
Earlier work this paper cites.
Longlora: Efficient fine-tuning of long-context large language models
Yukang Chen, Shengju Qian, Haotian Tang, Xin Lai, Zhijian Liu, Song Han, and Jiaya Jia · 2023
Earlier work this paper cites.
Flashattention-2: Faster attention with better parallelism and work partitioning
Tri Dao · 2023
Earlier work this paper cites.
Unmasking transformers: A theoretical approach to data recovery via attention weights
Yichuan Deng, Zhao Song, Shenghao Xie, and Chiwun Yang · 2023
Earlier work this paper cites.
Sparsegpt: Massive language models can be accurately pruned in one-shot
Elias Frantar and Dan Alistarh · 2023
Earlier work this paper cites.
Mamba: Linear-time sequence modeling with selective state spaces
Albert Gu and Tri Dao · 2023
Earlier work this paper cites.
Llama-adapter v2: Parameter-efficient visual instruction model
Peng Gao, Jiaming Han, Renrui Zhang, Ziyi Lin, Shijie Geng, Aojun Zhou, Wei Zhang, Pan Lu, Conghui He, Xiangyu Yue, et al · 2023
Cited alongside, same era.
An over-parameterized exponential regression
Yeqi Gao, Sridhar Mahadevan, and Zhao Song · 2023
Cited alongside, same era.
Yeqi Gao, Zhao Song, Weixin Wang, and Junze Yin · 2023
Cited alongside, same era.
Yeqi Gao, Zhao Song, and Shenghao Xie · 2023
Cited alongside, same era.
Na Liu, Liangyu Chen, Xiaoyu Tian, Wei Zou, Kaijiang Chen, and Ming Cui · 2024
Closest in time.
Fast john ellipsoid computation with differential privacy optimization
Xiaoyu Li, Yingyu Liang, Zhenmei Shi, Zhao Song, and Junwei Yu · 2024
Closest in time.
Fine-grained attention i/o complexity: Comprehensive analysis for backward passes
Xiaoyu Li, Yingyu Liang, Zhenmei Shi, Zhao Song, and Yufa Zhou · 2024
Closest in time.
Yingyu Liang, Heshan Liu, Zhenmei Shi, Zhao Song, and Junze Yin · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Distilling step-by-step! outperforming larger language models with less training data and smaller model sizes
Cheng-Yu Hsieh, Chun-Liang Li, Chih-kuan Yeh, Hootan Nakhost, Yasuhisa Fujii, Alex Ratner, Ranjay Krishna, Chen-Yu Lee, and Tomas Pfister · 2023
Cited alongside, same era.
On sparse modern hopfield model
Jerry Yao-Chieh Hu, Donglin Yang, Dennis Wu, Chenwei Xu, Bo-Yu Chen, and Han Liu · 2023
Cited alongside, same era.
Lion: Adversarial distillation of proprietary large language models
Yuxin Jiang, Chunkit Chan, Mingyang Chen, and Wei Wang · 2023
Cited alongside, same era.
Sparse finetuning for inference acceleration of large language models
Eldar Kurtic, Denis Kuznedelev, Elias Frantar, Michael Goin, and Dan Alistarh · 2023
Cited alongside, same era.
Polysketchformer: Fast transformers via sketches for polynomial kernels
Praneeth Kacham, Vahab Mirrokni, and Peilin Zhong · 2023
Cited alongside, same era.
Amama Mahmood, Junxiang Wang, Bingsheng Yao, Dakuo Wang, and Chien-Ming Huang · 2023
Cited alongside, same era.
Yarn: Efficient context window extension of large language models
Bowen Peng, Jeffrey Quesnelle, Honglu Fan, and Enrico Shippole · 2023
Cited alongside, same era.
The trade-off between universality and label efficiency of representations from contrastive learning
Zhenmei Shi, Jiefeng Chen, Kunyang Li, Jayaram Raghuram, Xi Wu, Yingyu Liang, and Somesh Jha · 2023
Cited alongside, same era.
Xiaoyu Li, Yingyu Liang, Zhenmei Shi, and Zhao Song · 2024
Closest in time.
Fast second-order method for neural network under small treewidth setting
Xiaoyu Li, Jiangxuan Long, Zhao Song, and Tianyi Zhou · 2024
Closest in time.
Multi-layer transformers gradient can be approximated in almost linear time
Yingyu Liang, Zhizhou Sha, Zhenmei Shi, Zhao Song, and Yufa Zhou · 2024
Closest in time.
Differential privacy mechanisms in neural tangent kernel regression
Yingyu Liang, Zhizhou Sha, Zhenmei Shi, and Zhao Song · 2024
Closest in time.
Toward infinite-long prefix in transformer
Yingyu Liang, Zhenmei Shi, Zhao Song, and Chiwun Yang · 2024
Closest in time.
Differential privacy of cross-attention with provable guarantee
Yingyu Liang, Zhenmei Shi, Zhao Song, and Yufa Zhou · 2024
Closest in time.
Tensor attention training: Provably efficient learning of higher-order transformers
Yingyu Liang, Zhenmei Shi, Zhao Song, and Yufa Zhou · 2024
Closest in time.
Unraveling the smoothness properties of diffusion models: A gaussian mixture perspective
Yingyu Liang, Zhenmei Shi, Zhao Song, and Yufa Zhou · 2024
Closest in time.
Quantum speedups for approximating the john ellipsoid
Xiaoyu Li, Zhao Song, and Junwei Yu · 2024
Closest in time.
AI @ Meta Llama Team · 2024
Closest in time.
Llama 3.2: Revolutionizing edge ai and vision with open, customizable models
Meta · 2024
Closest in time.
Linearizing large language models
Jean Mercat, Igor Vasiljevic, Sedrick Keh, Kushal Arora, Achal Dave, Adrien Gaidon, and Thomas Kollar · 2024
Closest in time.
Megalodon: Efficient llm pretraining and inference with unlimited context length
Xuezhe Ma, Xiaomeng Yang, Wenhan Xiong, Beidi Chen, Lili Yu, Hao Zhang, Jonathan May, Luke Zettlemoyer, Omer Levy, and Chunting Zhou · 2024
Closest in time.
Hello gpt-4o
OpenAI · 2024
Closest in time.
Introducing openai o1-preview
OpenAI · 2024
Closest in time.
Lut-gemm: Quantized matrix multiplication based on luts for efficient inference in large-scale generative language models
Gunho Park, Minsub Kim, Sungjae Lee, Jeonghoon Kim, Beomseok Kwon, Se Jung Kwon, Byeongwook Kim, Youngjoo Lee, Dongsoo Lee, et al · 2024
Closest in time.
Flashattention-3: Fast and accurate attention with asynchrony and low-precision
Jay Shah, Ganesh Bikshandi, Ying Zhang, Vijay Thakkar, Pradeep Ramani, and Tri Dao · 2024
Closest in time.
Representational strengths and limitations of transformers
Clayton Sanford, Daniel J Hsu, and Matus Telgarsky · 2024
Closest in time.
A simple and effective pruning approach for large language models
Mingjie Sun, Zhuang Liu, Anna Bair, and J Zico Kolter · 2024
Closest in time.
Zhenmei Shi, Yifei Ming, Xuan-Phi Nguyen, Yingyu Liang, and Shafiq Joty · 2024
Closest in time.
Lazydit: Lazy learning for the acceleration of diffusion transformers
Xuan Shen, Zhao Song, Yufa Zhou, Bo Chen, Yanyu Li, Yifan Gong, Kai Zhang, Hao Tan, Jason Kuen, Henghui Ding, Zhihao Shu, Wei Niu, Pu Zhao, Yanzhi Wang, and Jiuxiang Gu · 2024
Closest in time.
Numerical pruning for efficient autoregressive models
Xuan Shen, Zhao Song, Yufa Zhou, Bo Chen, Jing Liu, Ruiyi Zhang, Ryan A. Rossi, Hao Tan, Tong Yu, Xiang Chen, Yufan Zhou, Tong Sun, Pu Zhao, Yanzhi Wang, and Jiuxiang Gu · 2024
Closest in time.
Provable guarantees for neural networks via gradient feature learning
Zhenmei Shi, Junyi Wei, and Yingyu Liang · 2024
Closest in time.
Training multi-layer over-parametrized neural network in subquadratic time
Zhao Song, Lichen Zhang, and Ruizhe Zhang · 2024
Closest in time.
Uniform memory retrieval with larger capacity for modern hopfield models
Dennis Wu, Jerry Yao-Chieh Hu, Teng-Yun Hsiao, and Han Liu · 2024
Closest in time.
STanhop: Sparse tandem hopfield model for memory-enhanced time series prediction
Dennis Wu, Jerry Yao-Chieh Hu, Weijian Li, Bo-Yu Chen, and Han Liu · 2024
Closest in time.
Transformers are deep optimizers: Provable in-context learning for deep model training
Weimin Wu, Maojiang Su, Jerry Yao-Chieh Hu, Zhao Song, and Han Liu · 2024
Closest in time.
Sheared llama: Accelerating language model pre-training via structured pruning
Mengzhou Xia, Tianyu Gao, Zhiyuan Zeng, and Danqi Chen · 2024
Closest in time.
Bishop: Bi-directional cellular learning for tabular data with generalized sparse modern hopfield model
Chenwei Xu, Yu-Chao Huang, Jerry Yao-Chieh Hu, Weijian Li, Ammar Gilani, Hsi-Sheng Goan, and Han Liu · 2024
Closest in time.
Do large language models have compositional ability? an investigation into limitations and scalability
Zhuoyan Xu, Zhenmei Shi, and Yingyu Liang · 2024
Closest in time.
Towards few-shot adaptation of foundation models via multitask finetuning
Zhuoyan Xu, Zhenmei Shi, Junyi Wei, Fangzhou Mu, Yin Li, and Yingyu Liang · 2024
Closest in time.
Logicmp: A neuro-symbolic approach for encoding first-order logic constraints
Weidi Xu, Jingwei Wang, Lele Xie, Jianshan He, Hongting Zhou, Taifeng Wang, Xiaopei Wan, Jingdong Chen, Chao Qu, and Wei Chu · 2024
Closest in time.
The hedgehog & the porcupine: Expressive linear attentions with softmax mimicry
Michael Zhang, Kush Bhatia, Hermann Kumbong, and Christopher Ré · 2024
Closest in time.
Qjl: 1-bit quantized jl transform for kv cache quantization with zero overhead
Amir Zandieh, Majid Daliri, and Insu Han · 2024
Closest in time.
Graph unlearning with efficient partial retraining
Jiahao Zhang · 2024
Closest in time.
The expressive power of low-rank adaptation
Yuchen Zeng and Kangwook Lee · 2024
Closest in time.
Step-back prompting enables reasoning via abstraction in large language models
Huaixiu Steven Zheng, Swaroop Mishra, Xinyun Chen, Heng-Tze Cheng, Ed H Chi, Quoc V Le, and Denny Zhou · 2024
Closest in time.
Linear-time graph neural networks for scalable recommendations
Jiahao Zhang, Rui Xue, Wenqi Fan, Xin Xu, Qing Li, Jian Pei, and Xiaorui Liu · 2024
Closest in time.
Dynamic sparse no training: Training-free fine-tuning for sparse llms
Yuxin Zhang, Lirui Zhao, Mingbao Lin, Sun Yunyun, Yiwu Yao, Xingjia Han, Jared Tanner, Shiwei Liu, and Rongrong Ji · 2024
Closest in time.
Bypassing the exponential dependency: Looped transformers efficiently learn in-context by multi-step gradient descent
Bo Chen, Xiaoyu Li, Yingyu Liang, Zhenmei Shi, and Zhao Song · 2025
Closest in time.
The computational limits of state-space models and mamba via the lens of circuit complexity
Yifang Chen, Xiaoyu Li, Yingyu Liang, Zhenmei Shi, and Zhao Song · 2025
Closest in time.
Universal approximation of visual autoregressive transformers
Yifang Chen, Xiaoyu Li, Yingyu Liang, Zhenmei Shi, and Zhao Song · 2025
Closest in time.
Hsr-enhanced sparse attention acceleration
Bo Chen, Yingyu Liang, Zhizhou Sha, Zhenmei Shi, and Zhao Song · 2025
Closest in time.
Streaming kernel pca algorithm with small space
Yichuan Deng, Jiangxuan Long, Zhao Song, Zifan Wang, and Han Zhang · 2025
Closest in time.
Yekun Ke, Xiaoyu Li, Yingyu Liang, Zhizhou Sha, Zhenmei Shi, and Zhao Song · 2025
Closest in time.
Curse of attention: A kernel-based perspective for why transformers fail to generalize on time series forecasting and beyond
Yekun Ke, Yingyu Liang, Zhenmei Shi, Zhao Song, and Chiwun Yang · 2025
Closest in time.
Neural algorithmic reasoning for hypergraphs with looped transformers
Xiaoyu Li, Yingyu Liang, Jiangxuan Long, Zhenmei Shi, Zhao Song, and Zhen Zhuang · 2025
Closest in time.
Fourier circuits in neural networks and transformers: A case study of modular arithmetic with multiple inputs
Chenyang Li, Yingyu Liang, Zhenmei Shi, Zhao Song, and Tianyi Zhou · 2025
Closest in time.
On the computational capability of graph neural networks: A circuit complexity bound perspective
Xiaoyu Li, Yingyu Liang, Zhenmei Shi, Zhao Song, Wei Wang, and Jiahao Zhang · 2025
Closest in time.
When can we solve the weighted low rank approximation problem in truly subquadratic time?
Chenyang Li, Yingyu Liang, Zhenmei Shi, and Zhao Song · 2025
Closest in time.
Looped relu mlps may be all you need as practical programmable computers
Yingyu Liang, Zhizhou Sha, Zhenmei Shi, Zhao Song, and Yufa Zhou · 2025
Closest in time.
Tensor product attention is all you need
Yifan Zhang, Yifeng Liu, Huizhuo Yuan, Zhen Qin, Yang Yuan, Quanquan Gu, and Andrew Chi-Chih Yao · 2025
Closest in time.