Fetching the paper…
Reading the bibliography…
Recently, Large Language Models (LLMs) have achieved remarkable success.
An isserlis’ theorem for mixed gaussian variables: Application to the auto-bispectral density
JV Michalowicz, JM Nichols, F Bucholtz, and CC Olson · 2009
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom B Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Earlier work this paper cites.
Deep learning scaling is predictable, empirically
Joel Hestness, Sharan Narang, Newsha Ardalani, Gregory Diamos, Heewoo Jun, Hassan Kianinejad, Md Mostofa Ali Patwary, Yang Yang, and Yanqi Zhou · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Gradient descent provably optimizes over-parameterized neural networks
Simon S Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh · 2018
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Earlier work this paper cites.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Yuanzhi Li and Yingyu Liang · 2018
Earlier work this paper cites.
On exact computation with an infinitely wide neural net
Sanjeev Arora, Simon S Du, Wei Hu, Zhiyuan Li, Ruslan Salakhutdinov, and Ruosong Wang · 2019
Earlier work this paper cites.
Learning and generalization in overparameterized neural networks, going beyond two layers
Zeyuan Allen-Zhu, Yuanzhi Li, and Yingyu Liang · 2019
Earlier work this paper cites.
A convergence theory for deep learning via over-parameterization
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2019
Earlier work this paper cites.
On the convergence rate of training recurrent neural networks
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2019
Earlier work this paper cites.
Reconciling modern machine-learning practice and the classical bias–variance trade-off
Mikhail Belkin, Daniel Hsu, Siyuan Ma, and Soumik Mandal · 2019
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
Surprises in high-dimensional ridgeless least squares interpolation
Trevor Hastie, Andrea Montanari, Saharon Rosset, and Ryan J Tibshirani · 2019
Earlier work this paper cites.
The generalization error of random features regression: Precise asymptotics and double descent curve
Song Mei and Andrea Montanari · 2019
Earlier work this paper cites.
A constructive prediction of the generalization error across scales
Jonathan S Rosenfeld, Amir Rosenfeld, Yonatan Belinkov, and Nir Shavit · 2019
Earlier work this paper cites.
High-dimensional statistics: A non-asymptotic viewpoint
Martin J Wainwright · 2019
Earlier work this paper cites.
Benign overfitting in linear regression
Peter L Bartlett, Philip M Long, Gábor Lugosi, and Alexander Tsigler · 2020
Earlier work this paper cites.
Language models are few-shot learners
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Scaling laws for autoregressive generative modeling
Tom Henighan, Jared Kaplan, Mor Katz, Mark Chen, Christopher Hesse, Jacob Jackson, Heewoo Jun, Tom B Brown, Prafulla Dhariwal, Scott Gray, et al · 2020
Earlier work this paper cites.
Instahide’s sample complexity when mixing two private images
Baihe Huang, Zhao Song, Runzhou Tao, Junze Yin, Ruizhe Zhang, and Danyang Zhuo · 2020
Earlier work this paper cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu · 2020
Earlier work this paper cites.
A neural scaling law from the dimension of the data manifold
Utkarsh Sharma and Jared Kaplan · 2020
Earlier work this paper cites.
Explaining neural scaling laws
Yasaman Bahri, Ethan Dyer, Jared Kaplan, Jaehoon Lee, and Utkarsh Sharma · 2021
Earlier work this paper cites.
On the opportunities and risks of foundation models
Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al · 2021
Earlier work this paper cites.
Generalization error rates in kernel regression: The crossover from the noiseless to noisy regime
Stefano Camuto and Nicolas Flammarion · 2021
Earlier work this paper cites.
Making pre-trained language models better few-shot learners
Tianyu Gao, Adam Fisch, and Danqi Chen · 2021
Earlier work this paper cites.
The power of scale for parameter-efficient prompt tuning
Brian Lester, Rami Al-Rfou, and Noah Constant · 2021
Earlier work this paper cites.
The power of scale for parameter-efficient prompt tuning
Brian Lester, Rami Al-Rfou, and Noah Constant · 2021
Earlier work this paper cites.
Prediction of shear wave velocity using machine learning technique, multiple regression and well logs
Lin Shi and Jiachen Zhang · 2021
Earlier work this paper cites.
Calibrate before use: Improving few-shot performance of language models
Zihao Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh · 2021
Earlier work this paper cites.
Optimal-degree polynomial approximations for exponentials and gaussian kernel density estimation
Amol Aggarwal and Josh Alman · 2022
Earlier work this paper cites.
A nearly optimal size coreset algorithm with nearly linear time
Yichuan Deng, Zhao Song, Yitan Wang, and Yuanyuan Yang · 2022
Earlier work this paper cites.
Training compute-optimal large language models
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al · 2022
Earlier work this paper cites.
Sublinear time algorithm for online weighted bipartite matching
Hang Hu, Zhao Song, Runzhou Tao, Zhaozhuo Xu, Junze Yin, and Danyang Zhuo · 2022
Earlier work this paper cites.
A dynamic fast gaussian transform
Baihe Huang, Zhao Song, Omri Weinstein, Junze Yin, Hengjie Zhang, and Ruizhe Zhang · 2022
Earlier work this paper cites.
LoRA: Low-rank adaptation of large language models
Edward J Hu, yelong shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2022
Earlier work this paper cites.
A faster k k -means++ algorithm
Jiehao Liang, Somdeb Sarkhel, Zhao Song, Chenbo Yin, Junze Yin, and Danyang Zhuo · 2022
Earlier work this paper cites.
Dynamic maintenance of kernel density estimation data structure: From practice to theory
Jiehao Liang, Zhao Song, Zhaozhuo Xu, Junze Yin, and Danyang Zhuo · 2022
Earlier work this paper cites.
Evaluation of multiple linear regression and machine learning approaches to predict soil compaction and shear stress based on electrical parameters
Katarzyna Pentoś, Jasper Tembeck Mbah, Krzysztof Pieczarka, Gniewko Niedbała, and Tomasz Wojciechowski · 2022
Earlier work this paper cites.
Multiple regression model to analyze the total los for patients undergoing laparoscopic appendectomy
Teresa Angela Trunfio, Arianna Scala, Cristiana Giglio, Giovanni Rossi, Anna Borrelli, Maria Romano, and Giovanni Improta · 2022
Earlier work this paper cites.
Multiple regression
David Weisburd, David B Wilson, Alese Wooditch, Chester Britt, David Weisburd, David B Wilson, Alese Wooditch, and Chester Britt · 2022
Earlier work this paper cites.
The power and limitation of pretraining-finetuning for linear regression under covariate shift
Jingfeng Wu, Difan Zou, Vladimir Braverman, Quanquan Gu, and Sham Kakade · 2022
Earlier work this paper cites.
Clean energy investment and financial development as determinants of environment and sustainable economic growth: evidence from china
Zahid Zahoor, Irfan Khan, and Fujun Hou · 2022
Earlier work this paper cites.
Scaling vision transformers
Xiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, and Lucas Beyer · 2022
Cited alongside, same era.
A multi-layer extreme learning machine refined by sparrow search algorithm and weighted mean filter for short-term multi-step wind speed forecasting
Haochen Zhang, Zhiyun Peng, Junjie Tang, Ming Dong, Ke Wang, and Wenyuan Li · 2022
Cited alongside, same era.
The effect of trust, perception of risk and security on consumer purchase interest in lazada (empirical study on students of the faculty of economics and business, ibn sina university)
M Arpah, Septa Diana Nabella, et al · 2023
Cited alongside, same era.
Fast attention requires bounded entries
Josh Alman and Zhao Song · 2023
Cited alongside, same era.
Sparks of artificial general intelligence: Early experiments with gpt-4
Sébastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, et al · 2023
Hsr-enhanced sparse attention acceleration
Bo Chen, Yingyu Liang, Zhizhou Sha, Zhenmei Shi, and Zhao Song · 2024
Later among the works it cites.
Zero-th order algorithm for softmax attention optimization
Yichuan Deng, Zhihang Li, Sridhar Mahadevan, and Zhao Song · 2024
Later among the works it cites.
Low rank matrix completion via robust alternating minimization in nearly linear time
Yuzhou Gu, Zhao Song, Junze Yin, and Lichen Zhang · 2024
Later among the works it cites.
On computational limits of modern hopfield models: A fine-grained complexity analysis
Jerry Yao-Chieh Hu, Thomas Lin, Zhao Song, and Han Liu · 2024
Later among the works it cites.
Computational limits of low-rank adaptation (lora) for transformer-based models
Jerry Yao-Chieh Hu, Maojiang Su, En-Jui Kuo, Zhao Song, and Han Liu · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Federated empirical risk minimization via second-order method
Song Bian, Zhao Song, and Junze Yin · 2023
Cited alongside, same era.
Palm: Scaling language modeling with pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al · 2023
Cited alongside, same era.
Query complexity of active learning for function family with nearly orthogonal basis
Xiang Chen, Zhao Song, Baocheng Sun, Junze Yin, and Danyang Zhuo · 2023
Cited alongside, same era.
Attention scheme inspired softmax regression
Yichuan Deng, Zhihang Li, and Zhao Song · 2023
Cited alongside, same era.
Yichuan Deng, Sridhar Mahadevan, and Zhao Song · 2023
Cited alongside, same era.
Faster robust tensor power method for arbitrary order
Yichuan Deng, Zhao Song, and Junze Yin · 2023
Cited alongside, same era.
Llama-adapter v2: Parameter-efficient visual instruction model
Peng Gao, Jiaming Han, Renrui Zhang, Ziyi Lin, Shijie Geng, Aojun Zhou, Wei Zhang, Pan Lu, Conghui He, Xiangyu Yue, et al · 2023
Cited alongside, same era.
On statistical rates and provably efficient criteria of latent diffusion transformers (dits)
Jerry Yao-Chieh Hu, Weimin Wu, Zhao Song, and Han Liu · 2024
Later among the works it cites.
On sparse modern hopfield model
Jerry Yao-Chieh Hu, Donglin Yang, Dennis Wu, Chenwei Xu, Bo-Yu Chen, and Han Liu · 2024
Later among the works it cites.
Faster sampling algorithms for polytopes with small treewidth
Yekun Ke, Xiaoyu Li, Zhao Song, and Tianyi Zhou · 2024
Later among the works it cites.
Theoretical constraints on the expressive power of rope-based tensor attention transformers
Xiaoyu Li, Yingyu Liang, Zhenmei Shi, Zhao Song, and Mingda Wan · 2024
Later among the works it cites.
Yingyu Liang, Heshan Liu, Zhenmei Shi, Zhao Song, and Junze Yin · 2024
Later among the works it cites.
Beyond linear approximations: A novel pruning approach for attention matrix
Yingyu Liang, Jiangxuan Long, Zhenmei Shi, Zhao Song, and Yufa Zhou · 2024
Later among the works it cites.
A tighter complexity analysis of sparsegpt
Xiaoyu Li, Yingyu Liang, Zhenmei Shi, and Zhao Song · 2024
Later among the works it cites.
Fast second-order method for neural networks under small treewidth setting
Xiaoyu Li, Jiangxuan Long, Zhao Song, and Tianyi Zhou · 2024
Later among the works it cites.
Uniform last-iterate guarantee for bandits and reinforcement learning
Junyan Liu, Yunfan Li, Ruosong Wang, and Lin Yang · 2024
Later among the works it cites.
Achieving near-optimal regret for bandit algorithms with uniform last-iterate guarantee
Junyan Liu, Yunfan Li, and Lin Yang · 2024
Later among the works it cites.
Comparison of multiple linear regression and machine learning methods in predicting cognitive function in older chinese type 2 diabetes patients
Chi-Hao Liu, Chung-Hsin Peng, Li-Ying Huang, Fang-Yu Chen, Chun-Heng Kuo, Chung-Ze Wu, and Yu-Fang Cheng · 2024
Later among the works it cites.
Looped relu mlps may be all you need as practical programmable computers
Yingyu Liang, Zhizhou Sha, Zhenmei Shi, Zhao Song, and Yufa Zhou · 2024
Later among the works it cites.
Multi-layer transformers gradient can be approximated in almost linear time
Yingyu Liang, Zhizhou Sha, Zhenmei Shi, Zhao Song, and Yufa Zhou · 2024
Later among the works it cites.
Differential privacy of cross-attention with provable guarantee
Yingyu Liang, Zhenmei Shi, Zhao Song, and Yufa Zhou · 2024
Later among the works it cites.
Tensor attention training: Provably efficient learning of higher-order transformers
Yingyu Liang, Zhenmei Shi, Zhao Song, and Yufa Zhou · 2024
Later among the works it cites.
How to inverting the leverage score distribution?
Zhihang Li, Zhao Song, Weixin Wang, Junze Yin, and Zheng Yu · 2024
Later among the works it cites.
Inverting the leverage score gradient: An efficient approximate newton method
Chenyang Li, Zhao Song, Zhaoxing Xu, and Junze Yin · 2024
Later among the works it cites.
Scaling laws in linear regression: Compute, parameters, and data
Licong Lin, Jingfeng Wu, Sham M Kakade, Peter L Bartlett, and Jason D Lee · 2024
Later among the works it cites.
On the model-misspecification in reinforcement learning
Yunfan Li and Lin Yang · 2024
Later among the works it cites.
Overfitting behaviour of gaussian kernel ridgeless regression: Varying bandwidth or dimensionality
Nicolas Macris, Mehrnaz Vasheghani Farahani, and Florent Krzakala · 2024
Later among the works it cites.
Introducing openai o1-preview
OpenAI · 2024
Later among the works it cites.
Fast dynamic sampling for determinantal point processes
Zhao Song, Junze Yin, Lichen Zhang, and Ruizhe Zhang · 2024
Later among the works it cites.
STanhop: Sparse tandem hopfield model for memory-enhanced time series prediction
Dennis Wu, Jerry Yao-Chieh Hu, Weijian Li, Bo-Yu Chen, and Han Liu · 2024
Later among the works it cites.
Bishop: Bi-directional cellular learning for tabular data with generalized sparse modern hopfield model
Chenwei Xu, Yu-Chao Huang, Jerry Yao-Chieh Hu, Weijian Li, Ammar Gilani, Hsi-Sheng Goan, and Han Liu · 2024
Later among the works it cites.
Yang Cao, Bo Chen, Xiaoyu Li, Yingyu Liang, Zhizhou Sha, Zhenmei Shi, Zhao Song, and Mingda Wan · 2025
Closest in time.
Yuefan Cao, Xiaoyu Li, Yingyu Liang, Zhizhou Sha, Zhenmei Shi, Zhao Song, and Jiahao Zhang · 2025
Closest in time.
Universal approximation of visual autoregressive transformers
Yifang Chen, Xiaoyu Li, Yingyu Liang, Zhenmei Shi, and Zhao Song · 2025
Closest in time.
On computational limits of flowar models: Expressivity and efficiency
Chengyue Gong, Yekun Ke, Xiaoyu Li, Yingyu Liang, Zhizhou Sha, Zhenmei Shi, and Zhao Song · 2025
Closest in time.
Differential privacy mechanisms in neural tangent kernel regression
Jiuxiang Gu, Yingyu Liang, Zhizhou Sha, Zhenmei Shi, and Zhao Song · 2025
Closest in time.
An iterative algorithm for rescaled hyperbolic functions regression
Yeqi Gao, Zhao Song, and Junze Yin · 2025
Closest in time.
Yekun Ke, Xiaoyu Li, Yingyu Liang, Zhizhou Sha, Zhenmei Shi, and Zhao Song · 2025
Closest in time.
Circuit complexity bounds for visual autoregressive model
Yekun Ke, Xiaoyu Li, Yingyu Liang, Zhenmei Shi, and Zhao Song · 2025
Closest in time.
Curse of attention: A kernel-based perspective for why transformers fail to generalize on time series forecasting and beyond
Yekun Ke, Yingyu Liang, Zhenmei Shi, Zhao Song, and Chiwun Yang · 2025
Closest in time.
Fourier circuits in neural networks and transformers: A case study of modular arithmetic with multiple inputs
Chenyang Li, Yingyu Liang, Zhenmei Shi, Zhao Song, and Tianyi Zhou · 2025
Closest in time.
On the computational capability of graph neural networks: A circuit complexity bound perspective
Xiaoyu Li, Yingyu Liang, Zhenmei Shi, Zhao Song, Wei Wang, and Jiahao Zhang · 2025
Closest in time.
When can we solve the weighted low rank approximation problem in truly subquadratic time?
Chenyang Li, Yingyu Liang, Zhenmei Shi, and Zhao Song · 2025
Closest in time.
Lazydit: Lazy learning for the acceleration of diffusion transformers
Xuan Shen, Zhao Song, Yufa Zhou, Bo Chen, Yanyu Li, Yifan Gong, Kai Zhang, Hao Tan, Jason Kuen, Henghui Ding, Zhihao Shu, Wei Niu, Pu Zhao, Yanzhi Wang, and Jiuxiang Gu · 2025
Closest in time.
Numerical pruning for efficient autoregressive models
Xuan Shen, Zhao Song, Yufa Zhou, Bo Chen, Jing Liu, Ruiyi Zhang, Ryan A. Rossi, Hao Tan, Tong Yu, Xiang Chen, Yufan Zhou, Tong Sun, Pu Zhao, Yanzhi Wang, and Jiuxiang Gu · 2025
Closest in time.
Fast and efficient matching algorithm with deadline instances
Zhao Song, Weixin Wang, Chenbo Yin, and Junze Yin · 2025
Closest in time.
Efficient alternating minimization with applications to weighted low rank approximation
Zhao Song, Mingquan Ye, Junze Yin, and Lichen Zhang · 2025
Closest in time.
Statistical guarantees for lifelong reinforcement learning using pac-bayesian theory
Zhi Zhang, Chris Chow, Yasi Zhang, Yanchao Sun, Haochen Zhang, Eric Hanchen Jiang, Han Liu, Furong Huang, Yuchen Cui, and Oscar Hernan Madrid Padilla · 2025
Closest in time.