Fetching the paper…
Reading the bibliography…
We investigate the fundamental limits of transformer-based foundation models, extending our analysis to include Visual Autoregressive (VAR) transformers.
Conditional image generation with pixelcnn decoders
Aaron Van den Oord, Nal Kalchbrenner, Lasse Espeholt, Oriol Vinyals, Alex Graves, et al · 2016
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, 𝖫𝗈𝗌𝗌 \mathsf{Loss} ukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Pixelsnail: An improved autoregressive generative model
Xi Chen, Nikhil Mishra, Mostafa Rohaninejad, and Pieter Abbeel · 2018
Earlier work this paper cites.
Generating diverse high-fidelity images with vq-vae-2
Ali Razavi, Aaron Van den Oord, and Oriol Vinyals · 2019
Earlier work this paper cites.
Generative modeling by estimating gradients of the data distribution
Yang Song and Stefano Ermon · 2019
Earlier work this paper cites.
Denoising diffusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel · 2020
Earlier work this paper cites.
Instahide’s sample complexity when mixing two private images
Baihe Huang, Zhao Song, Runzhou Tao, Junze Yin, Ruizhe Zhang, and Danyang Zhuo · 2020
Earlier work this paper cites.
Denoising diffusion implicit models
Jiaming Song, Chenlin Meng, and Stefano Ermon · 2020
Earlier work this paper cites.
Are transformers universal approximators of sequence-to-sequence functions?
Chulhee Yun, Srinadh Bhojanapalli, Ankit Singh Rawat, Sashank Reddi, and Sanjiv Kumar · 2020
Earlier work this paper cites.
Cogview: Mastering text-to-image generation via transformers
Ming Ding, Zhuoyi Yang, Wenyi Hong, Wendi Zheng, Chang Zhou, Da Yin, Junyang Lin, Xu Zou, Zhou Shao, Hongxia Yang, et al · 2021
Earlier work this paper cites.
Taming transformers for high-resolution image synthesis
Patrick Esser, Robin Rombach, and Bjorn Ommer · 2021
Earlier work this paper cites.
Optimal-degree polynomial approximations for exponentials and gaussian kernel density estimation
Amol Aggarwal and Josh Alman · 2022
Earlier work this paper cites.
A nearly optimal size coreset algorithm with nearly linear time
Yichuan Deng, Zhao Song, Yitan Wang, and Yuanyuan Yang · 2022
Earlier work this paper cites.
Cogview2: Faster and better text-to-image generation via hierarchical transformers
Ming Ding, Wendi Zheng, Wenyi Hong, and Jie Tang · 2022
Earlier work this paper cites.
Cascaded diffusion models for high fidelity image generation
Jonathan Ho, Chitwan Saharia, William Chan, David J Fleet, Mohammad Norouzi, and Tim Salimans · 2022
Earlier work this paper cites.
Sublinear time algorithm for online weighted bipartite matching
Hang Hu, Zhao Song, Runzhou Tao, Zhaozhuo Xu, Junze Yin, and Danyang Zhuo · 2022
Earlier work this paper cites.
A dynamic fast gaussian transform
Baihe Huang, Zhao Song, Omri Weinstein, Junze Yin, Hengjie Zhang, and Ruizhe Zhang · 2022
Earlier work this paper cites.
Provable memorization capacity of transformers
Junghwan Kim, Michelle Kim, and Barzan Mozafari · 2022
Earlier work this paper cites.
Autoregressive image generation using residual quantization
Doyup Lee, Chiheon Kim, Saehoon Kim, Minsu Cho, and Wook-Shin Han · 2022
Earlier work this paper cites.
A faster k k -means++ algorithm
Jiehao Liang, Somdeb Sarkhel, Zhao Song, Chenbo Yin, Junze Yin, and Danyang Zhuo · 2022
Earlier work this paper cites.
Dynamic maintenance of kernel density estimation data structure: From practice to theory
Jiehao Liang, Zhao Song, Zhaozhuo Xu, Junze Yin, and Danyang Zhuo · 2022
Earlier work this paper cites.
A fast ode solver for diffusion probabilistic model sampling in around 10 steps
C Lu, Y Zhou, F Bao, J Chen, and C Li · 2022
Earlier work this paper cites.
Dreamfusion: Text-to-3d using 2d diffusion
Ben Poole, Ajay Jain, Jonathan T Barron, and Ben Mildenhall · 2022
Earlier work this paper cites.
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer · 2022
Earlier work this paper cites.
A multi-layer extreme learning machine refined by sparrow search algorithm and weighted mean filter for short-term multi-step wind speed forecasting
Haochen Zhang, Zhiyun Peng, Junjie Tang, Ming Dong, Ke Wang, and Wenyuan Li · 2022
Earlier work this paper cites.
Sumformer: Universal approximation for efficient transformers
Silas Alberti, Niclas Dern, Laura Thesing, and Gitta Kutyniok · 2023
Earlier work this paper cites.
Fast attention requires bounded entries
Josh Alman and Zhao Song · 2023
Earlier work this paper cites.
All are worth words: A vit backbone for diffusion models
Fan Bao, Shen Nie, Kaiwen Xue, Yue Cao, Chongxuan Li, Hang Su, and Jun Zhu · 2023
Earlier work this paper cites.
Federated empirical risk minimization via second-order method
Song Bian, Zhao Song, and Junze Yin · 2023
Earlier work this paper cites.
Query complexity of active learning for function family with nearly orthogonal basis
Xiang Chen, Zhao Song, Baocheng Sun, Junze Yin, and Danyang Zhuo · 2023
Earlier work this paper cites.
Yichuan Deng, Sridhar Mahadevan, and Zhao Song · 2023
Earlier work this paper cites.
Faster robust tensor power method for arbitrary order
Yichuan Deng, Zhao Song, and Junze Yin · 2023
Earlier work this paper cites.
An over-parameterized exponential regression
Yeqi Gao, Sridhar Mahadevan, and Zhao Song · 2023
Earlier work this paper cites.
Yeqi Gao, Zhao Song, Weixin Wang, and Junze Yin · 2023
Earlier work this paper cites.
Yeqi Gao, Zhao Song, and Shenghao Xie · 2023
Earlier work this paper cites.
Gradientcoin: A peer-to-peer decentralized large language models
Yeqi Gao, Zhao Song, and Junze Yin · 2023
Earlier work this paper cites.
An iterative algorithm for rescaled hyperbolic functions regression
Yeqi Gao, Zhao Song, and Junze Yin · 2023
Earlier work this paper cites.
Approximation theory of transformer networks for sequence modeling
Haotian Jiang and Qianxiao Li · 2023
Earlier work this paper cites.
Local convergence of approximate newton method for two layer nonlinear regression
Zhihang Li, Zhao Song, Zifan Wang, and Junze Yin · 2023
Cited alongside, same era.
Solving regularized exp, cosh and sinh regression problems
Zhihang Li, Zhao Song, and Tianyi Zhou · 2023
Cited alongside, same era.
Low-switching policy gradient with exploration via online sensitivity sampling
Yunfan Li, Yiran Wang, Yu Cheng, and Lin Yang · 2023
Cited alongside, same era.
Scalable diffusion models with transformers
William Peebles and Saining Xie · 2023
Cited alongside, same era.
Ritwik Sinha, Zhao Song, and Tianyi Zhou · 2023
Cited alongside, same era.
A tighter complexity analysis of sparsegpt
Xiaoyu Li, Yingyu Liang, Zhenmei Shi, and Zhao Song · 2024
Later among the works it cites.
Fast second-order method for neural networks under small treewidth setting
Xiaoyu Li, Jiangxuan Long, Zhao Song, and Tianyi Zhou · 2024
Later among the works it cites.
Uniform last-iterate guarantee for bandits and reinforcement learning
Junyan Liu, Yunfan Li, Ruosong Wang, and Lin Yang · 2024
Later among the works it cites.
Achieving near-optimal regret for bandit algorithms with uniform last-iterate guarantee
Junyan Liu, Yunfan Li, and Lin Yang · 2024
Later among the works it cites.
Looped relu mlps may be all you need as practical programmable computers
Yingyu Liang, Zhizhou Sha, Zhenmei Shi, Zhao Song, and Yufa Zhou · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A unified scheme of resnet and softmax
Zhao Song, Weixin Wang, and Junze Yin · 2023
Cited alongside, same era.
Fast and efficient matching algorithm with deadline instances
Zhao Song, Weixin Wang, Chenbo Yin, and Junze Yin · 2023
Cited alongside, same era.
The expressibility of polynomial based attention scheme
Zhao Song, Guangyi Xu, and Junze Yin · 2023
Cited alongside, same era.
An automatic learning rate schedule algorithm for achieving faster convergence and steeper descent
Zhao Song and Chiwun Yang · 2023
Cited alongside, same era.
A nearly-optimal bound for fast regression with ℓ ∞ \ell_{\infty} guarantee
Zhao Song, Mingquan Ye, Junze Yin, and Lichen Zhang · 2023
Cited alongside, same era.
Zhao Song, Junze Yin, and Ruizhe Zhang · 2023
Cited alongside, same era.
Dolfin: Diffusion layout transformers without autoencoder
Yilin Wang, Zeyuan Chen, Liangjun Zhong, Zheng Ding, Zhizhou Sha, and Zhuowen Tu · 2023
Cited alongside, same era.
Later among the works it cites.
Multi-layer transformers gradient can be approximated in almost linear time
Yingyu Liang, Zhizhou Sha, Zhenmei Shi, Zhao Song, and Yufa Zhou · 2024
Later among the works it cites.
Differential privacy mechanisms in neural tangent kernel regression
Yingyu Liang, Zhizhou Sha, Zhenmei Shi, and Zhao Song · 2024
Later among the works it cites.
Differential privacy of cross-attention with provable guarantee
Yingyu Liang, Zhenmei Shi, Zhao Song, and Yufa Zhou · 2024
Later among the works it cites.
Tensor attention training: Provably efficient learning of higher-order transformers
Yingyu Liang, Zhenmei Shi, Zhao Song, and Yufa Zhou · 2024
Later among the works it cites.
How to inverting the leverage score distribution?
Zhihang Li, Zhao Song, Weixin Wang, Junze Yin, and Zheng Yu · 2024
Later among the works it cites.
Inverting the leverage score gradient: An efficient approximate newton method
Chenyang Li, Zhao Song, Zhaoxing Xu, and Junze Yin · 2024
Later among the works it cites.
AI @ Meta Llama Team · 2024
Later among the works it cites.
On the model-misspecification in reinforcement learning
Yunfan Li and Lin Yang · 2024
Later among the works it cites.
Score-based generative diffusion models for social recommendations
Chengyi Liu, Jiahao Zhang, Shijie Wang, Wenqi Fan, and Qing Li · 2024
Later among the works it cites.
Introducing openai o1-preview
OpenAI · 2024
Later among the works it cites.
Flowar: Scale-wise autoregressive image generation meets flow matching
Sucheng Ren, Qihang Yu, Ju He, Xiaohui Shen, Alan Yuille, and Liang-Chieh Chen · 2024
Later among the works it cites.
Solving attention kernel regression problem via pre-conditioner
Zhao Song, Junze Yin, and Lichen Zhang · 2024
Later among the works it cites.
Fast dynamic sampling for determinantal point processes
Zhao Song, Junze Yin, Lichen Zhang, and Ruizhe Zhang · 2024
Later among the works it cites.
Visual autoregressive modeling: Scalable image generation via next-scale prediction
Keyu Tian, Yi Jiang, Zehuan Yuan, Bingyue Peng, and Liwei Wang · 2024
Later among the works it cites.
STanhop: Sparse tandem hopfield model for memory-enhanced time series prediction
Dennis Wu, Jerry Yao-Chieh Hu, Weijian Li, Bo-Yu Chen, and Han Liu · 2024
Later among the works it cites.
Prolificdreamer: High-fidelity and diverse text-to-3d generation with variational score distillation
Zhengyi Wang, Cheng Lu, Yikai Wang, Fan Bao, Chongxuan Li, Hang Su, and Jun Zhu · 2024
Later among the works it cites.
Tokencompose: Text-to-image diffusion with token-level supervision
Zirui Wang, Zhizhou Sha, Zheng Ding, Yilin Wang, and Zhuowen Tu · 2024
Later among the works it cites.
Omnicontrolnet: Dual-stage integration for conditional image generation
Yilin Wang, Haiyang Xu, Xiang Zhang, Zeyuan Chen, Zhizhou Sha, Zirui Wang, and Zhuowen Tu · 2024
Later among the works it cites.
Bishop: Bi-directional cellular learning for tabular data with generalized sparse modern hopfield model
Chenwei Xu, Yu-Chao Huang, Jerry Yao-Chieh Hu, Weijian Li, Ammar Gilani, Hsi-Sheng Goan, and Han Liu · 2024
Later among the works it cites.
Bayesian diffusion models for 3d shape reconstruction
Haiyang Xu, Yu Lei, Zeyuan Chen, Xiang Zhang, Yue Zhao, Yilin Wang, and Zhuowen Tu · 2024
Later among the works it cites.
Raphael: Text-to-image generation via large mixture of diffusion paths
Zeyue Xue, Guanglu Song, Qiushan Guo, Boxiao Liu, Zhuofan Zong, Yu Liu, and Ping Luo · 2024
Later among the works it cites.
Richspace: Enriching text-to-video prompt space via text embedding interpolation
Yuefan Cao, Chengyue Gong, Xiaoyu Li, Yingyu Liang, Zhizhou Sha, Zhenmei Shi, and Zhao Song · 2025
Closest in time.
High-order matching for one-step shortcut diffusion models
Bo Chen, Chengyue Gong, Xiaoyu Li, Yingyu Liang, Zhizhou Sha, Zhenmei Shi, Zhao Song, and Mingda Wan · 2025
Closest in time.
Yuefan Cao, Xiaoyu Li, Yingyu Liang, Zhizhou Sha, Zhenmei Shi, Zhao Song, and Jiahao Zhang · 2025
Closest in time.
Yekun Ke, Xiaoyu Li, Yingyu Liang, Zhizhou Sha, Zhenmei Shi, and Zhao Song · 2025
Closest in time.
Circuit complexity bounds for visual autoregressive model
Yekun Ke, Xiaoyu Li, Yingyu Liang, Zhenmei Shi, and Zhao Song · 2025
Closest in time.
Neural algorithmic reasoning for hypergraphs with looped transformers
Xiaoyu Li, Yingyu Liang, Jiangxuan Long, Zhenmei Shi, Zhao Song, and Zhen Zhuang · 2025
Closest in time.
Fourier circuits in neural networks and transformers: A case study of modular arithmetic with multiple inputs
Chenyang Li, Yingyu Liang, Zhenmei Shi, Zhao Song, and Tianyi Zhou · 2025
Closest in time.
On the computational capability of graph neural networks: A circuit complexity bound perspective
Xiaoyu Li, Yingyu Liang, Zhenmei Shi, Zhao Song, Wei Wang, and Jiahao Zhang · 2025
Closest in time.
Lazydit: Lazy learning for the acceleration of diffusion transformers
Xuan Shen, Zhao Song, Yufa Zhou, Bo Chen, Yanyu Li, Yifan Gong, Kai Zhang, Hao Tan, Jason Kuen, Henghui Ding, Zhihao Shu, Wei Niu, Pu Zhao, Yanzhi Wang, and Jiuxiang Gu · 2025
Closest in time.
Numerical pruning for efficient autoregressive models
Xuan Shen, Zhao Song, Yufa Zhou, Bo Chen, Jing Liu, Ruiyi Zhang, Ryan A. Rossi, Hao Tan, Tong Yu, Xiang Chen, Yufan Zhou, Tong Sun, Pu Zhao, Yanzhi Wang, and Jiuxiang Gu · 2025
Closest in time.
Efficient alternating minimization with applications to weighted low rank approximation
Zhao Song, Mingquan Ye, Junze Yin, and Lichen Zhang · 2025
Closest in time.
Statistical guarantees for lifelong reinforcement learning using pac-bayesian theory
Zhi Zhang, Chris Chow, Yasi Zhang, Yanchao Sun, Haochen Zhang, Eric Hanchen Jiang, Han Liu, Furong Huang, Yuchen Cui, and Oscar Hernan Madrid Padilla · 2025
Closest in time.