Fetching the paper…
Reading the bibliography…
The application of transformer-based models on time series forecasting (TSF) tasks has long been popular to study.
Distribution of residual autocorrelations in autoregressive-integrated moving average time series models
George EP Box and David A Pierce · 1970
Earlier work this paper cites.
Exponential smoothing: The state of the art
Everette S Gardner Jr · 1985
Earlier work this paper cites.
Probability inequalities for sums of bounded random variables
Wassily Hoeffding · 1994
Earlier work this paper cites.
Designing a neural network for forecasting financial and economic time series
Iebeling Kaastra and Milton Boyd · 1996
Earlier work this paper cites.
Combining kohonen maps with arima time series models to forecast traffic flow
Mascha Van Der Voort, Mark Dougherty, and Susan Watson · 1996
Earlier work this paper cites.
Long short-term memory
S Hochreiter · 1997
Earlier work this paper cites.
On time series analysis of public health and biomedical data
Scott L Zeger, Rafael Irizarry, and Roger D Peng · 2006
Earlier work this paper cites.
Short-term traffic flow forecasting: An experimental comparison of time-series analysis and supervised learning
Marco Lippi, Matteo Bertini, and Paolo Frasconi · 2013
Earlier work this paper cites.
Dynamic covariance models for multivariate financial time series
Yue Wu, José Miguel Hernández-Lobato, and Ghahramani Zoubin · 2013
Earlier work this paper cites.
On particle methods for parameter estimation in state-space models
Nikolas Kantas, Arnaud Doucet, Sumeetpal S Singh, Jan Maciejowski, and Nicolas Chopin · 2015
Earlier work this paper cites.
Forecasting traffic time series with multivariate predicting method
Yi Yin and Pengjian Shang · 2016
Earlier work this paper cites.
Multi-step ahead time series forecasting for different data patterns based on lstm recurrent neural network
Liu Yunpeng, Hou Di, Bao Junpeng, and Qi Yong · 2017
Earlier work this paper cites.
Time series forecasting for healthcare diagnosis and prognostics with the focus on cardiovascular diseases
C Bui, N Pham, A Vo, A Tran, A Nguyen, and T Le · 2018
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Earlier work this paper cites.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Yuanzhi Li and Yingyu Liang · 2018
Earlier work this paper cites.
On exact computation with an infinitely wide neural net
Sanjeev Arora, Simon S Du, Wei Hu, Zhiyuan Li, Russ R Salakhutdinov, and Ruosong Wang · 2019
Earlier work this paper cites.
Learning and generalization in overparameterized neural networks, going beyond two layers
Zeyuan Allen-Zhu, Yuanzhi Li, and Yingyu Liang · 2019
Earlier work this paper cites.
A convergence theory for deep learning via over-parameterization
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2019
Earlier work this paper cites.
On the convergence rate of training recurrent neural networks
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2019
Earlier work this paper cites.
Enhancing the locality and breaking the memory bottleneck of transformer on time series forecasting
Shiyang Li, Xiaoyong Jin, Yao Xuan, Xiyou Zhou, Wenhu Chen, Yu-Xiang Wang, and Xifeng Yan · 2019
Earlier work this paper cites.
Transformer dissection: a unified understanding of transformer’s attention via the lens of kernel
Yao-Hung Hubert Tsai, Shaojie Bai, Makoto Yamada, Louis-Philippe Morency, and Ruslan Salakhutdinov · 2019
Earlier work this paper cites.
An improved analysis of training over-parameterized deep neural networks
Difan Zou and Quanquan Gu · 2019
Earlier work this paper cites.
Why adam beats sgd for attention models
Jingzhao Zhang, Sai Praneeth Karimireddy, Andreas Veit, Seungyeon Kim, Sashank J Reddi, Sanjiv Kumar, and Suvrit Sra · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom B Brown · 2020
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy · 2020
Earlier work this paper cites.
Ai in healthcare: time-series forecasting using statistical, neural, and ensemble architectures
Shruti Kaushik, Abhinav Choudhury, Pankaj Kumar Sheron, Nataraj Dasgupta, Sayee Natarajan, Larry A Pickett, and Varun Dutt · 2020
Earlier work this paper cites.
Intellicode compose: Code generation using transformer
Alexey Svyatkovskiy, Shao Kun Deng, Shengyu Fu, and Neel Sundaresan · 2020
Earlier work this paper cites.
Financial time series forecasting with deep learning: A systematic literature review: 2005–2019
Omer Berat Sezer, Mehmet Ugur Gudelek, and Ahmet Murat Ozbayoglu · 2020
Earlier work this paper cites.
Quadratic suffices for over-parametrization via matrix chernoff bound, 2020
Zhao Song and Xin Yang · 2020
Earlier work this paper cites.
Towards theoretically understanding why sgd generalizes better than adam in deep learning
Pan Zhou, Jiashi Feng, Chao Ma, Caiming Xiong, Steven Chu Hong Hoi, et al · 2020
Earlier work this paper cites.
Training (overparametrized) neural networks in near-linear time
Jan van den Brand, Binghui Peng, Zhao Song, and Omri Weinstein · 2021
Earlier work this paper cites.
Fl-ntk: A neural tangent kernel-based framework for federated learning analysis
Baihe Huang, Xiaoxiao Li, Zhao Song, and Xin Yang · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2021
Earlier work this paper cites.
Temporal fusion transformers for interpretable multi-horizon time series forecasting
Bryan Lim, Sercan Arik, Nicolas Loeff, and Tomas Pfister · 2021
Earlier work this paper cites.
Research progress in attention mechanism in deep learning
Jian-wei LIU, Jun-wen LIU, and Xiong-lin LUO · 2021
Earlier work this paper cites.
Time-series forecasting with deep learning: a survey
Bryan Lim and Stefan Zohren · 2021
Earlier work this paper cites.
Roformer: Enhanced transformer with rotary position embedding, 2021
Jianlin Su, Yu Lu, Shengfeng Pan, Bo Wen, and Yunfeng Liu · 2021
Earlier work this paper cites.
Does preprocessing help training over-parameterized neural networks?
Zhao Song, Shuo Yang, and Ruizhe Zhang · 2021
Earlier work this paper cites.
Efficient attention: Attention with linear complexities
Zhuoran Shen, Mingyuan Zhang, Haiyu Zhao, Shuai Yi, and Hongsheng Li · 2021
Earlier work this paper cites.
Training multi-layer over-parametrized neural network in subquadratic time
Zhao Song, Lichen Zhang, and Ruizhe Zhang · 2021
Earlier work this paper cites.
Training data-efficient image transformers & distillation through attention
Hugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa, Alexandre Sablayrolles, and Hervé Jégou · 2021
Earlier work this paper cites.
Autoformer: Decomposition transformers with auto-correlation for long-term series forecasting
Haixu Wu, Jiehui Xu, Jianmin Wang, and Mingsheng Long · 2021
Earlier work this paper cites.
Informer: Beyond efficient transformer for long sequence time-series forecasting
Haoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang, Jianxin Li, Hui Xiong, and Wancai Zhang · 2021
Earlier work this paper cites.
Feature purification: How adversarial training performs robust deep learning
Zeyuan Allen-Zhu and Yuanzhi Li · 2022
Earlier work this paper cites.
Training overparametrized neural networks in sublinear time
Yichuan Deng, Hang Hu, Zhao Song, Omri Weinstein, and Danyang Zhuo · 2022
Earlier work this paper cites.
How to train your hippo: State space models with generalized orthogonal basis projections, 2022
Albert Gu, Isys Johnson, Aman Timalsina, Atri Rudra, and Christopher Ré · 2022
Earlier work this paper cites.
Training overparametrized neural networks in sublinear time
Hang Hu, Zhao Song, Omri Weinstein, and Danyang Zhuo · 2022
Earlier work this paper cites.
Image creation based on transformer and generative adversarial networks
Hangyu Liu and Qicheng Liu · 2022
Earlier work this paper cites.
Generative time series forecasting with diffusion, denoise, and disentanglement
Yan Li, Xinjiang Lu, Yaqing Wang, and Dejing Dou · 2022
Earlier work this paper cites.
Pyraformer: Low-complexity pyramidal attention for long-range time series modeling and forecasting
Shizhan Liu, Hang Yu, Cong Liao, Jianguo Li, Weiyao Lin, Alex X Liu, and Schahram Dustdar · 2022
Earlier work this paper cites.
Bounding the width of neural networks via coupled initialization a worst case analysis
Alexander Munteanu, Simon Omlor, Zhao Song, and David Woodruff · 2022
Cited alongside, same era.
A time series is worth 64 words: Long-term forecasting with transformers
Yuqi Nie, Nam H Nguyen, Phanwadee Sinthong, and Jayant Kalagnanam · 2022
Cited alongside, same era.
Zhenmei Shi, Junyi Wei, and Yingyu Liang · 2022
Cited alongside, same era.
Emergent abilities of large language models
Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, et al · 2022
Cited alongside, same era.
Fedformer: Frequency enhanced decomposed transformer for long-term series forecasting
Differential privacy mechanisms in neural tangent kernel regression
Jiuxiang Gu, Yingyu Liang, Zhizhou Sha, Zhenmei Shi, and Zhao Song · 2024
Closest in time.
Outlier-efficient hopfield layers for large transformer-based models
Jerry Yao-Chieh Hu, Pei-Hsuan Chang, Haozheng Luo, Hong-Yu Chen, Weijian Li, Wei-Po Wang, and Han Liu · 2024
Closest in time.
On computational limits of modern hopfield models: A fine-grained complexity analysis
Jerry Yao-Chieh Hu, Thomas Lin, Zhao Song, and Han Liu · 2024
Closest in time.
Neural network-based score estimation in diffusion models: Optimization and generalization
Yinbin Han, Meisam Razaviyayn, and Renyuan Xu · 2024
Closest in time.
Computational limits of low-rank adaptation (lora) for transformer-based models
Jerry Yao-Chieh Hu, Maojiang Su, En-Jui Kuo, Zhao Song, and Han Liu · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tian Zhou, Ziqing Ma, Qingsong Wen, Xue Wang, Liang Sun, and Rong Jin · 2022
Cited alongside, same era.
Llm-based interaction for content generation: A case study on the perception of employees in an it department
Alexandre Agossah, Frédérique Krupa, Matthieu Perreira Da Silva, and Patrick Le Callet · 2023
Cited alongside, same era.
Bypass exponential time preprocessing: Fast neural network training via weight-data correlation preprocessing
Josh Alman, Jiehao Liang, Zhao Song, Ruizhe Zhang, and Danyang Zhuo · 2023
Cited alongside, same era.
Llm based generation of item-description for recommendation system
Arkadeep Acharya, Brijraj Singh, and Naoyuki Onoe · 2023
Cited alongside, same era.
Physics of language models: Part 1, context-free grammar
Zeyuan Allen-Zhu and Yuanzhi Li · 2023
Cited alongside, same era.
Hierarchical attention network for multivariate time series long-term forecasting
Hongjing Bi, Lilei Lu, and Yizhen Meng · 2023
Cited alongside, same era.
Algorithm and hardness for dynamic attention maintenance in large language models
Jan van den Brand, Zhao Song, and Tianyi Zhou · 2023
Cited alongside, same era.
Tsmixer: An all-mlp architecture for time series forecasting
Si-An Chen, Chun-Liang Li, Nate Yoder, Sercan O Arik, and Tomas Pfister · 2023
Cited alongside, same era.
Closest in time.
Fundamental limits of prompt tuning transformers: Universality, capacity and efficiency
Jerry Yao-Chieh Hu, Wei-Po Wang, Ammar Gilani, Chenyang Li, Zhao Song, and Han Liu · 2024
Closest in time.
Provably optimal memory capacity for modern hopfield models: Transformer-compatible dense associative memories as spherical codes
Jerry Yao-Chieh Hu, Dennis Wu, and Han Liu · 2024
Closest in time.
Jerry Yao-Chieh Hu, Weimin Wu, Yi-Chen Lee, Yu-Chao Huang, Minshuo Chen, and Han Liu · 2024
Closest in time.
On statistical rates and provably efficient criteria of latent diffusion transformers (dits)
Jerry Yao-Chieh Hu, Weimin Wu, Zhuoru Li, Sophia Pi, , Zhao Song, and Han Liu · 2024
Closest in time.
Bliva: A simple multimodal llm for better handling of text-rich visual questions
Wenbo Hu, Yifan Xu, Yi Li, Weiyue Li, Zeyuan Chen, and Zhuowen Tu · 2024
Closest in time.
Bliva: A simple multimodal llm for better handling of text-rich visual questions
Wenbo Hu, Yifan Xu, Yi Li, Weiyue Li, Zeyuan Chen, and Zhuowen Tu · 2024
Closest in time.
On the expressive power of modern hopfield networks
Xiaoyu Li, Yuanpeng Li, Yingyu Liang, Zhenmei Shi, and Zhao Song · 2024
Closest in time.
Theoretical constraints on the expressive power of rope-based tensor attention transformers
Xiaoyu Li, Yingyu Liang, Zhenmei Shi, Zhao Song, and Mingda Wan · 2024
Closest in time.
Fine-grained attention i/o complexity: Comprehensive analysis for backward passes
Xiaoyu Li, Yingyu Liang, Zhenmei Shi, Zhao Song, and Yufa Zhou · 2024
Closest in time.
Yingyu Liang, Heshan Liu, Zhenmei Shi, Zhao Song, Zhuoyan Xu, and Junze Yin · 2024
Closest in time.
Chenyang Li, Yingyu Liang, Zhenmei Shi, and Zhao Song · 2024
Closest in time.
A tighter complexity analysis of sparsegpt
Xiaoyu Li, Yingyu Liang, Zhenmei Shi, and Zhao Song · 2024
Closest in time.
Multi-layer transformers gradient can be approximated in almost linear time
Yingyu Liang, Zhizhou Sha, Zhenmei Shi, Zhao Song, and Yufa Zhou · 2024
Closest in time.
Towards infinite-long prefix in transformer
Yingyu Liang, Zhenmei Shi, Zhao Song, and Chiwun Yang · 2024
Closest in time.
Tensor attention training: Provably efficient learning of higher-order transformers
Yingyu Liang, Zhenmei Shi, Zhao Song, and Yufa Zhou · 2024
Closest in time.
U-mamba: Enhancing long-range dependency for biomedical image segmentation
Jun Ma, Feifei Li, and Bo Wang · 2024
Closest in time.
Gpt-4 technical report, 2024
OpenAI, Josh Achiam, Steven Adler, and Sandhini Agarwal et al · 2024
Closest in time.
Vm-unet: Vision mamba unet for medical image segmentation
Jiacheng Ruan and Suncheng Xiang · 2024
Closest in time.
Learning to (learn at test time): Rnns with expressive hidden states
Yu Sun, Xinhao Li, Karan Dalal, Jiarui Xu, Arjun Vikram, Genghan Zhang, Yann Dubois, Xinlei Chen, Xiaolong Wang, Sanmi Koyejo, et al · 2024
Closest in time.
Zhenmei Shi, Yifei Ming, Xuan-Phi Nguyen, Yingyu Liang, and Shafiq Joty · 2024
Closest in time.
A graph-theoretic framework for understanding open-world semi-supervised learning
Yiyou Sun, Zhenmei Shi, and Yixuan Li · 2024
Closest in time.
Solving attention kernel regression problem via pre-conditioner
Zhao Song, Junze Yin, and Lichen Zhang · 2024
Closest in time.
Training multi-layer over-parametrized neural network in subquadratic time
Zhao Song, Lichen Zhang, and Ruizhe Zhang · 2024
Closest in time.
Uniform memory retrieval with larger capacity for modern hopfield models
Dennis Wu, Jerry Yao-Chieh Hu, Teng-Yun Hsiao, and Han Liu · 2024
Closest in time.
STanhop: Sparse tandem hopfield model for memory-enhanced time series prediction
Dennis Wu, Jerry Yao-Chieh Hu, Weijian Li, Bo-Yu Chen, and Han Liu · 2024
Closest in time.
Transformers are deep optimizers: Provable in-context learning for deep model training
Weimin Wu, Maojiang Su, Jerry Yao-Chieh Hu, Zhao Song, and Han Liu · 2024
Closest in time.
Do large language models have compositional ability? an investigation into limitations and scalability
Zhuoyan Xu, Zhenmei Shi, and Yingyu Liang · 2024
Closest in time.
Segmamba: Long-range sequential modeling mamba for 3d medical image segmentation
Zhaohu Xing, Tian Ye, Yijun Yang, Guang Liu, and Lei Zhu · 2024
Closest in time.
Tianzhu Ye, Li Dong, Yuqing Xia, Yutao Sun, Yi Zhu, Gao Huang, and Furu Wei · 2024
Closest in time.
Frequency-domain mlps are more effective learners in time series forecasting
Kun Yi, Qi Zhang, Wei Fan, Shoujin Wang, Pengyang Wang, Hui He, Ning An, Defu Lian, Longbing Cao, and Zhendong Niu · 2024
Closest in time.
Vision mamba: Efficient visual representation learning with bidirectional state space model
Lianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang, Wenyu Liu, and Xinggang Wang · 2024
Closest in time.
Convolution meets lora: Parameter efficient finetuning for segment anything model
Zihan Zhong, Zhiqiang Tang, Tong He, Haoyang Fang, and Chun Yuan · 2024
Closest in time.
Bypassing the exponential dependency: Looped transformers efficiently learn in-context by multi-step gradient descent
Bo Chen, Xiaoyu Li, Yingyu Liang, Zhenmei Shi, and Zhao Song · 2025
Closest in time.
The computational limits of state-space models and mamba via the lens of circuit complexity
Yifang Chen, Xiaoyu Li, Yingyu Liang, Zhenmei Shi, and Zhao Song · 2025
Closest in time.
Universal approximation of visual autoregressive transformers
Yifang Chen, Xiaoyu Li, Yingyu Liang, Zhenmei Shi, and Zhao Song · 2025
Closest in time.
Hsr-enhanced sparse attention acceleration
Bo Chen, Yingyu Liang, Zhizhou Sha, Zhenmei Shi, and Zhao Song · 2025
Closest in time.
On computational limits of flowar models: Expressivity and efficiency
Chengyue Gong, Yekun Ke, Xiaoyu Li, Yingyu Liang, Zhizhou Sha, Zhenmei Shi, and Zhao Song · 2025
Closest in time.
Yekun Ke, Xiaoyu Li, Yingyu Liang, Zhizhou Sha, Zhenmei Shi, and Zhao Song · 2025
Closest in time.
Circuit complexity bounds for visual autoregressive model
Yekun Ke, Xiaoyu Li, Yingyu Liang, Zhenmei Shi, and Zhao Song · 2025
Closest in time.
Neural algorithmic reasoning for hypergraphs with looped transformers
Xiaoyu Li, Yingyu Liang, Jiangxuan Long, Zhenmei Shi, Zhao Song, and Zhen Zhuang · 2025
Closest in time.
Fourier circuits in neural networks and transformers: A case study of modular arithmetic with multiple inputs
Chenyang Li, Yingyu Liang, Zhenmei Shi, Zhao Song, and Tianyi Zhou · 2025
Closest in time.
On the computational capability of graph neural networks: A circuit complexity bound perspective
Xiaoyu Li, Yingyu Liang, Zhenmei Shi, Zhao Song, Wei Wang, and Jiahao Zhang · 2025
Closest in time.
Beyond linear approximations: A novel pruning approach for attention matrix
Yingyu Liang, Jiangxuan Long, Zhenmei Shi, Zhao Song, and Yufa Zhou · 2025
Closest in time.
Looped relu mlps may be all you need as practical programmable computers
Yingyu Liang, Zhizhou Sha, Zhenmei Shi, Zhao Song, and Yufa Zhou · 2025
Closest in time.