Fetching the paper…
Reading the bibliography…
Transformers have been actively studied for time-series forecasting in recent years.
Forecasting sales by exponentially weighted moving averages
Peter R Winters · 1960
Earlier work this paper cites.
Decomposition of seasonal time series: a model for the census x-11 program
William P Cleveland and George C Tiao · 1976
Earlier work this paper cites.
Stl: A seasonal-trend decomposition
Robert B Cleveland, William S Cleveland, Jean E McRae, and Irma Terpenning · 1990
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
The theta model: a decomposition approach to forecasting
Vassilis Assimakopoulos and Konstantinos Nikolopoulos · 2000
Earlier work this paper cites.
Forecasting seasonals and trends by exponentially weighted moving averages
Charles C Holt · 2004
Earlier work this paper cites.
On periodicity detection and structural periodic similarity
Michail Vlachos, Philip Yu, and Vittorio Castelli · 2005
Earlier work this paper cites.
Forecasting with exponential smoothing: the state space approach
Rob Hyndman, Anne B Koehler, J Keith Ord, and Ralph D Snyder · 2008
Earlier work this paper cites.
Automatic time series forecasting: the forecast package for r
Rob J Hyndman and Yeasmin Khandakar · 2008
Earlier work this paper cites.
Damped trend exponential smoothing: a modelling viewpoint
Eddie McKenzie and Everette S Gardner Jr · 2010
Earlier work this paper cites.
Forecasting time series with complex seasonal patterns using exponential smoothing
Alysha M De Livera, Rob J Hyndman, and Ralph D Snyder · 2011
Earlier work this paper cites.
Forecasting monthly and quarterly time series using stl decomposition
Marina Theodosiou · 2011
Earlier work this paper cites.
Fast training of convolutional networks through ffts
Michaël Mathieu, Mikael Henaff, and Yann LeCun · 2014
Earlier work this paper cites.
Dropout: A simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Cited alongside, same era.
Learning theory and algorithms for forecasting non-stationary time series
Vitaly Kuznetsov and Mehryar Mohri · 2015
Cited alongside, same era.
Layer normalization, 2016
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Complex exponential smoothing
Ivan Svetunkov · 2016
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
End-to-end object detection with transformers
Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko · 2020
Later among the works it cites.
Reformer: The efficient transformer
Nikita Kitaev, Lukasz Kaiser, and Anselm Levskaya · 2020
Later among the works it cites.
Fixed encoder self-attention patterns in transformer-based machine translation
Alessandro Raganato, Yves Scherrer, and Jörg Tiedemann · 2020
Later among the works it cites.
Adversarial sparse transformer for time series forecasting
Sifan Wu, Xi Xiao, Qianggang Ding, Peilin Zhao, Ying Wei, and Junzhou Huang · 2020
Later among the works it cites.
Hard-coded gaussian attention for neural machine translation
Weiqiu You, Simeng Sun, and Mohit Iyyer · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Modeling long-and short-term temporal patterns with deep neural networks
Guokun Lai, Wei-Cheng Chang, Yiming Yang, and Hanxiao Liu · 2018
Cited alongside, same era.
Forecasting at scale
Sean J Taylor and Benjamin Letham · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
Enhancing the locality and breaking the memory bottleneck of transformer on time series forecasting
Shiyang Li, Xiaoyong Jin, Yao Xuan, Xiyou Zhou, Wenhu Chen, Yu-Xiang Wang, and Xifeng Yan · 2019
Cited alongside, same era.
N-beats: Neural basis expansion analysis for interpretable time series forecasting
Boris N Oreshkin, Dmitri Carpov, Nicolas Chapados, and Yoshua Bengio · 2019
Cited alongside, same era.
Deepar: Probabilistic forecasting with autoregressive recurrent networks
David Salinas, Valentin Flunkert, Jan Gasthaus, and Tim Januschowski · 2019
Cited alongside, same era.
Aadyot Bhatnagar, Paul Kassianik, Chenghao Liu, Tian Lan, Wenzhuo Yang, Rowan Cassius, Doyen Sahoo, Devansh Arpit, Sri Subramanian, Gerald Woo, et al · 2021
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Later among the works it cites.
Fnet: Mixing tokens with fourier transforms
James Lee-Thorp, Joshua Ainslie, Ilya Eckstein, and Santiago Ontanon · 2021
Later among the works it cites.
Synthesizer: Rethinking self-attention for transformer models
Yi Tay, Dara Bahri, Donald Metzler, Da-Cheng Juan, Zhe Zhao, and Che Zheng · 2021
Later among the works it cites.
Autoformer: Decomposition transformers with Auto-Correlation for long-term series forecasting
Haixu Wu, Jiehui Xu, Jianmin Wang, and Mingsheng Long · 2021
Later among the works it cites.
A transformer-based framework for multivariate time series representation learning
George Zerveas, Srideepika Jayaraman, Dhaval Patel, Anuradha Bhamidipaty, and Carsten Eickhoff · 2021
Later among the works it cites.
Informer: Beyond efficient transformer for long sequence time-series forecasting
Haoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang, Jianxin Li, Hui Xiong, and Wancai Zhang · 2021
Later among the works it cites.
CoST: Contrastive learning of disentangled seasonal-trend representations for time series forecasting
Gerald Woo, Chenghao Liu, Doyen Sahoo, Akshat Kumar, and Steven Hoi · 2022
Closest in time.