Fetching the paper…
Reading the bibliography…
Although Transformer-based methods have significantly improved state-of-the-art results for long-term series forecasting, they are not only computationally expensive but more importantly, are unable to capture the global view of time series (e.g.
Some recent advances in forecasting and control
Box, G. E. P. and Jenkins, G. M · 1968
Earlier work this paper cites.
Distribution of residual autocorrelations in autoregressive-integrated moving average time series models
Box, G. E. P. and Pierce, D. A · 1970
Earlier work this paper cites.
Extensions of lipschitz mappings into hilbert space
Johnson, W. B · 1984
Earlier work this paper cites.
Stl: A seasonal-trend decomposition
Cleveland, R. B., Cleveland, W. S., McRae, J. E., and Terpenning, I · 1990
Earlier work this paper cites.
Long Short-Term Memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
An algorithmic theory of learning: Robust concepts and random projection
Arriaga, R. I. and Vempala, S. S · 2006
Earlier work this paper cites.
Relative-error CUR matrix decompositions
Drineas, P., Mahoney, M. W., and Muthukrishnan, S · 2007
Earlier work this paper cites.
Random features for large-scale kernel machines
Rahimi, A. and Recht, B · 2008
Earlier work this paper cites.
On the difficulty of training recurrent neural networks
Pascanu, R., Mikolov, T., and Bengio, Y · 2013
Earlier work this paper cites.
Empirical evaluation of gated recurrent neural networks on sequence modeling
Chung, J., Gülçehre, Ç., Cho, K., and Bengio, Y · 2014
Earlier work this paper cites.
Fast training of convolutional networks through ffts
Mathieu, M., Henaff, M., and LeCun, Y · 2014
Earlier work this paper cites.
Deepar: Probabilistic forecasting with autoregressive recurrent networks
Flunkert, V., Salinas, D., and Gasthaus, J · 2017
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization
Kingma, D. P. and Ba, J · 2017
Earlier work this paper cites.
A dual-stage attention-based recurrent neural network for time series prediction
Qin, Y., Song, D., Chen, H., Cheng, W., Jiang, G., and Cottrell, G. W · 2017
Earlier work this paper cites.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Cited alongside, same era.
Modeling long-and short-term temporal patterns with deep neural networks
Lai, G., Chang, W.-C., Yang, Y., and Liu, H · 2018
Cited alongside, same era.
Deep state space models for time series forecasting
Rangapuram, S. S., Seeger, M. W., Gasthaus, J., Stella, L., Wang, Y., and Januschowski, T · 2018
Cited alongside, same era.
BERT: pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M., Lee, K., and Toutanova, K · 2019
Cited alongside, same era.
Enhancing the locality and breaking the memory bottleneck of transformer on time series forecasting
Li, S., Jin, X., Xuan, Y., Zhou, X., Chen, W., Wang, Y.-X., and Yan, X · 2019
Cited alongside, same era.
Blockwise self-attention for long document understanding
Qiu, J., Ma, H., Levy, O., Yih, W., Wang, S., and Tang, J · 2020
Later among the works it cites.
Sparse sinkhorn attention
Tay, Y., Bahri, D., Yang, L., Metzler, D., and Juan, D · 2020
Later among the works it cites.
Linformer: Self-attention with linear complexity
Wang, S., Li, B. Z., Khabsa, M., Fang, H., and Ma, H · 2020
Later among the works it cites.
Big bird: Transformers for longer sequences
Zaheer, M., Guruganesh, G., Dubey, K. A., Ainslie, J., Alberti, C., Ontañón, S., Pham, P., Ravula, A., Wang, Q., Yang, L., and Ahmed, A · 2020
Later among the works it cites.
Rethinking attention with performers
Choromanski, K. M., Likhosherstov, V., Dohan, D., Song, X., Gane, A., Sarlós, T., Hawkins, P., Davis, J. Q., Mohiuddin, A., Kaiser, L., Belanger, D. B., Colwell, L. J., and Weller, A · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
N-BEATS: Neural basis expansion analysis for interpretable time series forecasting
Oreshkin, B. N., Carpov, D., Chapados, N., and Bengio, Y · 2019
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Kopf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S · 2019
Cited alongside, same era.
Sampled softmax with random fourier features
Rawat, A. S., Chen, J., Yu, F. X., Suresh, A. T., and Kumar, S · 2019
Cited alongside, same era.
Think globally, act locally: A deep neural network approach to high-dimensional time series forecasting
Sen, R., Yu, H., and Dhillon, I. S · 2019
Cited alongside, same era.
RobustSTL: A robust seasonal-trend decomposition algorithm for long time series
Wen, Q., Gao, J., Song, X., Sun, L., Xu, H., and Zhu, S · 2019
Cited alongside, same era.
Longformer: The long-document transformer
Beltagy, I., Peters, M. E., and Cohan, A · 2020
Cited alongside, same era.
An autocorrelation-based lstm-autoencoder for anomaly detection on time-series data
Homayouni, H., Ghosh, S., Ray, I., Gondalia, S., Duggan, J., and Kahn, M. G · 2020
Cited alongside, same era.
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N · 2021
Later among the works it cites.
Multiwavelet-based operator learning for differential equations, 2021
Gupta, G., Xiao, X., and Bogdan, P · 2021
Later among the works it cites.
Luna: Linear unified nested attention
Ma, X., Kong, X., Wang, S., Zhou, C., May, J., Ma, H., and Zettlemoyer, L · 2021
Later among the works it cites.
Global filter networks for image classification
Rao, Y., Zhao, W., Zhu, Z., Lu, J., and Zhou, J · 2021
Later among the works it cites.
Autoformer: Decomposition transformers with auto-correlation for long-term series forecasting
Wu, H., Xu, J., Wang, J., and Long, M · 2021
Later among the works it cites.
Nyströmformer: A nyström-based algorithm for approximating self-attention
Xiong, Y., Zeng, Z., Chakraborty, R., Tan, M., Fung, G., Li, Y., and Singh, V · 2021
Later among the works it cites.
Informer: Beyond efficient transformer for long sequence time-series forecasting
Zhou, H., Zhang, S., Peng, J., Zhang, S., Li, J., Xiong, H., and Zhang, W · 2021
Later among the works it cites.
H-transformer-1d: Fast one-dimensional hierarchical attention for sequences
Zhu, Z. and Soricut, R · 2021
Later among the works it cites.