Fetching the paper…
Reading the bibliography…
Transformer-based architectures achieved breakthrough performance in natural language processing and computer vision, yet they remain inferior to simpler linear baselines in multivariate long-term forecasting.
Some Recent Advances in Forecasting and Control
Box, G. E. P., Jenkins, G. M., and MacGregor, J. F · 1974
Earlier work this paper cites.
A method for solving the convex programming problem with convergence rate o ( 1 / k 2 ) o(1/k^{2})
Nesterov, Y · 1983
Earlier work this paper cites.
Time Series Analysis, Forecasting and Control
Box, G. E. P. and Jenkins, G · 1990
Earlier work this paper cites.
Topics in Matrix Analysis
Horn, R. A. and Johnson, C. R · 1991
Earlier work this paper cites.
Python reference manual
Van Rossum, G. and Drake Jr, F. L · 1995
Earlier work this paper cites.
Methodology for long-term prediction of time series
Sorjamaa, A., Hao, J., Reyhani, N., Ji, Y., and Lendasse, A · 2005
Earlier work this paper cites.
Matplotlib: A 2d graphics environment
Hunter, J. D · 2007
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Glorot, X. and Bengio, Y · 2010
Earlier work this paper cites.
Guaranteed minimum-rank solutions of linear matrix equations via nuclear norm minimization
Recht, B., Fazel, M., and Parrilo, P. A · 2010
Earlier work this paper cites.
Scikit-learn: Machine learning in Python
Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., Vanderplas, J., Passos, A., Cournapeau, D., Brucher, M., Perrot, M., and Duchesnay, E · 2011
Earlier work this paper cites.
A simpler approach to matrix completion
Recht, B · 2011
Earlier work this paper cites.
Exact matrix completion via convex optimization
Candès, E. and Recht, B · 2012
Earlier work this paper cites.
TensorFlow: Large-scale machine learning on heterogeneous systems, 2015
Abadi, M., Agarwal, A., Barham, P., Brevdo, E., Chen, Z., Citro, C., Corrado, G. S., Davis, A., Dean, J., Devin, M., Ghemawat, S., Goodfellow, I., Harp, A., Irving, G., Isard, M., Jia, Y., Jozefowicz, R., Kaiser, L., Kudlur, M., Levenberg, J., Mané, D., Monga, R., Moore, S., Murray, D., Olah, C., Schuster, M., Shlens, J., Steiner, B., Sutskever, I., Talwar, K., Tucker, P., Vanhoucke, V., Vasudevan, V., Viégas, F., Vinyals, O., Warden, P., Wattenberg, M., Wicke, M., Yu, Y., and Zheng, X · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. and Ba, J · 2015
Earlier work this paper cites.
Electricity dataset, 2015
UCI · 2015
Earlier work this paper cites.
Electrocardiogram time series forecasting and optimization using ant colony optimization algorithm
Čepulionis, P. and Lukoševičiūtė, K · 2016
Earlier work this paper cites.
Entropy-SGD: Biasing gradient descent into wide valleys
Chaudhari, P., Choromanska, A., Soatto, S., LeCun, Y., Baldassi, C., Borgs, C., Chayes, J., Sagun, L., and Zecchina, R · 2017
Earlier work this paper cites.
Computing nonvacuous generalization bounds for deep (stochastic) neural networks with many more parameters than training data
Dziugaite, G. K. and Roy, D. M · 2017
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Keskar, N. S., Mudigere, D., Nocedal, J., Smelyanskiy, M., and Tang, P. T. P · 2017
Earlier work this paper cites.
SGDR: Stochastic gradient descent with warm restarts
Loshchilov, I. and Hutter, F · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L. u., and Polosukhin, I · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding, 2018
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Radford, A., Narasimhan, K., Salimans, T., and Sutskever, I · 2018
Cited alongside, same era.
Deep state space models for time series forecasting
Rangapuram, S. S., Seeger, M. W., Gasthaus, J., Stella, L., Wang, Y., and Januschowski, T · 2018
Cited alongside, same era.
Multi-horizon time series forecasting with temporal attention learning
Fan, C., Zhang, Y., Pan, Y., Li, X., Zhang, C., Yuan, R., Wu, D., Wang, W., Pei, J., and Huang, H · 2019
Cited alongside, same era.
Enhancing the locality and breaking the memory bottleneck of transformer on time series forecasting
Li, S., Jin, X., Xuan, Y., Zhou, X., Chen, W., Wang, Y.-X., and Yan, X · 2019
Cited alongside, same era.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2019
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
Signal propagation in transformers: Theoretical perspectives and the role of rank collapse
Anagnostidis, S., Biggio, L., Noci, L., Orvieto, A., Singh, S. P., and Lucchi, A · 2022
Later among the works it cites.
When vision transformers outperform resnets without pre-training or strong data augmentations
Chen, X., Hsieh, C.-J., and Gong, B · 2022
Later among the works it cites.
Triformer: Triangular, variable-specific attentions for long sequence multivariate time series forecasting
Cirstea, R.-G., Guo, C., Yang, B., Kieu, T., Dong, X., and Pan, S · 2022
Later among the works it cites.
Pyraformer: Low-complexity pyramidal attention for long-range time series modeling and forecasting
Liu, S., Yu, H., Liao, C., Li, J., Lin, W., Liu, A. X., and Dustdar, S · 2022
Later among the works it cites.
Toward understanding why adam converges faster than SGD for transformers
Pan, Y. and Li, Y · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Kopf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S · 2019
Cited alongside, same era.
Deepar: Probabilistic forecasting with autoregressive recurrent networks
Salinas, D., Flunkert, V., Gasthaus, J., and Januschowski, T · 2019
Cited alongside, same era.
Think globally, act locally: a deep neural network approach to high-dimensional time series forecasting
Sen, R., Yu, H.-F., and Dhillon, I · 2019
Cited alongside, same era.
Batch normalization provably avoids rank collapse for randomly initialised deep networks
Daneshmand, H., Kohler, J., Bach, F., Hofmann, T., and Lucchi, A · 2020
Cited alongside, same era.
Reformer: The efficient transformer
Kitaev, N., Kaiser, L., and Levskaya, A · 2020
Cited alongside, same era.
Understanding the difficulty of training transformers
Liu, L., Liu, X., Gao, J., Chen, W., and Han, J · 2020
Cited alongside, same era.
Why are adaptive methods good for attention models?
Zhang, J., Karimireddy, S. P., Veit, A., Kim, S., Reddi, S., Kumar, S., and Sra, S · 2020
Cited alongside, same era.
Restormer: Efficient transformer for high-resolution image restoration
Zamir, S. W., Arora, A., Khan, S., Hayat, M., Khan, F. S., and Yang, M.-H · 2022
Later among the works it cites.
Resnest: Split-attention networks
Zhang, H., Wu, C., Zhang, Z., Zhu, Y., Lin, H., Zhang, Z., Sun, Y., He, T., Mueller, J., Manmatha, R., Li, M., and Smola, A · 2022
Later among the works it cites.
FEDformer: Frequency enhanced decomposed transformer for long-term series forecasting
Zhou, T., Ma, Z., Wen, Q., Wang, X., Sun, L., and Jin, R · 2022
Later among the works it cites.
Linear attention is (maybe) all you need (to understand transformer optimization), 2023
Ahn, K., Cheng, X., Song, M., Yun, C., Jadbabaie, A., and Sra, S · 2023
Later among the works it cites.
TSMixer: An all-MLP architecture for time series forecasting
Chen, S.-A., Li, C.-L., Arik, S. O., Yoder, N. C., and Pfister, T · 2023
Later among the works it cites.
Deep transformers without shortcuts: Modifying self-attention for faithful signal propagation
He, B., Martens, J., Zhang, G., Botev, A., Brock, A., Smith, S. L., and Teh, Y. W · 2023
Later among the works it cites.
A time series is worth 64 words: Long-term forecasting with transformers
Nie, Y., Nguyen, N. H., Sinthong, P., and Kalagnanam, J · 2023
Later among the works it cites.
Gpt-4 technical report, 2023
OpenAI · 2023
Later among the works it cites.
Forecasting stock market prices using machine learning and deep learning models: A systematic review, performance analysis and discussion of implications
Sonkavde, G., Dharrao, D. S., Bongale, A. M., Deokate, S. T., Doreswamy, D., and Bhat, S. K · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models, 2023
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., Rodriguez, A., Joulin, A., Grave, E., and Lample, G · 2023
Later among the works it cites.
Mimetic initialization of self-attention layers
Trockman, A. and Kolter, J. Z · 2023
Later among the works it cites.
Are transformers effective for time series forecasting?
Zeng, A., Chen, M., Zhang, L., and Xu, Q · 2023
Later among the works it cites.
Stabilizing transformer training by preventing attention entropy collapse
Zhai, S., Likhomanenko, T., Littwin, E., Busbridge, D., Ramapuram, J., Zhang, Y., Gu, J., and Susskind, J. M · 2023
Later among the works it cites.
itransformer: Inverted transformers are effective for time series forecasting
Liu, Y., Hu, T., Zhang, H., Wu, H., Wang, S., Ma, L., and Long, M · 2024
Closest in time.
Unified training of universal time series forecasting transformers, 2024
Woo, G., Liu, C., Kumar, A., Xiong, C., Savarese, S., and Sahoo, D · 2024
Closest in time.
Deep learning for time series forecasting: Advances and open problems
Casolaro, A., Capone, V., Iannuzzo, G., and Camastra, F · 2078
Closest in time.