Fetching the paper…
Reading the bibliography…
This technical report describes the Time Series Optimized Transformer for Observability (Toto), a new state of the art foundation model for time series forecasting developed by Datadog.
Long-range Forecasting: From Crystal Ball to Computer
J. Scott Armstrong · 1985
Earlier work this paper cites.
Robust mixture modelling using the t distribution
D. Peel and G.J. McLachlan · 2000
Earlier work this paper cites.
Glu variants improve transformer, 2020
Noam Shazeer · 2002
Earlier work this paper cites.
Another look at measures of forecast accuracy
R. J Hyndman and A. B. Koehler · 2006
Earlier work this paper cites.
A comparison of the forecasting ability of arima models
Simon Stevenson · 2007
Earlier work this paper cites.
A student t -mixture autoregressive model with applications to heavy-tailed financial data
C. S. WONG, W. S. CHAN, and P. L. KAM · 2009
Earlier work this paper cites.
Forecasting with limited data: Combining arima and diffusion models
Charisios Christodoulos, Christos Michalakelis, and Dimitris Varoutas · 2010
Earlier work this paper cites.
Deep Learning
Ian Goodfellow, Yoshua Bengio, and Aaron Courville · 2016
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Exploring the hidden dimension in accelerating convolutional neural networks, 2018
Zhihao Jia, Sina Lin, Charles R Qi, and Alex Aiken · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford and Karthik Narasimhan · 2018
Earlier work this paper cites.
A mixture autoregressive model based on student’s t–distribution
Mika Meitz, Daniel P. A. Preve, and Pentti Saikkonen · 2018
Earlier work this paper cites.
Deepar: Probabilistic forecasting with autoregressive recurrent networks
David Salinas, Valentin Flunkert, Jan Gasthaus, and Tim Januschowski · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Earlier work this paper cites.
Root Mean Square Layer Normalization
Biao Zhang and Rico Sennrich · 2019
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2019
Earlier work this paper cites.
Informer: Beyond efficient transformer for long sequence time-series forecasting
Haoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang, Jianxin Li, Hui Xiong, and Wan Zhang · 2020
Cited alongside, same era.
On layer normalization in the transformer architecture, 2020
Ruibin Xiong, Yunchang Yang, Di He, Kai Zheng, Shuxin Zheng, Huishuai Zhang, Yanyan Lan, Liwei Wang, and Tie-Yan Liu · 2020
Cited alongside, same era.
On the relationship between self-attention and convolutional layers
Jean-Baptiste Cordonnier, Andreas Loukas, and Martin Jaggi · 2020
Cited alongside, same era.
Forecasting: Principles and Practice
Rob J Hyndman and George Athanasopoulos · 2021
Cited alongside, same era.
Parallelizing dnn training on gpus: Challenges and opportunities
Weizheng Xu, Youtao Zhang, and Xulong Tang · 2021
Cited alongside, same era.
Autoformer: Decomposition transformers with auto-correlation for long-term series forecasting
Generative adversarial networks in time series: A systematic literature review
Eoin Brophy, Zhengwei Wang, Qi She, and Tomás Ward · 2023
Later among the works it cites.
A time series is worth 64 words: Long-term forecasting with transformers
Yuqi Nie, Nam H Nguyen, Phanwadee Sinthong, and Jayant Kalagnanam · 2023
Later among the works it cites.
Crossformer: Transformer utilizing cross-dimension dependency for multivariate time series forecasting
Yunhao Zhang and Junchi Yan · 2023
Later among the works it cites.
Timegpt-1, 2023
Azul Garza and Max Mergenthaler-Canseco · 2023
Later among the works it cites.
Lag-llama: Towards foundation models for time series forecasting
Kashif Rasul, Arjun Ashok, Andrew Robert Williams, Arian Khorasani, George Adamopoulos, Rishika Bhagwatkar, Marin Biloš, Hena Ghonia, Nadhir Hassen, Anderson Schneider, Sahil Garg, Alexandre Drouin, Nicolas Chapados, Yuriy Nevmyvaka, and Irina Rish · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Haixu Wu, Jiehui Xu, Jianmin Wang, and Mingsheng Long · 2021
Cited alongside, same era.
Tianyang Lin, Yuxin Wang, Xiangyang Liu, and Xipeng Qiu · 2021
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Cited alongside, same era.
Msa transformer
Roshan M Rao, Jason Liu, Robert Verkuil, Joshua Meier, John Canny, Pieter Abbeel, Tom Sercu, and Alexander Rives · 2021
Cited alongside, same era.
Vivit: A video vision transformer
Anurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun, Mario Lučić, and Cordelia Schmid · 2021
Cited alongside, same era.
Roformer: Enhanced transformer with rotary position embedding, 2021
Jianlin Su, Yu Lu, Shengfeng Pan, Bo Wen, and Yunfeng Liu · 2021
Cited alongside, same era.
A length-extrapolatable transformer
Yutao Sun, Li Dong, Barun Patra, Shuming Ma, Shaohan Huang, Alon Benhaim, Vishrav Chaudhary, Xia Song, and Furu Wei · 2022
Cited alongside, same era.
Nate Gruver, Marc Anton Finzi, Shikai Qiu, and Andrew Gordon Wilson · 2023
Later among the works it cites.
Long-term forecasting with tiDE: Time-series dense encoder
Abhimanyu Das, Weihao Kong, Andrew Leach, Shaan K Mathur, Rajat Sen, and Rose Yu · 2023
Later among the works it cites.
Timesnet: Temporal 2d-variation modeling for general time series analysis
Haixu Wu, Tengge Hu, Yong Liu, Hang Zhou, Jianmin Wang, and Mingsheng Long · 2023
Later among the works it cites.
Are transformers effective for time series forecasting?
Ailing Zeng, Muxi Chen, Lei Zhang, and Qiang Xu · 2023
Later among the works it cites.
Unified training of universal time series forecasting transformers
Gerald Woo, Chenghao Liu, Akshat Kumar, Caiming Xiong, Silvio Savarese, and Doyen Sahoo · 2024
Closest in time.
itransformer: Inverted transformers are effective for time series forecasting
Yong Liu, Tengge Hu, Haoran Zhang, Haixu Wu, Shiyu Wang, Lintao Ma, and Mingsheng Long · 2024
Closest in time.
SAMformer: Unlocking the potential of transformers in time series forecasting with sharpness-aware minimization and channel-wise attention
Romain Ilbert, Ambroise Odonnat, Vasilii Feofanov, Aladin Virmaux, Giuseppe Paolo, Themis Palpanas, and Ievgen Redko · 2024
Closest in time.
A decoder-only foundation model for time-series forecasting
Abhimanyu Das, Weihao Kong, Rajat Sen, and Yichen Zhou · 2024
Closest in time.
Chronos: Learning the language of time series, 2024
Abdul Fatir Ansari, Lorenzo Stella, Caner Turkmen, Xiyuan Zhang, Pedro Mercado, Huibin Shen, Oleksandr Shchur, Syama Sundar Rangapuram, Sebastian Pineda Arango, Shubham Kapoor, Jasper Zschiegner, Danielle C. Maddix, Hao Wang, Michael W. Mahoney, Kari Torkkola, Andrew Gordon Wilson, Michael Bohlke-Schneider, and Yuyang Wang · 2024
Closest in time.
Generalising about univariate forecasting methods: further empirical evidence
Robert Fildes, Michèle Hibon, Spyros Makridakis, and Nigel Meade · 2070
Closest in time.