Fetching the paper…
Reading the bibliography…
Over the past years, foundation models have caused a paradigm shift in machine learning due to their unprecedented capabilities for zero-shot and few-shot generalization.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 1901
Earlier work this paper cites.
The probable error of a mean
Student · 1908
Earlier work this paper cites.
Forecasting and stock control for intermittent demands
Croston, J. D · 1972
Earlier work this paper cites.
Time Series Analysis: Forecasting and Control
Box, G. and Jenkins, G · 1976
Earlier work this paper cites.
Scoring Rules for Continuous Probability Distributions
Matheson, J. E. and Winkler, R. L · 1976
Earlier work this paper cites.
Learning to Learn: Introduction and Overview , pp. 3–17
Thrun, S. and Pratt, L · 1998
Earlier work this paper cites.
The accuracy of intermittent demand estimates
Syntetos, A. A. and Boylan, J. E · 2004
Earlier work this paper cites.
A Modern Introduction to Probability and Statistics: Understanding why and how , volume 488
Dekking, F. M., Kraaikamp, C., Lopuhaä, H. P., and Meester, L. E · 2005
Earlier work this paper cites.
Strictly proper scoring rules, prediction, and estimation
Gneiting, T. and Raftery, A. E · 2007
Earlier work this paper cites.
Matplotlib: A 2D graphics environment
Hunter, J. D · 2007
Earlier work this paper cites.
Automatic time series forecasting: The forecast package for R
Hyndman, R. J. and Khandakar, Y · 2008
Earlier work this paper cites.
Highly comparative time-series analysis: the empirical structure of time series and their methods
Fulcher, B. D., Little, M. A., and Jones, N. S · 2013
Earlier work this paper cites.
Air Quality
Vito, S · 2016
Earlier work this paper cites.
Beijing PM2.5 Data
Chen, S · 2017
Earlier work this paper cites.
An Introduction to Decision Theory
Peterson, M · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Earlier work this paper cites.
Forecasting: Principles and Practice
Hyndman, R. and Athanasopoulos, G · 2018
Earlier work this paper cites.
A multi-horizon quantile recurrent forecaster, 2018
Wen, R., Torkkola, K., Narayanaswamy, B., and Madeka, D · 2018
Earlier work this paper cites.
Beijing Multi-Site Air-Quality Data
Chen, S · 2019
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Earlier work this paper cites.
Enhancing the locality and breaking the memory bottleneck of transformer on time series forecasting
Li, S., Jin, X., Xuan, Y., Zhou, X., Chen, W., Wang, Y.-X., and Yan, X · 2019
Earlier work this paper cites.
catch22: Canonical time-series characteristics: Selected through highly comparative time-series analysis
Lubba, C. H., Sethi, S. S., Knaute, P., Schultz, S. R., Fulcher, B. D., and Jones, N. S · 2019
Earlier work this paper cites.
PyTorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Kopf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I · 2019
Earlier work this paper cites.
Root mean square layer normalization
Zhang, B. and Sennrich, R · 2019
Cited alongside, same era.
GluonTS: Probabilistic and Neural Time Series Modeling in Python
Alexandrov, A., Benidis, K., Bohlke-Schneider, M., Flunkert, V., Gasthaus, J., Januschowski, T., Maddix, D. C., Rangapuram, S., Salinas, D., Schulz, J., Stella, L., Türkmen, A. C., and Wang, Y · 2020
Cited alongside, same era.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Cited alongside, same era.
Conformer: Convolution-augmented transformer for speech recognition, 2020
Gulati, A., Qin, J., Chiu, C.-C., Parmar, N., Zhang, Y., Yu, J., Han, W., Wang, S., Zhang, Z., Wu, Y., and Pang, R · 2020
Cited alongside, same era.
Array programming with NumPy
Harris, C. R., Millman, K. J., van der Walt, S. J., Gommers, R., Virtanen, P., Cournapeau, D., Wieser, E., Taylor, J., Berg, S., Smith, N. J., Kern, R., Picus, M., Hoyer, S., van Kerkwijk, M. H., Brett, M., Haldane, A., del R’ıo, J. F., Wiebe, M., Peterson, P., G’erard-Marchant, P., Sheppard, K., Reddy, T., Weckesser, W., Abbasi, H., Gohlke, C., and Oliphant, T. E · 2020
Scaling laws vs model architectures: How does inductive bias influence scaling?, 2022
Tay, Y., Dehghani, M., Abnar, S., Chung, H. W., Fedus, W., Rao, J., Narang, S., Tran, V. Q., Yogatama, D., and Metzler, D · 2022
Later among the works it cites.
Image as a foreign language: Beit pretraining for all vision and vision-language tasks, 2022
Wang, W., Bao, H., Dong, L., Bjorck, J., Peng, Z., Liu, Q., Aggarwal, K., Mohammed, O. K., Singhal, S., Som, S., and Wei, F · 2022
Later among the works it cites.
Autoformer: Decomposition transformers with auto-correlation for long-term series forecasting, 2022
Wu, H., Xu, J., Wang, J., and Long, M · 2022
Later among the works it cites.
FEDformer: Frequency enhanced decomposed transformer for long-term series forecasting
Zhou, T., Ma, Z., Wen, Q., Wang, X., Sun, L., and Jin, R · 2022
Later among the works it cites.
Tactis-2: Better, faster, simpler attentional copulas for multivariate time series, 2023
Ashok, A., Étienne Marcotte, Zantedeschi, V., Chapados, N., and Drouin, A · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Cited alongside, same era.
pandas-dev/pandas: Pandas, February 2020
Pandas development team, T · 2020
Cited alongside, same era.
DeepAR: Probabilistic forecasting with autoregressive recurrent networks
Salinas, D., Flunkert, V., Gasthaus, J., and Januschowski, T · 2020
Cited alongside, same era.
Adversarial sparse transformer for time series forecasting
Wu, S., Xiao, X., Ding, Q., Zhao, P., Wei, Y., and Huang, J · 2020
Cited alongside, same era.
Monash time series forecasting archive
Godahewa, R., Bergmeir, C., Webb, G. I., Hyndman, R. J., and Montero-Manso, P · 2021
Cited alongside, same era.
Forecasting: Principles and practice
Hyndman, R. and Athanasopoulos, G · 2021
Cited alongside, same era.
Temporal fusion transformers for interpretable multi-horizon time series forecasting
Lim, B., Arık, S. O., Loeff, N., and Pfister, T · 2021
Cited alongside, same era.
Caballero, E., Gupta, K., Rish, I., and Krueger, D · 2023
Closest in time.
Llm4ts: Two-stage fine-tuning for time-series forecasting with pre-trained llms, 2023
Chang, C., Peng, W.-C., and Chen, T.-F · 2023
Closest in time.
Fraug: Frequency domain augmentation for time series forecasting, 2023
Chen, M., Xu, Z., Zeng, A., and Xu, Q · 2023
Closest in time.
Mitigating cold-start forecasting using cold causal demand forecasting model
Fatemi, Z., Huynh, M.-T. T., Zheleva, E., Syed, Z., and Di, X · 2023
Closest in time.
Time-llm: Time series forecasting by reprogramming large language models, 2023
Jin, M., Wang, S., Ma, L., Chu, Z., Zhang, J. Y., Shi, X., Chen, P.-Y., Liang, Y., Li, Y.-F., Pan, S., and Wen, Q · 2023
Closest in time.
How does it function? characterizing long-term trends in production serverless workloads
Joosen, A., Hassan, A., Asenov, M., Singh, R., Darlow, L., Wang, J., and Barker, A · 2023
Closest in time.
Ti-MAE: Self-supervised masked time series autoencoders, 2023
Li, Z., Wang, P., Rao, Z., Pan, L., and Xu, Z · 2023
Closest in time.
Unitime: A language-empowered unified model for cross-domain time series forecasting, 2023
Liu, X., Hu, J., Li, Y., Diao, S., Liang, Y., Hooi, B., and Zimmermann, R · 2023
Closest in time.
Basisformer: Attention-based time series forecasting with learnable and interpretable basis
Ni, Z., Yu, H., Liu, S., Li, J., and Lin, W · 2023
Closest in time.
Gpt-4 technical report, 2023
OpenAI · 2023
Closest in time.
Autogluon–timeseries: AutoML for probabilistic time series forecasting
Shchur, O., Turkmen, A. C., Erickson, N., Shen, H., Shirkov, A., Hu, T., and Wang, B · 2023
Closest in time.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., Rodriguez, A., Joulin, A., Grave, E., and Lample, G · 2023
Closest in time.
ETSformer: Exponential smoothing transformers for time-series forecasting, 2023
Woo, G., Liu, C., Sahoo, D., Kumar, A., and Hoi, S · 2023
Closest in time.
Toward a foundation model for time series data, 2023
Yeh, C.-C. M., Dai, X., Chen, H., Zheng, Y., Fan, Y., Der, A., Lai, V., Zhuang, Z., Wang, J., Wang, L., and Zhang, W · 2023
Closest in time.
TEMPO: Prompt-based generative pre-trained transformer for time series forecasting
Anonymous · 2024
Closest in time.
Cold start (recommender systems) — Wikipedia, the free encyclopedia
Wikipedia · 2024
Closest in time.
The theta model: a decomposition approach to forecasting
Assimakopoulos, V. and Nikolopoulos, K · 2070
Closest in time.
DeepAR: Probabilistic forecasting with autoregressive recurrent networks
Salinas, D., Flunkert, V., Gasthaus, J., and Januschowski, T · 2070
Closest in time.