Fetching the paper…
Reading the bibliography…
Many real-world applications require the prediction of long sequence time-series, such as electricity consumption planning.
Transformer-xl: Attentive language models beyond a fixed-length context
Dai, Z.; Yang, Z.; Yang, Y.; Carbonell, J.; Le, Q. V.; and Salakhutdinov, R. 2019 · 1901
Earlier work this paper cites.
Generating Long Sequences with Sparse Transformers
Child, R.; Gray, S.; Radford, A.; and Sutskever, I. 2019 · 1904
Earlier work this paper cites.
Liu, Y.; Gong, C.; Yang, L.; and Chen, Y. 2019 · 1904
Earlier work this paper cites.
Adaptively Truncating Backpropagation Through Time to Control Gradient Bias
Aicher, C.; Foti, N. J.; and Fox, E. B. 2019 · 1905
Earlier work this paper cites.
Better Long-Range Dependency By Bootstrapping A Mutual Information Regularizer
Cao, Y.; and Xu, P. 2019 · 1905
Earlier work this paper cites.
CDSA: Cross-Dimensional Self-Attention for Multivariate, Geo-tagged Time Series Imputation
Ma, J.; Shou, Z.; Zareian, A.; Mansour, H.; Vetro, A.; and Chang, S.-F. 2019 · 1905
Earlier work this paper cites.
Enhancing the Locality and Breaking the Memory Bottleneck of Transformer on Time Series Forecasting
Li, S.; Jin, X.; Xuan, Y.; Zhou, X.; Chen, W.; Wang, Y.-X.; and Yan, X. 2019 · 1907
Earlier work this paper cites.
Blockwise Self-Attention for Long Document Understanding
Qiu, J.; Ma, H.; Levy, O.; Yih, S. W.-t.; Wang, S.; and Tang, J. 2019 · 1911
Earlier work this paper cites.
Compressive transformers for long-range sequence modelling
Rae, J. W.; Potapenko, A.; Jayakumar, S. M.; and Lillicrap, T. P. 2019 · 1911
Earlier work this paper cites.
Seq-U-Net: A One-Dimensional Causal U-Net for Efficient Sequence Modelling
Stoller, D.; Tian, M.; Ewert, S.; and Dixon, S. 2019 · 1911
Earlier work this paper cites.
The meaning and measurement of size hierarchies in plant populations
Weiner, J.; and Solbrig, O. T. 1984 · 1984
Earlier work this paper cites.
Time series: theory and methods
Ray, W. 1990 · 1990
Earlier work this paper cites.
Long short-term memory
Hochreiter, S.; and Schmidhuber, J. 1997 · 1997
Earlier work this paper cites.
Bidirectional recurrent neural networks
Schuster, M.; and Paliwal, K. K. 1997 · 1997
Earlier work this paper cites.
StatStream: Statistical Monitoring of Thousands of Data Streams in Real Time
Zhu, Y.; and Shasha, D. E. 2002 · 2002
Earlier work this paper cites.
Broad distribution effects in sums of lognormal random variables
Romeo, M.; Da Costa, V.; and Bardou, F. 2003 · 2003
Earlier work this paper cites.
Longformer: The Long-Document Transformer
Beltagy, I.; Peters, M. E.; and Cohan, A. 2020 · 2004
Earlier work this paper cites.
Language Models are Few-Shot Learners
Brown, T. B.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; Agarwal, S.; Herbert-Voss, A.; Krueger, G.; Henighan, T.; Child, R.; Ramesh, A.; Ziegler, D. M.; Wu, J.; Winter, C.; Hesse, C.; Chen, M.; Sigler, E.; Litwin, M.; Gray, S.; Chess, B.; Clark, J.; Berner, C.; McCandlish, S.; Radford, A.; Sutskever, I.; and Amodei, D. 2020 · 2005
Earlier work this paper cites.
Change of Support of Transformations: Conservation of Lognormality Revisited
Vargasguzman, J. A. 2005 · 2005
Earlier work this paper cites.
Optimal multi-scale patterns in time series streams
Papadimitriou, S.; and Yu, P. 2006 · 2006
Cited alongside, same era.
Linformer: Self-Attention with Linear Complexity
Wang, S.; Li, B.; Khabsa, M.; Fang, H.; and Ma, H. 2020 · 2006
Cited alongside, same era.
Sums of lognormals
Dufresne, D. 2008 · 2008
Cited alongside, same era.
An extended limit theorem for correlated lognormal sums
Beaulieu, N. C. 2011 · 2011
Cited alongside, same era.
The sum and difference of two lognormal random variables
Lo, C.-F. 2012 · 2012
Cited alongside, same era.
Stock price prediction using the ARIMA model
Ariyo, A. A.; Adewumi, A. O.; and Ayo, C. K. 2014 · 2014
Cited alongside, same era.
Seeger, M.; Rangapuram, S.; Wang, Y.; Salinas, D.; Gasthaus, J.; Januschowski, T.; and Flunkert, V. 2017 · 2017
Later among the works it cites.
Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Later among the works it cites.
A multi-horizon quantile recurrent forecaster
Wen, R.; Torkkola, K.; Narayanaswamy, B.; and Madeka, D. 2017 · 2017
Later among the works it cites.
Dilated residual networks
Yu, F.; Koltun, V.; and Funkhouser, T. 2017 · 2017
Later among the works it cites.
Long-term forecasting using tensor-train rnns
Yu, R.; Zheng, S.; Anandkumar, A.; and Yue, Y. 2017 · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
On the Properties of Neural Machine Translation: Encoder-Decoder Approaches
Cho, K.; van Merrienboer, B.; Bahdanau, D.; and Bengio, Y. 2014 · 2014
Cited alongside, same era.
Empirical evaluation of gated recurrent neural networks on sequence modeling
Chung, J.; Gulcehre, C.; Cho, K.; and Bengio, Y. 2014 · 2014
Cited alongside, same era.
FUNNEL: automatic mining of spatially coevolving epidemics
Matsubara, Y.; Sakurai, Y.; van Panhuis, W. G.; and Faloutsos, C. 2014 · 2014
Cited alongside, same era.
Sequence to sequence learning with neural networks
Sutskever, I.; Vinyals, O.; and Le, Q. V. 2014 · 2014
Cited alongside, same era.
Neural Machine Translation by Jointly Learning to Align and Translate
Bahdanau, D.; Cho, K.; and Bengio, Y. 2015 · 2015
Cited alongside, same era.
Time series analysis: forecasting and control
Box, G. E.; Jenkins, G. M.; Reinsel, G. C.; and Ljung, G. M. 2015 · 2015
Cited alongside, same era.
Zilly, J. G.; Srivastava, R. K.; Koutník, J.; and Schmidhuber, J. 2017 · 2017
Later among the works it cites.
Convolutional sequence modeling revisited
Bai, S.; Kolter, J. Z.; and Koltun, V. 2018 · 2018
Later among the works it cites.
Log-sum-exp neural networks and posynomial models for convex and log-log-convex data
Calafiore, G. C.; Gaubert, S.; and Possieri, C. 2018 · 2018
Later among the works it cites.
A Memory-Network Based Solution for Multivariate Time-Series Forecasting
Chang, Y.-Y.; Sun, F.-Y.; Wu, Y.-H.; and Lin, S.-D. 2018 · 2018
Later among the works it cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2018 · 2018
Later among the works it cites.
Modeling long-and short-term temporal patterns with deep neural networks
Lai, G.; Chang, W.-C.; Yang, Y.; and Liu, H. 2018 · 2018
Later among the works it cites.
Diffusion Convolutional Recurrent Neural Network: Data-Driven Traffic Forecasting
Li, Y.; Yu, R.; Shahabi, C.; and Liu, Y. 2018 · 2018
Later among the works it cites.
Armdn: Associative and recurrent mixture density networks for eretail demand forecasting
Mukherjee, S.; Shankar, D.; Ghosh, A.; Tathawadekar, N.; Kompalli, P.; Sarawagi, S.; and Chaudhury, K. 2018 · 2018
Later among the works it cites.
Attend and diagnose: Clinical time series analysis using attention models
Song, H.; Rajan, D.; Thiagarajan, J. J.; and Spanias, A. 2018 · 2018
Later among the works it cites.
Forecasting at scale
Taylor, S. J.; and Letham, B. 2018 · 2018
Later among the works it cites.
Learning longer-term dependencies in rnns with auxiliary losses
Trinh, T. H.; Dai, A. M.; Luong, M.-T.; and Le, Q. V. 2018 · 2018
Later among the works it cites.
Reformer: The Efficient Transformer
Kitaev, N.; Kaiser, L.; and Levskaya, A. 2019 · 2019
Later among the works it cites.
Transformer Dissection: An Unified Understanding for Transformer’s Attention via the Lens of Kernel
Tsai, Y.-H. H.; Bai, S.; Yamada, M.; Morency, L.-P.; and Salakhutdinov, R. 2019 · 2019
Later among the works it cites.