Fetching the paper…
Reading the bibliography…
Sequence models assign probabilities to variable-length sequences such as natural language texts.
https://doi.org/10.1016/0031-8914(50)90072-3
L. van Hove, “Sur L’intégrale de Configuration Pour Les Systèmes De Particules À Une Dimension,” Physica · 1950
Earlier work this paper cites.
https://projecteuclid.org:443/euclid.cmp/1103841344
F. J. Dyson, “Existence of a phase-transition in a one-dimensional Ising ferromagnet,” Comm. Math. Phys · 1969
Earlier work this paper cites.
F. Jelinek and R. Mercer, “Interpolated estimation of Markov source parameters from sparse data,” in Proceedings of the Workshop on Pattern Recognition in Practice
1980
Earlier work this paper cites.
https://www.complex-systems.com/abstracts/v01_i01_a08/
W. Li, “Power Spectra of Regular Languages and Cellular Automata,” Complex Systems · 1987
Earlier work this paper cites.
https://doi.org/10.2307/2938368
A. W. Lo, “Long-Term Memory in Stock Market Prices,” Econometrica · 1991
Earlier work this paper cites.
https://doi.org/10.1209/0295-5075/17/7/014
W. Li and K. Kaneko, “Long-Range Correlation and Partial 1 / f α 1/f^{\alpha} Spectrum in a Noncoding DNA Sequence,” Europhysics Letters (EPL) · 1992
Earlier work this paper cites.
https://doi.org/10.1038/356168a0
C.-K. Peng, S. V. Buldyrev, A. L. Goldberger, S. Havlin, F. Sciortino, M. Simons, and H. E. Stanley, “Long-range correlations in nucleotide sequences,” Nature · 1992
Earlier work this paper cites.
https:/doi.org/10.1103/PhysRevLett.68.3805
R. F. Voss, “Evolution of long-range fractal correlations and 1 / f 1/f noise in DNA base sequences,” Phys. Rev. Lett · 1992
Earlier work this paper cites.
https://doi.org/10.1142/S0218348X93000083
A. SCHENKEL, J. ZHANG, and Y.-C. ZHANG, “LONG RANGE CORRELATION IN HUMAN WRITINGS,” Fractals · 1993
Earlier work this paper cites.
https://doi.org/10.1016/0927-5398(93)90006-D
Z. Ding, C. W. Granger, and R. F. Engle, “A long memory property of stock market returns and a new model,” Journal of Empirical Finance · 1993
Earlier work this paper cites.
https://doi.org/10.1209/0295-5075/26/4/001
W. Ebeling and T. Pöschel, “Entropy and Long-Range Correlations in Literary English,” Europhysics Letters (EPL) · 1994
Earlier work this paper cites.
https://doi.org/10.1162/neco.1997.9.8.1735
S. Hochreiter and J. Schmidhuber, “Long Short-Term Memory,” Neural Computation · 1997
Earlier work this paper cites.
https://doi.org/10.1002/(SICI)1097-4571(1999)50:14<1295::AID-ASI4>3.0.CO;2-5
P. Kokol, V. Podgorelec, M. Zorman, T. Kokol, and T. Njivar, “Computer and natural language texts—A comparison based on long-range correlations,” Journal of the American Society for Information Science · 1999
Earlier work this paper cites.
World Scientific, 1999
D. Ruelle, Statistical Mechanics: Rigorous Results · 1999
Earlier work this paper cites.
https://doi.org/10.1109/9780470544037.ch14
S. Hochreiter, Y. Bengio, and P. Frasconi, “Gradient Flow in Recurrent Nets: the Difficulty of Learning Long-Term Dependencies,” in Field Guide to Dynamical Recurrent Networks · 2001
Earlier work this paper cites.
https://doi.org/10.1109/NNSP.2002.1030094
D. Eck and J. Schmidhuber, “Finding temporal structure in music: blues improvisation with LSTM recurrent networks,” in Proceedings of the 12th IEEE Workshop on Neural Networks for Signal Processing · 2002
Earlier work this paper cites.
https://doi.org/10.1142/S0218348X02001257
M. A. MONTEMURRO and P. A. PURY, “LONG-RANGE FRACTAL CORRELATIONS IN LITERARY CORPORA,” Fractals · 2002
Earlier work this paper cites.
http://jmlr.org/papers/v3/bengio03a.html
Y. Bengio, R. Ducharme, P. Vincent, and C. Jauvin, “A Neural Probabilistic Language Model,” Journal of Machine Learning Research · 2003
Earlier work this paper cites.
https://doi.org/10.1023/B:JOSS.0000022373.63640.4e
J. A. Cuesta and A. Sánchez, “General Non-Existence Theorem for Phase Transitions in One-Dimensional Systems with Short Range Interactions, and Physical Examples of Such Transitions,” Journal of Statistical Physics · 2004
Earlier work this paper cites.
https://www.aclweb.org/anthology/J93-2004
M. P. Marcus, B. Santorini, and M. A. Marcinkiewicz, “Building a Large Annotated Corpus of English: The Penn Treebank,” Computational Linguistics · 2004
Earlier work this paper cites.
https://doi.org/10.1162/comj.2005.29.1.55
B. Manaris, J. Romero, P. Machado, D. Krehbiel, T. Hirzel, W. Pharr, and R. B. Davis, “Zipf’s Law, Music Classification, and Aesthetics,” Computer Music Journal · 2005
Cited alongside, same era.
http://mattmahoney.net/dc/textdata.html
M. Mahoney, “Relationship of Wikipedia Text to Clean Text,” 2006 · 2006
Cited alongside, same era.
https://arxiv.org/abs/physics/0307138
P. Grassberger, “Entropy Estimates from Insufficient Samplings,” 2008 · 2008
Cited alongside, same era.
Cambridge University Press, 2009
P. Flajolet and R. Sedgewick, Analytic Combinatorics · 2009
Cited alongside, same era.
http://proceedings.mlr.press/v9/glorot10a.html
X. Glorot and Y. Bengio, “Understanding the difficulty of training deep feedforward neural networks,” in Proceedings of the Thirteenth International Conference on Artificial Intelligence and Statistics · 2010
Cited alongside, same era.
https://doi.org/10.1073/pnas.1113828109
D. J. Levitin, P. Chordia, and V. Menon, “Musical rhythm spectra from Bach to Joplin obey a 1 / f 1/f power law,” Proceedings of the National Academy of Sciences · 2012
http://papers.nips.cc/paper/5846-end-to-end-memory-networks.pdf
S. Sukhbaatar, A. Szlam, J. Weston, and R. Fergus, “End-To-End Memory Networks,” in Advances in Neural Information Processing Systems 28 · 2015
Later among the works it cites.
http://papers.nips.cc/paper/5857-inferring-algorithmic-patterns-with-stack-augmented-recurrent-nets.pdf
A. Joulin and T. Mikolov, “Inferring Algorithmic Patterns with Stack-Augmented Recurrent Nets,” in Advances in Neural Information Processing Systems 28 · 2015
Later among the works it cites.
https://doi.org/10.18653/v1/K16-1028
R. Nallapati, B. Zhou, C. dos Santos, c. Gu̇lçehre, and B. Xiang, “Abstractive Text Summarization using Sequence-to-sequence RNNs and Beyond,” in Proceedings of The 20th SIGNLL Conference on Computational Natural Language Learning · 2016
Later among the works it cites.
https://arxiv.org/abs/1601.06759
A. van den Oord, N. Kalchbrenner, and K. Kavukcuoglu, “Pixel Recurrent Neural Networks,” 2016 · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
G. Hinton, N. Srivastava, and K. Swersky, “Lecture 6e—RmsProp: Divide the gradient by a running average of its recent magnitude.” Coursera: Neural Networks for Machine Learning, 2012
2012
Cited alongside, same era.
https://doi.org/10.1109/ICASSP.2013.6638947
A. Graves, A. Mohamed, and G. Hinton, “Speech recognition with deep recurrent neural networks,” in 2013 IEEE International Conference on Acoustics, Speech and Signal Processing · 2013
Cited alongside, same era.
http://papers.nips.cc/paper/5166-training-and-analysing-deep-recurrent-neural-networks.pdf
M. Hermans and B. Schrauwen, “Training and Analysing Deep Recurrent Neural Networks,” in Advances in Neural Information Processing Systems 26 · 2013
Cited alongside, same era.
Elsevier Science, 2013
L. Landau and E. Lifshitz, Statistical Physics, Volume 5 · 2013
Cited alongside, same era.
https://arxiv.org/abs/1312.6120
A. M. Saxe, J. L. McClelland, and S. Ganguli, “Exact solutions to the nonlinear dynamics of learning in deep linear neural networks,” 2013 · 2013
Cited alongside, same era.
http://papers.nips.cc/paper/5346-sequence-to-sequence-learning-with-neural-networks.pdf
I. Sutskever, O. Vinyals, and Q. V. Le, “Sequence to Sequence Learning with Neural Networks,” in Advances in Neural Information Processing Systems 27 · 2014
Cited alongside, same era.
S. Merity, C. Xiong, J. Bradbury, and R. Socher, “Pointer Sentinel Mixture Models,” 2016 · 2016
Later among the works it cites.
https://doi.org/10.18653/v1/P16-1162
R. Sennrich, B. Haddow, and A. Birch, “Neural Machine Translation of Rare Words with Subword Units,” in Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) · 2016
Later among the works it cites.
https://arxiv.org/abs/1611.01462
H. Inan, K. Khosravi, and R. Socher, “Tying Word Vectors and Word Classifiers: A Loss Framework for Language Modeling,” 2016 · 2016
Later among the works it cites.
https://arxiv.org/abs/1608.03983
I. Loshchilov and F. Hutter, “SGDR: Stochastic Gradient Descent with Warm Restarts,” 2016 · 2016
Later among the works it cites.
http://papers.nips.cc/paper/7181-attention-is-all-you-need.pdf
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is All you Need,” in Advances in Neural Information Processing Systems 30 · 2017
Later among the works it cites.
https://doi.org/10.3390/e19070299
H. Lin and M. Tegmark, “Critical Behavior in Physics and Probabilistic Formal Languages,” Entropy · 2017
Later among the works it cites.
https://arxiv.org/abs/1703.00810
R. Schwartz-ziv and N. Tishby, “Opening the black box of Deep Neural Networks via Information,” 2017 · 2017
Later among the works it cites.
https://doi.org/10.1371/journal.pone.0189326
S. Takahashi and K. Tanaka-Ishii, “Do neural nets learn statistical laws behind natural language?,” PLOS ONE · 2017
Later among the works it cites.
https://arxiv.org/abs/1707.05589
G. Melis, C. Dyer, and P. Blunsom, “On the State of the Art of Evaluation in Neural Language Models,” 2017 · 2017
Later among the works it cites.
https://arxiv.org/abs/1612.04426
E. Grave, A. Joulin, and N. Usunier, “Improving Neural Language Models with a Continuous Cache,” 2017 · 2017
Later among the works it cites.
https://arxiv.org/abs/1708.02182
S. Merity, N. S. Keskar, and R. Socher, “Regularizing and Optimizing LSTM Language Models,” aug 2017 · 2017
Later among the works it cites.
https://www.aclweb.org/anthology/P18-1027
U. Khandelwal, H. He, P. Qi, and D. Jurafsky, “Sharp Nearby, Fuzzy Far Away: How Neural Language Models Use Context,” in Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) · 2018
Later among the works it cites.
https://arxiv.org/abs/1810.05728
Z. Goldfeld, E. V. D. Berg, K. Greenewald, I. Melnyk, N. Nguyen, B. Kingsbury, and Y. Polyanskiy, “Estimating Information Flow in Neural Networks,” 2018 · 2018
Later among the works it cites.
https://d4mucfpksywv.cloudfront.net/better-language-models/language_models_are_unsupervised_multitask_learners.pdf
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever, “Language Models are Unsupervised Multitask Learners,” 2019 · 2019
Closest in time.
https://www.aclweb.org/anthology/E17-2025
O. Press and L. Wolf, “Using the Output Embedding to Improve Language Models,” in Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 2, Short Papers · 2025
Closest in time.