Fetching the paper…
Reading the bibliography…
In order to build efficient deep recurrent neural architectures, it is essential to analyze the complexityof long distance dependencies (LDDs) of the dataset being modeled.
The mathematical theory of communication
Claude Elwood Shannon and Warren Weaver · 1949
Earlier work this paper cites.
Three models for the description of language
N. Chomsky · 1956
Earlier work this paper cites.
On certain formal properties of grammars
Noam Chomsky · 1959
Earlier work this paper cites.
Transmission of Information: A Statistical Theory of Communication
Robert M. Fano · 1961
Earlier work this paper cites.
Implicit learning of artificial grammars
Arthur S. Reber · 1967
Earlier work this paper cites.
Learning of construction of finite automata from examples using hill-climbing : RR : Regular set Recognizer
Masaru Tomita · 1982
Earlier work this paper cites.
Encoding sequential structure: experience with the real-time recurrent learning algorithm
A. W. Smith and D. Zipser · 1989
Earlier work this paper cites.
Finding structure in time
Jeffrey L. Elman · 1990
Earlier work this paper cites.
Untersuchungen zu dynamischen neuronalen netzen
Sepp Hochreiter · 1991
Earlier work this paper cites.
Elements of Information Theory
Thomas M. Cover and Joy A. Thomas · 1991
Earlier work this paper cites.
Induction of finite-state automata using second-order recurrent networks
Raymond L. Watrous and Gary M. Kuhn · 1991
Earlier work this paper cites.
Long-range correlations in nucleotide sequences
C. K. Peng, S. V. Buldyrev, A. L. Goldberger, S. Havlin, F. Sciortino, M. Simons, and H. E. Stanley · 1992
Earlier work this paper cites.
Learning and extracting finite state automata with second-order recurrent neural networks
C. L. Giles, C. B. Miller, D. Chen, H. H. Chen, G. Z. Sun, and Y. C. Lee · 1992
Earlier work this paper cites.
Learning long-term dependencies with gradient descent is difficult
Y. Bengio, P. Simard, and P. Frasconi · 1994
Earlier work this paper cites.
The penn treebank: Annotating predicate argument structure
Mitchell Marcus, Grace Kim, Mary Ann Marcinkiewicz, Robert MacIntyre, Ann Bies, Mark Ferguson, Karen Katz, and Britta Schasberger · 1994
Earlier work this paper cites.
Entropy and long-range correlations in literary english
W Ebeling and T Pöschel · 1994
Earlier work this paper cites.
Hierarchical recurrent neural networks for long-term dependencies
Salah El Hihi and Yoshua Bengio · 1995
Earlier work this paper cites.
The dynamics of discrete-time computation, with application to recurrent neural networks and finite state machine extraction
M. Casey · 1996
Cited alongside, same era.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Cited alongside, same era.
The dynamics and light curves of beamed gamma-ray burst afterglows
James E. Rhoads · 1999
Cited alongside, same era.
Gradient flow in recurrent nets: the difficulty of learning long-term dependencies
Sepp Hochreiter, Yoshua Bengio, and Paolo Frasconi · 2001
Cited alongside, same era.
Long-range fractal correlations in literary corpora
Marcelo A. Montemurro and Pedro A. Pury · 2002
Cited alongside, same era.
Estimation of entropy and mutual information
Liam Paninski · 2003
Cited alongside, same era.
Entropy and long-range correlations in dna sequences
S. S. Melnik and O. V. Usatenko · 2014
Later among the works it cites.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2016
Later among the works it cites.
Critical behavior in physics and probabilistic formal languages
Henry W. Lin and Max Tegmark · 2017
Later among the works it cites.
Dilated recurrent neural networks
Shiyu Chang, Yang Zhang, Wei Han, Mo Yu, Xiaoxiao Guo, Wei Tan, Xiaodong Cui, Michael Witbrock, Mark A Hasegawa-Johnson, and Thomas S Huang · 2017
Later among the works it cites.
Attentive language models
Giancarlo D. Salton, Robert J. Ross, and John D. Kelleher · 2017
Later among the works it cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Entropy Estimates from Insufficient Samplings
P. Grassberger · 2003
Cited alongside, same era.
Afterglow light curves and broken power laws: A statistical study
Gudlaugur Jóhannesson, Gunnlaugur Bjöörnsson, and Einar H. Gudmundsson · 2006
Cited alongside, same era.
Normalized (pointwise) mutual information in collocation extraction
Gerlof Bouma · 2009
Cited alongside, same era.
Foma: a finite-state compiler and library
Mans Hulden · 2009
Cited alongside, same era.
On Languages Piecewise Testable in the Strict Sense
James Rogers, Jeffrey Heinz, Gil Bailey, Matt Edlefsen, Molly Visscher, David Wellcome, and Sean Wibel · 2010
Cited alongside, same era.
Estimating strictly piecewise distributions
Jeffrey Heinz and James Rogers · 2010
Cited alongside, same era.
Later among the works it cites.
Subregular complexity and deep learning
Enes Avcu, Chihiro Shibata, and Jeffrey Heinz · 2017
Later among the works it cites.
Skip rnn: Learning to skip state updates in recurrent neural networks
Víctor Campos, Brendan Jou, Xavier Giró-i Nieto, Jordi Torres, and Shih-Fu Chang · 2018
Closest in time.
Using regular languages to explore the representational capacity of recurrent neural architectures
Abhijit Mahalunkar and John D. Kelleher · 2018
Closest in time.
Frage: Frequency-agnostic word representation
Chengyue Gong, Di He, Xu Tan, Tao Qin, Liwei Wang, and Tie-Yan Liu · 2018
Closest in time.
Direct output connection for a high-rank language model
Sho Takase, Jun Suzuki, and Masaaki Nagata · 2018
Closest in time.
Breaking the softmax bottleneck: A high-rank RNN language model
Zhilin Yang, Zihang Dai, Ruslan Salakhutdinov, and William W. Cohen · 2018
Closest in time.
Dynamic evaluation of neural sequence models
Ben Krause, Emmanuel Kahembwe, Iain Murray, and Steve Renals · 2018
Closest in time.
Regularizing and optimizing LSTM language models
Stephen Merity, Nitish Shirish Keskar, and Richard Socher · 2018
Closest in time.
Fast parametric learning with activation memorization
Jack Rae, Chris Dyer, Peter Dayan, and Timothy Lillicrap · 2018
Closest in time.
Transformer-XL: Language modeling with longer-term dependency, 2019
Zihang Dai, Zhilin Yang, Yiming Yang, William W. Cohen, Jaime Carbonell, Quoc V. Le, and Ruslan Salakhutdinov · 2019
Closest in time.
Adaptive input representations for neural language modeling
Alexei Baevski and Michael Auli · 2019
Closest in time.