Fetching the paper…
Reading the bibliography…
Recent work in NLP shows that LSTM language models capture hierarchical structure in language data.
Studying the inductive biases of rnns with synthetic variations of natural languages
Shauli Ravfogel, Yoav Goldberg, and Tal Linzen. 2019 · 1903
Earlier work this paper cites.
Mathijs Mul and Willem H. Zuidema. 2019 · 1906
Earlier work this paper cites.
Jaap Jumelet, Willem Zuidema, and Dieuwke Hupkes. 2019 · 1909
Earlier work this paper cites.
A value for n-person games
Lloyd S Shapley. 1953 · 1953
Earlier work this paper cites.
Three models for the description of language
N. Chomsky. 1956 · 1956
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Transferring Inductive Biases through Knowledge Distillation
Samira Abnar, Mostafa Dehghani, and Willem Zuidema. 2020 · 2006
Earlier work this paper cites.
Children’s first language acquisition from a usage-based perspective
Elena Lieven and Michael Tomasello. 2008 · 2008
Earlier work this paper cites.
A gold standard dependency corpus for English
Natalia Silveira, Timothy Dozat, Marie-Catherine de Marneffe, Samuel Bowman, Miriam Connor, John Bauer, and Christopher D. Manning. 2014 · 2014
Earlier work this paper cites.
Transition-based dependency parsing with stack long short-term memory
Chris Dyer, Miguel Ballesteros, Wang Ling, Austin Matthews, and Noah A. Smith. 2015 · 2015
Earlier work this paper cites.
A fast unified model for parsing and sentence understanding
Samuel R. Bowman, Jon Gauthier, Abhinav Rastogi, Raghav Gupta, Christopher D. Manning, and Christopher Potts. 2016 · 2016
Earlier work this paper cites.
Recurrent neural network grammars
Chris Dyer, Adhiguna Kuncoro, Miguel Ballesteros, and Noah A. Smith. 2016 · 2016
Earlier work this paper cites.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. 2016 · 2016
Earlier work this paper cites.
What Do Recurrent Neural Network Grammars Learn About Syntax?
Adhiguna Kuncoro, Chris Dyer, Miguel Ballesteros, Graham Neubig, Lingpeng Kong, and Noah A. Smith. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Boosting Neural Machine Translation
Dakun Zhang, Jungi Kim, Josep Crego, and Jean Senellart. 2017 · 2017
Cited alongside, same era.
Deep RNNs Encode Soft Hierarchical Syntax
Terra Blevins, Omer Levy, and Luke Zettlemoyer. 2018 · 2018
Cited alongside, same era.
On the Learning Dynamics of Deep Neural Networks
Remi Tachet des Combes, Mohammad Pezeshki, Samira Shabanian, Aaron Courville, and Yoshua Bengio. 2018 · 2018
Cited alongside, same era.
Colorless green recurrent networks dream hierarchically
On the practical computational power of finite precision RNNs for language recognition
Gail Weiss, Yoav Goldberg, and Eran Yahav. 2018 · 2018
Later among the works it cites.
An Empirical Exploration of Curriculum Learning for Neural Machine Translation
Xuan Zhang, Gaurav Kumar, Huda Khayrallah, Kenton Murray, Jeremy Gwinnup, Marianna J. Martindale, Paul McNamee, Kevin Duh, and Marine Carpuat. 2018 · 2018
Later among the works it cites.
What does BERT look at? an analysis of BERT’s attention
Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D. Manning. 2019 · 2019
Later among the works it cites.
A Structural Probe for Finding Syntax in Word Representations
John Hewitt and Christopher D Manning. 2019 · 2019
Later among the works it cites.
Information-theoretic analysis of multivariate single-cell signaling responses
Tomasz Jetka, Karol Nienałtowski, Tomasz Winarski, Sławomir Błoński, and Michał Komorowski. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kristina Gulordava, Piotr Bojanowski, Edouard Grave, Tal Linzen, and Marco Baroni. 2018 · 2018
Cited alongside, same era.
Visualisation and ’Diagnostic Classifiers’ Reveal how Recurrent and Recursive Neural Networks Process Hierarchical Structure (Extended Abstract)
Dieuwke Hupkes and Willem Zuidema. 2018 · 2018
Cited alongside, same era.
h-detach: Modifying the LSTM Gradient Towards Better Optimization
Bhargav Kanuparthi, Devansh Arpit, Giancarlo Kerg, Nan Rosemary Ke, Ioannis Mitliagkas, and Yoshua Bengio. 2018 · 2018
Cited alongside, same era.
Memorize or generalize? searching for a compositional rnn in a haystack
Adam Liška, Germán Kruszewski, and Marco Baroni. 2018 · 2018
Cited alongside, same era.
LSTMs Exploit Linguistic Attributes of Data
Nelson F. Liu, Omer Levy, Roy Schwartz, Chenhao Tan, and Noah A. Smith. 2018 · 2018
Cited alongside, same era.
Beyond Word Importance: Contextual Decomposition to Extract Interactions from LSTMs
W. James Murdoch, Peter J. Liu, and Bin Yu. 2018 · 2018
Cited alongside, same era.
Closing brackets with recurrent neural networks
Natalia Skachkova, Thomas Trost, and Dietrich Klakow. 2018 · 2018
Cited alongside, same era.
Transcoding compositionally: Using attention to find more generalizable solutions
Kris Korrel, Dieuwke Hupkes, Verna Dankers, and Elia Bruni. 2019 · 2019
Later among the works it cites.
The emergence of number and syntax units in LSTM language models
Yair Lakretz, German Kruszewski, Theo Desbordes, Dieuwke Hupkes, Stanislas Dehaene, and Marco Baroni. 2019 · 2019
Later among the works it cites.
Multi-element long distance dependencies: Using SPk languages to explore the characteristics of long-distance dependencies
Abhijit Mahalunkar and John Kelleher. 2019 · 2019
Later among the works it cites.
Understanding learning dynamics of language models with SVCCA
Naomi Saphra and Adam Lopez. 2019 · 2019
Later among the works it cites.
A mathematical theory of semantic development in deep neural networks
Andrew M. Saxe, James L. McClelland, and Surya Ganguli. 2019 · 2019
Later among the works it cites.
LSTM networks can perform dynamic counting
Mirac Suzgun, Yonatan Belinkov, Stuart Shieber, and Sebastian Gehrmann. 2019 · 2019
Later among the works it cites.
Generating Hierarchical Explanations on Text Classification via Feature Interaction Detection
Hanjie Chen, Guangtao Zheng, and Yangfeng Ji. 2020 · 2020
Closest in time.
Compositionality Decomposed: How do Neural Networks Generalise? (Extended Abstract)
Dieuwke Hupkes, Verna Dankers, Elia Bruni, and Mathijs Mul. 2020 · 2020
Closest in time.