Fetching the paper…
Reading the bibliography…
Transformers are emerging as the new workhorse of NLP, showing great success across tasks.
Generating long sequences with sparse transformers
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever. 2019 · 1904
Earlier work this paper cites.
On the computational power of RNNs
Samuel A. Korsky and Robert C. Berwick. 2019 · 1906
Earlier work this paper cites.
R Thomas McCoy, Junghyun Min, and Tal Linzen. 2019 · 1911
Earlier work this paper cites.
Syntactic structures
Noam Chomsky. 1957 · 1957
Earlier work this paper cites.
The algebraic theory of context-free languages
Noam Chomsky and Marcel P Schützenberger. 1963 · 1963
Earlier work this paper cites.
Finitary models of language users
George A Miller and Noam Chomsky. 1963 · 1963
Earlier work this paper cites.
Counter-Free Automata (MIT research monograph no. 65)
Robert McNaughton and Seymour A Papert. 1971 · 1971
Earlier work this paper cites.
Parity, circuits, and the polynomial-time hierarchy
Merrick Furst, James B Saxe, and Michael Sipser. 1984 · 1984
Earlier work this paper cites.
Evidence against the context-freeness of natural language
Stuart M Shieber. 1985 · 1985
Earlier work this paper cites.
Regular languages in NC1
David A Mix Barrington, Kevin Compton, Howard Straubing, and Denis Thérien. 1992 · 1992
Earlier work this paper cites.
Learning long-term dependencies with gradient descent is difficult
Yoshua Bengio, Patrice Simard, Paolo Frasconi, et al. 1994 · 1994
Earlier work this paper cites.
Optimal depth, very small size circuits for symmetrical functions in AC0
Johan Hastad, Ingo Wegener, Norbert Wurm, and Sang-Zin Yi. 1994 · 1994
Earlier work this paper cites.
On the computational power of neural nets
Hava Siegelman and Eduardo D Sontag. 1995 · 1995
Earlier work this paper cites.
The average sensitivity of bounded-depth circuits
Ravi B Boppana. 1997 · 1997
Earlier work this paper cites.
Computation in recurrent neural networks: From counters to iterated function systems
Yvonne Kalinke and Helko Lehmann. 1998 · 1998
Earlier work this paper cites.
Memory limitations and structural forgetting: The perception of complex ungrammatical sentences as grammatical
Edward Gibson and James Thomas. 1999 · 1999
Earlier work this paper cites.
Fractal encoding of context-free grammars in connectionist networks
Whitney Tabor. 2000 · 2000
Earlier work this paper cites.
LSTM recurrent networks learn simple context-free and context-sensitive languages
Felix A Gers and Jürgen Schmidhuber. 2001 · 2001
Earlier work this paper cites.
An activation-based model of sentence processing as skilled memory retrieval
Richard L Lewis and Shravan Vasishth. 2005 · 2005
Earlier work this paper cites.
Stack-like and queue-like dynamics in recurrent neural networks
André Grüning. 2006 · 2006
Cited alongside, same era.
On the implicit acquisition of a context-free grammar by a simple recurrent neural network
Bo Cartling. 2008 · 2008
Cited alongside, same era.
Processing of nested and cross-serial dependencies: an automaton perspective on SRN behaviour
Christo Kirov and Robert Frank. 2012 · 2012
Cited alongside, same era.
Structures, not strings: linguistics as part of the cognitive sciences
Martin BH Everaert, Marinus AC Huybregts, Noam Chomsky, Robert C Berwick, and Johan J Bolhuis. 2015 · 2015
Cited alongside, same era.
Long short-term memory-networks for machine reading
Jianpeng Cheng, Li Dong, and Mirella Lapata. 2016 · 2016
Cited alongside, same era.
Smooth boolean functions are easy: Efficient algorithms for low-sensitivity functions
Evaluating the ability of LSTMs to learn context-free grammars
Luzi Sennhauser and Robert Berwick. 2018 · 2018
Later among the works it cites.
Closing brackets with recurrent neural networks
Natalia Skachkova, Thomas Trost, and Dietrich Klakow. 2018 · 2018
Later among the works it cites.
The importance of being recurrent for modeling hierarchical structure
Ke Tran, Arianna Bisazza, and Christof Monz. 2018 · 2018
Later among the works it cites.
On the practical computational power of finite precision rnns for language recognition
Gail Weiss, Yoav Goldberg, and Eran Yahav. 2018 · 2018
Later among the works it cites.
What does BERT look at? An analysis of BERT’s attention
Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D. Manning. 2019 · 2019
Closest in time.
Transformer-XL: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime G. Carbonell, Quoc Viet Le, and Ruslan Salakhutdinov. 2019 · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Parikshit Gopalan, Noam Nisan, Rocco A Servedio, Kunal Talwar, and Avi Wigderson. 2016 · 2016
Cited alongside, same era.
Assessing the ability of LSTMs to learn syntax-sensitive dependencies
Tal Linzen, Emmanuel Dupoux, and Yoav Goldberg. 2016 · 2016
Cited alongside, same era.
A decomposable attention model for natural language inference
Ankur Parikh, Oscar Täckström, Dipanjan Das, and Jakob Uszkoreit. 2016 · 2016
Cited alongside, same era.
A structured self-attentive sentence embedding
Zhouhan Lin, Minwei Feng, Cicero Nogueira dos Santos, Mo Yu, Bing Xiang, Bowen Zhou, and Yoshua Bengio. 2017 · 2017
Cited alongside, same era.
Probability and Computing , 2nd edition
Michael Mitzenmacher and Eli Upfal. 2017 · 2017
Cited alongside, same era.
The cue-based retrieval theory of sentence comprehension: New findings and new challenges
Dan Parker, Michael Shvartsman, and Julie A Van Dyke. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Closest in time.
Universal transformers
Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, and Lukasz Kaiser. 2019 · 2019
Closest in time.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Closest in time.
Modeling recurrence for transformer
Jie Hao, Xing Wang, Baosong Yang, Longyue Wang, Jinfeng Zhang, and Zhaopeng Tu. 2019 · 2019
Closest in time.
On the robustness of self-attentive models
Yu-Lun Hsieh, Minhao Cheng, Da-Cheng Juan, Wei Wei, Wen-Lian Hsu, and Cho-Jui Hsieh. 2019 · 2019
Closest in time.
Open Sesame: Getting inside BERT’s linguistic knowledge
Yongjie Lin, Yi Chern Tan, and Robert Frank. 2019 · 2019
Closest in time.
Sequential neural networks as automata
William Merrill. 2019 · 2019
Closest in time.
Stable recurrent models
John Miller and Moritz Hardt. 2019 · 2019
Closest in time.
On the Turing completeness of modern neural network architectures
Jorge Pérez, Javier Marinković, and Pablo Barceló. 2019 · 2019
Closest in time.
On evaluating the generalization of LSTM models in formal languages
Mirac Suzgun, Yonatan Belinkov, and Stuart M Shieber. 2019 · 2019
Closest in time.
BERT rediscovers the classical NLP pipeline
Ian Tenney, Dipanjan Das, and Ellie Pavlick. 2019 · 2019
Closest in time.
Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned
Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, and Ivan Titov. 2019 · 2019
Closest in time.
Assessing the ability of self-attention networks to learn word order
Baosong Yang, Longyue Wang, Derek F. Wong, Lidia S. Chao, and Zhaopeng Tu. 2019 · 2019
Closest in time.