Fetching the paper…
Reading the bibliography…
Natural languages are believed to be (mildly) context-sensitive.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D. Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 1901
Earlier work this paper cites.
Memory-augmented recurrent neural networks can learn generalized Dyck languages
Mirac Suzgun, Sebastian Gehrmann, Yonatan Belinkov, and Stuart M. Shieber. 2019 · 1911
Earlier work this paper cites.
Three models for the description of language
Noam Chomsky. 1956 · 1956
Earlier work this paper cites.
Syntactic Structures
Noam Chomsky. 1957 · 1957
Earlier work this paper cites.
Automatic syntactic analysis and the pushdown store
Anthony G. Oettinger. 1961 · 1961
Earlier work this paper cites.
The algebraic theory of context-free languages
Noam Chomsky and Marcel Paul Schützenberger. 1963 · 1963
Earlier work this paper cites.
Application of pushdown-store machines
R. James Evey. 1963 · 1963
Earlier work this paper cites.
On context-free languages and push-down automata
Marcel Paul Schützenberger. 1963 · 1963
Earlier work this paper cites.
On finite monoids having only trivial subgroups
M.P. Schützenberger. 1965 · 1965
Earlier work this paper cites.
Counter-free automata
Robert McNaughton and Seymour Papert. 1971 · 1971
Earlier work this paper cites.
Evidence Against the Context-Freeness of Natural Language , pages 320–334. Springer Netherlands
Stuart M. Shieber. 1987 · 1987
Earlier work this paper cites.
The induction of dynamical recognizers
Jordan B. Pollack. 1991 · 1991
Earlier work this paper cites.
Learning context-free grammars: Capabilities and limitations of a recurrent neural network with an external stack memory
Sreerupa Das, C. Lee Giles, and Guo-Zheng Sun. 1992 · 1992
Earlier work this paper cites.
A connectionist symbol manipulator that discovers the structure of context-free languages
Michael C. Mozer and Sreerupa Das. 1992 · 1992
Earlier work this paper cites.
The Penn Treebank: Annotating predicate argument structure
Mitchell Marcus, Grace Kim, Mary Ann Marcinkiewicz, Robert MacIntyre, Ann Bies, Mark Ferguson, Karen Katz, and Britta Schasberger. 1994 · 1994
Earlier work this paper cites.
Discrete recurrent neural networks for grammatical inference
Zheng Zeng, Rodney M. Goodman, and Padhraic Smyth. 1994 · 1994
Earlier work this paper cites.
Homotopy Type Theory: Univalent Foundations of Mathematics
The Univalent Foundations Program. 2013 · 2013
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
Learning to transduce with unbounded memory
Edward Grefenstette, Karl Moritz Hermann, Mustafa Suleyman, and Phil Blunsom. 2015 · 2015
Earlier work this paper cites.
Inferring algorithmic patterns with stack-augmented recurrent nets
Armand Joulin and Tomas Mikolov. 2015 · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba. 2015 · 2015
Cited alongside, same era.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton. 2016 · 2016
Cited alongside, same era.
Long short-term memory-networks for machine reading
Jianpeng Cheng, Li Dong, and Mirella Lapata. 2016 · 2016
Cited alongside, same era.
Recurrent neural network grammars
Chris Dyer, Adhiguna Kuncoro, Miguel Ballesteros, and Noah A. Smith. 2016 · 2016
Cited alongside, same era.
A decomposable attention model for natural language inference
Datasets: A community library for natural language processing
Quentin Lhoest, Albert Villanova del Moral, Yacine Jernite, Abhishek Thakur, Patrick von Platen, Suraj Patil, Julien Chaumond, Mariama Drame, Julien Plu, Lewis Tunstall, Joe Davison, Mario Šaško, Gunjan Chhablani, Bhavitvya Malik, Simon Brandeis, Teven Le Scao, Victor Sanh, Canwen Xu, Nicolas Patry, Angelina McMillan-Major, Philipp Schmid, Sylvain Gugger, Clément Delangue, Théo Matussière, Lysandre Debut, Stas Bekman, Pierric Cistac, Thibault Goehringer, Victor Mustar, François Lagunas, Alexander Rush, and Thomas Wolf. 2021 · 2021
Later among the works it cites.
Attention is Turing-complete
Jorge Pérez, Pablo Barceló, and Javier Marinkovic. 2021 · 2021
Later among the works it cites.
Thinking like transformers
Gail Weiss, Yoav Goldberg, and Eran Yahav. 2021 · 2021
Later among the works it cites.
Learning hierarchical structures with differentiable nondeterministic stacks
Brian DuSell and David Chiang. 2022 · 2022
Later among the works it cites.
Formal language recognition by hard attention transformers: Perspectives from circuit complexity
Yiding Hao, Dana Angluin, and Robert Frank. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ankur Parikh, Oscar Täckström, Dipanjan Das, and Jakob Uszkoreit. 2016 · 2016
Cited alongside, same era.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Context-free transductions with neural stacks
Yiding Hao, William Merrill, Dana Angluin, Robert Frank, Noah Amsel, Andrew Benz, and Simon Mendelsohn. 2018 · 2018
Cited alongside, same era.
Memory architectures in recurrent neural network language models
Dani Yogatama, Yishu Miao, Gabor Melis, Wang Ling, Adhiguna Kuncoro, Chris Dyer, and Phil Blunsom. 2018 · 2018
Cited alongside, same era.
Transformer-XL: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc Le, and Ruslan Salakhutdinov. 2019 · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Train short, test long: Attention with linear biases enables input length extrapolation
Ofir Press, Noah Smith, and Mike Lewis. 2022 · 2022
Later among the works it cites.
Transformer grammars: Augmenting transformer language models with syntactic inductive biases at scale
Laurent Sartran, Samuel Barrett, Adhiguna Kuncoro, Miloš Stanojević, Phil Blunsom, and Chris Dyer. 2022 · 2022
Later among the works it cites.
Masked hard-attention transformers and boolean RASP recognize exactly the star-free languages
Dana Angluin, David Chiang, and Andy Yang. 2023 · 2023
Later among the works it cites.
Tighter bounds on the expressivity of transformer encoders
David Chiang, Peter Cholak, and Anand Pillay. 2023 · 2023
Later among the works it cites.
Neural networks and the Chomsky hierarchy
Grégoire Delétang, Anian Ruoss, Jordi Grau-Moya, Tim Genewein, Li Kevin Wenliang, Elliot Catt, Chris Cundy, Marcus Hutter, Shane Legg, Joel Veness, and Pedro A. Ortega. 2023 · 2023
Later among the works it cites.
Towards revealing the mystery behind chain of thought: A theoretical perspective
Guhao Feng, Bohang Zhang, Yuntian Gu, Haotian Ye, Di He, and Liwei Wang. 2023 · 2023
Later among the works it cites.
A logic for expressing log-precision transformers
William Merrill and Ashish Sabharwal. 2023 · 2023
Later among the works it cites.
Pushdown layers: Encoding recursive structure in transformer language models
Shikhar Murty, Pratyusha Sharma, Jacob Andreas, and Christopher Manning. 2023 · 2023
Later among the works it cites.
RoFormer: Enhanced transformer with rotary position embedding
Jianlin Su, Murtadha Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu. 2023 · 2023
Later among the works it cites.
Logical languages accepted by transformer encoders with hard attention
Pablo Barcelo, Alexander Kozachinskiy, Anthony Widjaja Lin, and Vladimir Podolskii. 2024 · 2024
Closest in time.
Stack attention: Improving the ability of transformers to model hierarchical patterns
Brian DuSell and David Chiang. 2024 · 2024
Closest in time.
The expressive power of transformers with chain of thought
William Merrill and Ashish Sabharwal. 2024 · 2024
Closest in time.
Transformers can represent n n -gram language models
Anej Svete and Ryan Cotterell. 2024 · 2024
Closest in time.