Fetching the paper…
Reading the bibliography…
Large language models have become one of the most commonly deployed NLP inventions.
Elementary Principles in Statistical Mechanics
W. G. Gibbs. 1902 · 1902
Earlier work this paper cites.
Graphs and matrices: A translation of ”Graphok és matrixok” by Dénes Kőnig (1931)
Gábor Szárnyas. 2020 · 1931
Earlier work this paper cites.
The Psycho-Biology of Language
George Kingsley Zipf. 1935 · 1935
Earlier work this paper cites.
On Computable Numbers, with an Application to the Entscheidungsproblem
A. M. Turing. 1937 · 1937
Earlier work this paper cites.
The Organization of Behavior: A Neuropsychological Theory
D.O. Hebb. 1949 · 1949
Earlier work this paper cites.
“Cloze Procedure”: A new tool for measuring readability
Wilson L. Taylor. 1953 · 1953
Earlier work this paper cites.
Information theory and statistical mechanics
E. T. Jaynes. 1957 · 1957
Earlier work this paper cites.
Dropout: A simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. 2014 · 1958
Earlier work this paper cites.
On certain formal properties of grammars
Noam Chomsky. 1959 · 1959
Earlier work this paper cites.
The algebraic theory of context-free languages
N. Chomsky and M.P. Schützenberger. 1963 · 1963
Earlier work this paper cites.
Finitary models of language users
George A. Miller and Noam Chomsky. 1963 · 1963
Earlier work this paper cites.
Some methods of speeding up the convergence of iteration methods
B.T. Polyak. 1964 · 1964
Earlier work this paper cites.
Aspects of the Theory of Syntax , 50 edition
Noam Chomsky. 1965 · 1965
Earlier work this paper cites.
Counter machines and counter languages
Patrick C. Fischer, Albert R. Meyer, and Arnold L. Rosenberg. 1968 · 1968
Earlier work this paper cites.
Formal language theory: Refining the chomsky hierarchy
Gerhard Jäger and James Rogers. 2012 · 1970
Earlier work this paper cites.
Applying probability measures to abstract languages
T.L. Booth and R.A. Thompson. 1973 · 1973
Earlier work this paper cites.
Statistical analysis of non-lattice data
Julian Besag. 1975 · 1975
Earlier work this paper cites.
Continuous speech recognition by statistical methods
F. Jelinek. 1976 · 1976
Earlier work this paper cites.
Threshold matrices and the state assignment problem for neural nets
A. K. Dewdney. 1977 · 1977
Earlier work this paper cites.
Algebraic structures for transitive closure
Daniel J. Lehmann. 1977 · 1977
Earlier work this paper cites.
On the decidability of homomorphism equivalence for languages
K. Culik and Arto Salomaa. 1978 · 1978
Earlier work this paper cites.
A maximum likelihood approach to continuous speech recognition
Lalit R. Bahl, Frederick Jelinek, and Robert L. Mercer. 1983 · 1983
Earlier work this paper cites.
Van periferie naar kern
Riny Huybregts, Germen de Haan, Mieke Trommelen, and Wim Zonneveld. 1984 · 1984
Earlier work this paper cites.
Evidence against the context-freeness of natural language
Stuart M. Shieber. 1985 · 1985
Earlier work this paper cites.
Serial order: A parallel distributed processing approach
Michael I. Jordan. 1986 · 1986
Earlier work this paper cites.
Neural Nets and the brain model problem
Marvin Lee Minsky. 1986 · 1986
Earlier work this paper cites.
Real Analysis , 3 rd {}^{\text{rd}} edition
Halsey L. Royden. 1988 · 1988
Earlier work this paper cites.
Finding structure in time
Jeffrey L. Elman. 1990 · 1990
Earlier work this paper cites.
Self-Organized Language Modeling for Speech Recognition , page 450–506. Morgan Kaufmann Publishers Inc., San Francisco, CA, USA
F. Jelinek. 1990 · 1990
Earlier work this paper cites.
Efficient simulation of finite automata by neural nets
Noga Alon, A. K. Dewdney, and Teunis J. Ott. 1991 · 1991
Earlier work this paper cites.
On the computational power of neural nets
Hava T. Siegelmann and Eduardo D. Sontag. 1992 · 1992
Earlier work this paper cites.
On structuring probabilistic dependences in stochastic language modelling
Hermann Ney, Ute Essen, and Reinhard Kneser. 1994 · 1994
Earlier work this paper cites.
Probability and Measure , 3 rd {}^{\text{rd}} edition
Patrick Billingsley. 1995 · 1995
Earlier work this paper cites.
Good-turing frequency estimation without tears
William A. Gale and Geoffrey Sampson. 1995 · 1995
Earlier work this paper cites.
Optimal simulation of automata by neural nets
P. Indyk. 1995 · 1995
Earlier work this paper cites.
An empirical study of smoothing techniques for language modeling
Stanley F. Chen and Joshua Goodman. 1996 · 1996
Earlier work this paper cites.
Introduction to Probability , 2 nd {}^{\text{nd}} revised edition
Charles M. Grinstead and J. Laurie Snell. 1997 · 1997
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Chapter 25 - serial order: A parallel distributed processing approach
Michael I. Jordan. 1997 · 1997
Earlier work this paper cites.
Estimation of probabilistic context-free grammars
Zhiyi Chi and Stuart Geman. 1998 · 1998
Earlier work this paper cites.
Relating probabilistic grammars and automata
Steven Abney, David McAllester, and Fernando Pereira. 1999 · 1999
Earlier work this paper cites.
Statistical properties of probabilistic context-free grammars
Zhiyi Chi. 1999 · 1999
Earlier work this paper cites.
A neural probabilistic language model
Yoshua Bengio, Réjean Ducharme, and Pascal Vincent. 2000 · 2000
Earlier work this paper cites.
On the determinization of weighted finite automata
Adam L. Buchsbaum, Raffaele Giancarlo, and Jeffery R. Westbrook. 2000 · 2000
Cited alongside, same era.
Topology , 2 nd {}^{\text{nd}} edition
James R. Munkres. 2000 · 2000
Cited alongside, same era.
The Elements of Statistical Learning
Trevor Hastie, Robert Tibshirani, and Jerome Friedman. 2001 · 2001
Cited alongside, same era.
Fine-tuning pretrained language models: Weight initializations, data orders, and early stopping
Jesse Dodge, Gabriel Ilharco, Roy Schwartz, Ali Farhadi, Hannaneh Hajishirzi, and Noah A. Smith. 2020 · 2002
Cited alongside, same era.
Divergence measures and message passing
Thomas Minka. 2005 · 2005
Cited alongside, same era.
Neural Probabilistic Language Models , pages 137–186. Springer Berlin Heidelberg, Berlin, Heidelberg
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Later among the works it cites.
Recurrent neural networks as weighted language recognizers
Yining Chen, Sorcha Gilroy, Andreas Maletti, Jonathan May, and Kevin Knight. 2018 · 2018
Later among the works it cites.
FRAGE: Frequency-Agnostic word representation
Chengyue Gong, Di He, Xu Tan, Tao Qin, Liwei Wang, and Tie-Yan Liu. 2018 · 2018
Later among the works it cites.
Context-free transductions with neural stacks
Yiding Hao, William Merrill, Dana Angluin, Robert Frank, Noah Amsel, Andrew Benz, and Simon Mendelsohn. 2018 · 2018
Later among the works it cites.
Depth-bounding is effective: Improvements and evaluation of unsupervised PCFG induction
Lifeng Jin, Finale Doshi-Velez, Timothy Miller, William Schuler, and Lane Schwartz. 2018 · 2018
Later among the works it cites.
What is an information projection?
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yoshua Bengio, Holger Schwenk, Jean-Sébastien Senécal, Fréderic Morin, and Jean-Luc Gauvain. 2006 · 2006
Cited alongside, same era.
Pattern Recognition and Machine Learning
Christopher M. Bishop. 2006 · 2006
Cited alongside, same era.
Introduction to Automata Theory, Languages, and Computation (3rd Edition)
John E. Hopcroft, Rajeev Motwani, and Jeffrey D. Ullman. 2006 · 2006
Cited alongside, same era.
Constraints on multiple center-embedding of clauses
Fred Karlsson. 2007 · 2007
Cited alongside, same era.
Continuous space language models
Holger Schwenk. 2007 · 2007
Cited alongside, same era.
Weighted and probabilistic context-free grammars are equally expressive
Noah A. Smith and Mark Johnson. 2007 · 2007
Cited alongside, same era.
Evaluating Derivatives: Principles and Techniques of Algorithmic Differentiation , 2 nd {}^{\text{nd}} edition
Andreas Griewank and Andrea Walther. 2008 · 2008
Cited alongside, same era.
Frank Nielsen. 2018 · 2018
Later among the works it cites.
On the practical computational power of finite precision RNNs for language recognition
Gail Weiss, Yoav Goldberg, and Eran Yahav. 2018 · 2018
Later among the works it cites.
Learning classifiers with fenchel-young losses: Generalized entropies, margins, and algorithms
Mathieu Blondel, Andre Martins, and Vlad Niculae. 2019 · 2019
Later among the works it cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
Probability: Theory and Examples , 5 th {}^{\text{th}} edition
Rick Durrett. 2019 · 2019
Later among the works it cites.
A structural probe for finding syntax in word representations
John Hewitt and Christopher D. Manning. 2019 · 2019
Later among the works it cites.
What does BERT learn about the structure of language?
Ganesh Jawahar, Benoît Sagot, and Djamé Seddah. 2019 · 2019
Later among the works it cites.
On the computational power of rnns
Samuel A. Korsky and Robert C. Berwick. 2019 · 2019
Later among the works it cites.
Experimenting with power divergences for language modeling
Matthieu Labeau and Shay B. Cohen. 2019 · 2019
Later among the works it cites.
Linguistic knowledge and transferability of contextual representations
Nelson F. Liu, Matt Gardner, Yonatan Belinkov, Matthew E. Peters, and Noah A. Smith. 2019 · 2019
Later among the works it cites.
Sequential neural networks as automata
William Merrill. 2019 · 2019
Later among the works it cites.
On Losses for Modern Language Models
Stéphane Aroca-Ouellette and Frank Rudzicz. 2020 · 2020
Later among the works it cites.
On the Ability and Limitations of Transformers to Recognize Formal Languages
Satwik Bhattamishra, Kabir Ahuja, and Navin Goyal. 2020 · 2020
Later among the works it cites.
Theoretical limitations of self-attention in neural sequence models
Michael Hahn. 2020 · 2020
Later among the works it cites.
Emergent linguistic structure in artificial neural networks trained by self-supervision
Christopher D. Manning, Kevin Clark, John Hewitt, Urvashi Khandelwal, and Omer Levy. 2020 · 2020
Later among the works it cites.
Generalized entropy regularization or: There’s nothing special about label smoothing
Clara Meister, Elizabeth Salesky, and Ryan Cotterell. 2020 · 2020
Later among the works it cites.
A formal hierarchy of RNN architectures
William Merrill, Gail Weiss, Yoav Goldberg, Roy Schwartz, Noah A. Smith, and Eran Yahav. 2020 · 2020
Later among the works it cites.
oLMpics-On What Language Model Pre-training Captures
Alon Talmor, Yanai Elazar, Yoav Goldberg, and Jonathan Berant. 2020 · 2020
Later among the works it cites.
Consistency of a recurrent language model with respect to incomplete decoding
Sean Welleck, Ilia Kulikov, Jaedeok Kim, Richard Yuanzhe Pang, and Kyunghyun Cho. 2020 · 2020
Later among the works it cites.
Residual energy-based models for text
Anton Bakhtin, Yuntian Deng, Sam Gross, Myle Ott, Marc’Aurelio Ranzato, and Arthur Szlam. 2021 · 2021
Later among the works it cites.
Limitations of autoregressive models and their alternatives
Chu-Cheng Lin, Aaron Jaech, Xin Li, Matthew R. Gormley, and Jason Eisner. 2021a · 2021
Later among the works it cites.
Limitations of autoregressive models and their alternatives
Chu-Cheng Lin, Aaron Jaech, Xin Li, Matthew R. Gormley, and Jason Eisner. 2021b · 2021
Later among the works it cites.
Attention is turing-complete
Jorge Pérez, Pablo Barceló, and Javier Marinkovic. 2021 · 2021
Later among the works it cites.
A Primer in BERTology: What We Know About How BERT Works
Anna Rogers, Olga Kovaleva, and Anna Rumshisky. 2021 · 2021
Later among the works it cites.
Probing classifiers: Promises, shortcomings, and advances
Yonatan Belinkov. 2022 · 2022
Later among the works it cites.
Algorithms for weighted pushdown automata
Alexandra Butoi, Brian DuSell, Tim Vieira, Ryan Cotterell, and David Chiang. 2022 · 2022
Later among the works it cites.
Overcoming a theoretical limitation of self-attention
David Chiang and Peter Cholak. 2022 · 2022
Later among the works it cites.
Neural networks and the chomsky hierarchy
Gr’egoire Del’etang, Anian Ruoss, Jordi Grau-Moya, Tim Genewein, Li Kevin Wenliang, Elliot Catt, Marcus Hutter, Shane Legg, and Pedro A. Ortega. 2022 · 2022
Later among the works it cites.
A measure-theoretic characterization of tight language models
Li Du, Lucas Torroba Hennigen, Tiago Pimentel, Clara Meister, Jason Eisner, and Ryan Cotterell. 2022 · 2022
Later among the works it cites.
On the uncomputability of partition functions in energy-based sequence models
Chu-Cheng Lin and Arya D. McCarthy. 2022 · 2022
Later among the works it cites.
Grokking: Generalization beyond overfitting on small algorithmic datasets
Alethea Power, Yuri Burda, Harrison Edwards, Igor Babuschkin, and Vedant Misra. 2022 · 2022
Later among the works it cites.
The multiBERTs: BERT reproductions for robustness analysis
Thibault Sellam, Steve Yadlowsky, Ian Tenney, Jason Wei, Naomi Saphra, Alexander D’Amour, Tal Linzen, Jasmijn Bastings, Iulia Raluca Turc, Jacob Eisenstein, Dipanjan Das, and Ellie Pavlick. 2022 · 2022
Later among the works it cites.
The parallelism tradeoff: Limitations of log-precision transformers
William Merrill and Ashish Sabharwal. 2023 · 2023
Closest in time.
On the representational capacity of recurrent neural language models
Franz Nowak, Anej Svete, Li Du, and Ryan Cotterell. 2023 · 2023
Closest in time.
Distilling reasoning capabilities into smaller language models
Kumar Shridhar, Alessandro Stolfo, and Mrinmaya Sachan. 2023 · 2023
Closest in time.
Recurrent neural language models as probabilistic finite-state automata
Anej Svete and Ryan Cotterell. 2023b · 2023
Closest in time.
A theoretical result on the inductive bias of rnn language models
Anej Svete, Robin Shing Moon Chan, and Ryan Cotterell. 2024 · 2024
Closest in time.