Fetching the paper…
Reading the bibliography…
Standard autoregressive language models perform only polynomial-time computation to compute the probability of the next symbol.
Generating long sequences with sparse transformers
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever. 2019 · 1904
Earlier work this paper cites.
The complexity of theorem-proving procedures
Stephen A. Cook. 1971 · 1971
Earlier work this paper cites.
The circuit value problem is log \log space complete for P \mathrm{P}
Richard E. Ladner. 1975 · 1975
Earlier work this paper cites.
Some connections between nonuniform and uniform complexity classes
Richard M. Karp and Richard J. Lipton. 1980 · 1980
Earlier work this paper cites.
Average case complete problems
Leonid A. Levin. 1986 · 1986
Earlier work this paper cites.
On the computational power of neural nets
Hava T. Siegelmann and Eduardo D. Sontag. 1992 · 1992
Earlier work this paper cites.
On the hardness of approximate reasoning
Dan Roth. 1996 · 1996
Earlier work this paper cites.
Estimation of probabilistic context-free grammars
Zhiyi Chi and Stuart Geman. 1998 · 1998
Earlier work this paper cites.
Approximate nearest neighbors: Towards removing the curse of dimensionality
Piotr Indyk and Rajeev Motwani. 1998 · 1998
Earlier work this paper cites.
Conditional random fields: Probabilistic models for segmenting and labeling sequence data
John D. Lafferty, Andrew McCallum, and Fernando C. N. Pereira. 2001 · 2001
Earlier work this paper cites.
Whole-sentence exponential language models: A vehicle for linguistic-statistical integration
Ronald Rosenfeld, Stanley Chen, and Xiaojin Zhu. 2001 · 2001
Earlier work this paper cites.
Estimation of non-normalized statistical models by score matching
Aapo Hyvärinen. 2005 · 2005
Earlier work this paper cites.
Average-case complexity
Andrej Bogdanov and Luca Trevisan. 2006 · 2006
Earlier work this paper cites.
A simple proof that optimality theory is computationally intractable
W. Idsardi. 2006 · 2006
Earlier work this paper cites.
A tutorial on energy-based learning
Yann LeCun, Sumit Chopra, Raia Hadsell, Marc’Aurelio Ranzato, and Fu-Jie Huang. 2006 · 2006
Earlier work this paper cites.
Complexity of inference in graphical models
Venkat Chandrasekaran, Nathan Srebro, and Prahladh Harsha. 2008 · 2008
Earlier work this paper cites.
The complexity of phrase alignment problems
John DeNero and D. Klein. 2008 · 2008
Earlier work this paper cites.
Generic complexity of undecidable problems
Alexei G. Myasnikov and Alexander N. Rybalov. 2008 · 2008
Cited alongside, same era.
Refining generative language models using discriminative learning
Ben Sandbank. 2008 · 2008
Cited alongside, same era.
Computational Complexity: a Modern Approach
Sanjeev Arora and Boaz Barak. 2009 · 2009
Cited alongside, same era.
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
Michael U. Gutmann and Aapo Hyvärinen. 2010 · 2010
Cited alongside, same era.
RNNLM—Recurrent neural network language modeling toolkit
Tomas Mikolov, Stefan Kombrink, Anoop Deoras, Lukas Burget, and Jan Honza Cernocky. 2011 · 2011
Cited alongside, same era.
Noise-contrastive estimation of unnormalized statistical models, with applications to natural image statistics
Michael U. Gutmann and Aapo Hyvärinen. 2012 · 2012
Recurrent neural networks as weighted language recognizers
Yining Chen, Sorcha Gilroy, Andreas Maletti, Jonathan May, and Kevin Knight. 2018 · 2018
Later among the works it cites.
Whole sentence neural language models
Y. Huang, A. Sethy, K. Audhkhasi, and B. Ramabhadran. 2018 · 2018
Later among the works it cites.
Noise contrastive estimation and negative sampling for conditional models: Consistency and statistical efficiency
Zhuang Ma and Michael Collins. 2018 · 2018
Later among the works it cites.
Hard non-monotonic attention for character-level transduction
Shijie Wu, Pamela Shapiro, and Ryan Cotterell. 2018 · 2018
Later among the works it cites.
Neural finite-state transducers: Beyond rational relations
Chu-Cheng Lin, Hao Zhu, Matthew R. Gormley, and Jason Eisner. 2019 · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Smatch: an evaluation metric for semantic feature structures
Shu Cai and Kevin Knight. 2013 · 2013
Cited alongside, same era.
Auto-encoding variational Bayes
Diederik P. Kingma and Max Welling. 2014 · 2014
Cited alongside, same era.
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V. Le. 2014 · 2014
Cited alongside, same era.
Adaptive computation time for recurrent neural networks
A. Graves. 2016 · 2016
Cited alongside, same era.
WaveNet: A generative model for raw audio
Aaron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu. 2016 · 2016
Cited alongside, same era.
Weighting finite-state transductions with neural context
Pushpendre Rastogi, Ryan Cotterell, and Jason Eisner. 2016 · 2016
Cited alongside, same era.
Later among the works it cites.
Defending against neural fake news
Rowan Zellers, Ari Holtzman, Hannah Rashkin, Yonatan Bisk, Ali Farhadi, Franziska Roesner, and Yejin Choi. 2019 · 2019
Later among the works it cites.
Re: Teaching gpt-3 to identify nonsense
blixt. 2020 · 2020
Closest in time.
Language models are few-shot learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 2020
Closest in time.
Measuring systematic generalization in neural proof generation with transformers
Nicolas Gontier, Koustuv Sinha, Siva Reddy, and C. Pal. 2020 · 2020
Closest in time.
Generalization through memorization: Nearest neighbor language models
Urvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer, and Mike Lewis. 2020 · 2020
Closest in time.
My Hobby: Embedding NP \mathrm{NP} -Complete Problems in Restaurant Orders
Randall Munroe. 2009 · 2020
Closest in time.
Unsupervised commonsense question answering with self-talk
Vered Shwartz, Peter West, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2020 · 2020
Closest in time.
Consistency of a recurrent language model with respect to incomplete decoding
Sean Welleck, Ilia Kulikov, Jaedeok Kim, Richard Yuanzhe Pang, and Kyunghyun Cho. 2020 · 2020
Closest in time.
Residual energy-based models for text generation
Anton Bakhtin, Yuntian Deng, Sam Gross, Myle Ott, Marc’Aurelio Ranzato, and Arthur Szlam. 2021 · 2021
Closest in time.
On the dangers of stochastic parrots: Can language models be too big?
Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021 · 2021
Closest in time.
Adaptive semiparametric language models
Dani Yogatama, Cyprien de Masson d’Autume, and Lingpeng Kong. 2021 · 2021
Closest in time.