Fetching the paper…
Reading the bibliography…
To predict the next token, autoregressive models ordinarily examine the past.
A machine program for theorem-proving
Davis, M., Logemann, G., and Loveland, D · 1962
Earlier work this paper cites.
Logic and conversation
Grice, H. P · 1975
Earlier work this paper cites.
Information processing in dynamical systems: Foundations of harmony theory
Smolensky, P · 1986
Earlier work this paper cites.
On constraints and repair strategies
Paradis, C · 1988
Earlier work this paper cites.
Where the really hard problems are
Cheeseman, P., Kanefsky, B., and Taylor, W. M · 1991
Earlier work this paper cites.
Phonology in Generative Grammar
Kenstowicz, M · 1993
Earlier work this paper cites.
Optimality Theory: Constraint Interaction in Generative Grammar
Prince, A. and Smolensky, P · 1993
Earlier work this paper cites.
Analytic and algorithmic solution of random satisfiability problems
Mézard, M., Parisi, G., and Zecchina, R · 2002
Earlier work this paper cites.
A tutorial on energy-based learning
LeCun, Y., Chopra, S., Hadsell, R., and Huang, F. J · 2007
Earlier work this paper cites.
Structure compilation: Trading structure for features
Liang, P., Daumé III, H., and Klein, D · 2008
Earlier work this paper cites.
Thinking, Fast and Slow
Kahneman, D · 2011
Earlier work this paper cites.
A survey of Monte Carlo tree search methods
Browne, C., Powley, E. J., Whitehouse, D., Lucas, S. M., Cowling, P. I., Rohlfshagen, P., Tavener, S., Liebana, D. P., Samothrakis, S., and Colton, S · 2012
Cited alongside, same era.
Computational rationality: Linking mechanism and behavior through utility maximization
Lewis, R. L., Howes, A., and Singh, S · 2014
Cited alongside, same era.
Dropout: A simple way to prevent neural networks from overfitting
Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, D. and Ba, J · 2015
Cited alongside, same era.
Training data augmentation for low-resource morphological inflection
Bergmanis, T., Kann, K., Schütze, H., and Goldwater, S · 2017
Cited alongside, same era.
Stochastic beams and where to find them: The Gumbel-top-k trick for sampling sequences without replacement
Kool, W., Van Hoof, H., and Welling, M · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I · 2019
Later among the works it cites.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Later among the works it cites.
Language modeling through inverse reinforcement learning , 2020
Mehta, R., Winston, C., and Michael, P · 2020
Later among the works it cites.
Learning retrospective knowledge with reverse reinforcement learning
Zhang, S., Veeriah, V., and Whiteson, S · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
CoNLL-SIGMORPHON 2017 shared task: Universal morphological reinflection in 52 languages
Cotterell, R., Kirov, C., Sylak-Glassman, J., Walther, G., Vylomova, E., Xia, P., Faruqui, M., Kübler, S., Yarowsky, D., Eisner, J., and Hulden, M · 2017
Cited alongside, same era.
Hafez: An interactive poetry generation system
Ghazvininejad, M., Shi, X., Priyadarshi, J., and Knight, K · 2017
Cited alongside, same era.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Cited alongside, same era.
Neural particle smoothing for sampling from conditional sequence models
Lin, C.-C. and Eisner, J · 2018
Cited alongside, same era.
Toward diverse text generation with inverse reinforcement learning , 2018
Shi, Z., Chen, X., Qiu, X., and Huang, X · 2018
Cited alongside, same era.
Lin, C.-C., Jaech, A., Li, X., Gormley, M. R., and Eisner, J · 2021
Later among the works it cites.
FUDGE: Controlled text generation with future discriminators
Yang, K. and Klein, D · 2021
Later among the works it cites.
PPL-MCTS: Constrained textual generation through discriminator-guided MCTS decoding
Chaffin, A., Claveau, V., and Kijak, E · 2022
Later among the works it cites.
Proof of the satisfiability conjecture for large k k
Ding, J., Sly, A., and Sun, N · 2022
Later among the works it cites.
Standing on the shoulders of giant frozen language models
Levine, Y., Dalmedigos, I., Ram, O., Zeldes, Y., Jannai, D., Muhlgay, D., Osin, Y., Lieber, O., Lenz, B., Shalev-Shwartz, S., Shashua, A., Leyton-Brown, K., and Shoham, Y · 2022
Later among the works it cites.
A survey of controllable text generation using Transformer-based pre-trained language models
Zhang, H., Song, H., Li, S., Zhou, M., and Song, D · 2022
Later among the works it cites.