Fetching the paper…
Reading the bibliography…
Transformer language models have made tremendous strides in natural language understanding tasks.
HuggingFace’s Transformers: State-of-the-art Natural Language Processing
Wolf, T.; Debut, L.; Sanh, V.; Chaumond, J.; Delangue, C.; Moi, A.; Cistac, P.; Rault, T.; Louf, R.; Funtowicz, M.; Davison, J.; Shleifer, S.; von Platen, P.; Ma, C.; Jernite, Y.; Plu, J.; Xu, C.; Scao, T. L.; Gugger, S.; Drame, M.; Lhoest, Q.; and Rush, A. M. 2019 · 1910
Earlier work this paper cites.
Long short-term memory
Hochreiter, S.; and Schmidhuber, J. 1997 · 1997
Earlier work this paper cites.
Scaling Laws for Neural Language Models
Kaplan, J.; McCandlish, S.; Henighan, T.; Brown, T. B.; Chess, B.; Child, R.; Gray, S.; Radford, A.; Wu, J.; and Amodei, D. 2020 · 2001
Earlier work this paper cites.
Language Models are Few-Shot Learners
Brown, T. B.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; Agarwal, S.; Herbert-Voss, A.; Krueger, G.; Henighan, T.; Child, R.; Ramesh, A.; Ziegler, D. M.; Wu, J.; Winter, C.; Hesse, C.; Chen, M.; Sigler, E.; Litwin, M.; Gray, S.; Chess, B.; Clark, J.; Berner, C.; McCandlish, S.; Radford, A.; Sutskever, I.; and Amodei, D. 2020 · 2005
Earlier work this paper cites.
Introduction to Information Retrieval
Manning, C. D.; Raghavan, P.; and Schütze, H. 2008 · 2008
Earlier work this paper cites.
The Chess Transformer: Mastering Play using Generative Language Models
Noever, D.; Ciolino, M.; and Kalin, J. 2020 · 2008
Earlier work this paper cites.
python-chess: a chess library for Python
Fiekas, N. 2012 · 2012
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P.; and Ba, J. 2014 · 2014
Earlier work this paper cites.
Teaching Machines to Read and Comprehend
Hermann, K. M.; Kočiský, T.; Grefenstette, E.; Espeholt, L.; Kay, W.; Suleyman, M.; and Blunsom, P. 2015 · 2015
Earlier work this paper cites.
Predicting Moves in Chess using Convolutional Neural Networks
Oshri, B.; and Khandwala, N. 2015 · 2015
Earlier work this paper cites.
Towards AI-Complete Question Answering: A Set of Prerequisite Toy Tasks
Weston, J.; Bordes, A.; Chopra, S.; Rush, A. M.; van Merriënboer, B.; Joulin, A.; and Mikolov, T. 2015 · 2015
Earlier work this paper cites.
DeepChess: End-to-End Deep Neural Network for Automatic Learning in Chess
David, E.; Netanyahu, N. S.; and Wolf, L. 2016 · 2016
Earlier work this paper cites.
Probing for semantic evidence of composition by means of simple classification tasks
Ettinger, A.; Elgohary, A.; and Resnik, P. 2016 · 2016
Earlier work this paper cites.
The Goldilocks Principle: Reading Children’s Books with Explicit Memory Representations
Hill, F.; Bordes, A.; Chopra, S.; and Weston, J. 2016 · 2016
Cited alongside, same era.
A Corpus and Cloze Evaluation for Deeper Understanding of Commonsense Stories
Mostafazadeh, N.; Chambers, N.; He, X.; Parikh, D.; Batra, D.; Vanderwende, L.; Kohli, P.; and Allen, J. 2016 · 2016
Cited alongside, same era.
The LAMBADA dataset: Word prediction requiring a broad discourse context
Paperno, D.; Kruszewski, G.; Lazaridou, A.; Pham, N. Q.; Bernardi, R.; Pezzelle, S.; Baroni, M.; Boleda, G.; and Fernández, R. 2016 · 2016
Cited alongside, same era.
Fine-grained Analysis of Sentence Embeddings Using Auxiliary Prediction Tasks
Adi, Y.; Kermany, E.; Belinkov, Y.; Lavi, O.; and Goldberg, Y. 2017 · 2017
Cited alongside, same era.
Understanding intermediate layers using linear classifier probes
Alain, G.; and Bengio, Y. 2017 · 2017
Cited alongside, same era.
PyTorch Lightning
Falcon et al., W. 2019 · 2019
Later among the works it cites.
Designing and Interpreting Probes with Control Tasks
Hewitt, J.; and Liang, P. 2019 · 2019
Later among the works it cites.
PyTorch: An Imperative Style, High-Performance Deep Learning Library
Paszke, A.; Gross, S.; Massa, F.; Lerer, A.; Bradbury, J.; Chanan, G.; Killeen, T.; Lin, Z.; Gimelshein, N.; Antiga, L.; Desmaison, A.; Kopf, A.; Yang, E.; DeVito, Z.; Raison, M.; Tejani, A.; Chilamkurthy, S.; Steiner, B.; Fang, L.; Bai, J.; and Chintala, S. 2019 · 2019
Later among the works it cites.
Language Models as Knowledge Bases?
Petroni, F.; Rocktäschel, T.; Riedel, S.; Lewis, P.; Bakhtin, A.; Wu, Y.; and Miller, A. 2019 · 2019
Later among the works it cites.
Language Models are Unsupervised Multitask Learners
Radford, A.; Wu, J.; Child, R.; Luan, D.; Amodei, D.; and Sutskever, I. 2019 · 2019
Later among the works it cites.
What do you learn from context? Probing for sentence structure in contextualized word representations
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hermann, K. M.; Hill, F.; Green, S.; Wang, F.; Faulkner, R.; Soyer, H.; Szepesvari, D.; Czarnecki, W. M.; Jaderberg, M.; Teplyashin, D.; Wainwright, M.; Apps, C.; Hassabis, D.; and Blunsom, P. 2017 · 2017
Cited alongside, same era.
Understanding Grounded Language Learning Agents
Hill, F.; Hermann, K. M.; Blunsom, P.; and Clark, S. 2017 · 2017
Cited alongside, same era.
LSDSem 2017 Shared Task: The Story Cloze Test
Mostafazadeh, N.; Roth, M.; Louis, A.; Chambers, N.; and Allen, J. 2017 · 2017
Cited alongside, same era.
Attention is All you Need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, L. u.; and Polosukhin, I. 2017 · 2017
Cited alongside, same era.
TextWorld: A Learning Environment for Text-based Games
Côté, M.-A.; Kádár, A.; Yuan, X.; Kybartas, B.; Barnes, T.; Fine, E.; Moore, J.; Tao, R. Y.; Hausknecht, M.; Asri, L. E.; Adada, M.; Tay, W.; and Trischler, A. 2018 · 2018
Cited alongside, same era.
Mixed Precision Training
Micikevicius, P.; Narang, S.; Alben, J.; Diamos, G.; Elsen, E.; Garcia, D.; Ginsburg, B.; Houston, M.; Kuchaiev, O.; Venkatesh, G.; and Wu, H. 2018 · 2018
Cited alongside, same era.
A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play
Silver, D.; Hubert, T.; Schrittwieser, J.; Antonoglou, I.; Lai, M.; Guez, A.; Lanctot, M.; Sifre, L.; Kumaran, D.; Graepel, T.; Lillicrap, T.; Simonyan, K.; and Hassabis, D. 2018 · 2018
Cited alongside, same era.
Tenney, I.; Xia, P.; Chen, B.; Wang, A.; Poliak, A.; McCoy, R. T.; Kim, N.; Durme, B. V.; Bowman, S. R.; Das, D.; and Pavlick, E. 2019 · 2019
Later among the works it cites.
Transformers Play Chess
Cheng, R. 2020 · 2020
Later among the works it cites.
What BERT is Not: Lessons from a New Suite of Psycholinguistic Diagnostics for Language Models
Ettinger, A. 2020 · 2020
Later among the works it cites.
Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention
Katharopoulos, A.; Vyas, A.; Pappas, N.; and Fleuret, F. 2020 · 2020
Later among the works it cites.
Reformer: The Efficient Transformer
Kitaev, N.; Kaiser, L.; and Levskaya, A. 2020 · 2020
Later among the works it cites.
Information-Theoretic Probing for Linguistic Structure
Pimentel, T.; Valvoda, J.; Hall Maudslay, R.; Zmigrod, R.; Williams, A.; and Cotterell, R. 2020 · 2020
Later among the works it cites.
A Very Unlikely Chess Game
Presser, S.; and Branwen, G. 2020 · 2020
Later among the works it cites.
Rethinking Attention with Performers
Choromanski, K. M.; Likhosherstov, V.; Dohan, D.; Song, X.; Gane, A.; Sarlos, T.; Hawkins, P.; Davis, J. Q.; Mohiuddin, A.; Kaiser, L.; Belanger, D. B.; Colwell, L. J.; and Weller, A. 2021 · 2021
Closest in time.