Fetching the paper…
Reading the bibliography…
Developing NLP models traditionally involves two stages - training and application.
I.—COMPUTING MACHINERY AND INTELLIGENCE
Turing, A. M · 1950
Earlier work this paper cites.
Procedures as a representation for data in a computer program for understanding natural language
Winograd, T · 1971
Earlier work this paper cites.
Backpropagation through time: what it does and how to do it
Werbos, P. J · 1990
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Gradient-based learning applied to document recognition
LeCun, Y., Bottou, L., Bengio, Y., and Haffner, P · 1998
Earlier work this paper cites.
Game of life cellular automata , volume 1
Adamatzky, A · 2010
Earlier work this paper cites.
Parallel and serial grouping of image elements in visual perception
Houtkamp, R. and Roelfsema, P · 2010
Earlier work this paper cites.
Recurrent neural network based language model
Mikolov, T., Karafiát, M., Burget, L., Černockỳ, J., and Khudanpur, S · 2010
Earlier work this paper cites.
The Feynman lectures on physics, Vol. I: The new millennium edition: mainly mechanics, radiation, and heat , volume 1
Feynman, R. P., Leighton, R. B., and Sands, M · 2011
Earlier work this paper cites.
A three-way model for collective learning on multi-relational data
Nickel, M., Tresp, V., and Kriegel, H.-P · 2011
Earlier work this paper cites.
Towards ai-complete question answering: A set of prerequisite toy tasks
Weston, J., Bordes, A., Chopra, S., Rush, A. M., van Merriënboer, B., Joulin, A., and Mikolov, T · 2015
Cited alongside, same era.
Tracking the world state with recurrent entity networks
Henaff, M., Weston, J., Szlam, A., Bordes, A., and LeCun, Y · 2016
Cited alongside, same era.
Learning to understand phrases by embedding the dictionary
Hill, F., Cho, K., Korhonen, A., and Bengio, Y · 2016
Cited alongside, same era.
Conditional image generation with pixelcnn decoders, 2016
van den Oord, A., Kalchbrenner, N., Vinyals, O., Espeholt, L., Graves, A., and Kavukcuoglu, K · 2016
Cited alongside, same era.
Learning to compute word embeddings on the fly
Bahdanau, D., Bosc, T., Jastrzębski, S., Grefenstette, E., Vincent, P., and Bengio, Y · 2017
Reviving and improving recurrent back-propagation
Liao, R., Xiong, Y., Fetaya, E., Zhang, L., Yoon, K., Pitkow, X., Urtasun, R., and Zemel, R · 2018
Later among the works it cites.
Transformer-xl: Attentive language models beyond a fixed-length context
Dai, Z., Yang, Z., Yang, Y., Carbonell, J., Le, Q. V., and Salakhutdinov, R · 2019
Later among the works it cites.
Learning long-range spatial dependencies with horizontal gated-recurrent units, 2019
Linsley, D., Kim, J., Veerabadran, V., and Serre, T · 2019
Later among the works it cites.
Zero-shot entity linking by reading entity descriptions
Logeswaran, L., Chang, M.-W., Lee, K., Toutanova, K., Devlin, J., and Lee, H · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2017
Cited alongside, same era.
Unbiasing truncated backpropagation through time
Tallec, C. and Ollivier, Y · 2017
Cited alongside, same era.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Cited alongside, same era.
Still not systematic after all these years: On the compositional skills of sequence-to-sequence recurrent networks
Lake, B. and Baroni, M · 2018
Cited alongside, same era.
Rethinking attention with performers
Choromanski, K., Likhosherstov, V., Dohan, D., Song, X., Gane, A., Sarlos, T., Hawkins, P., Davis, J., Mohiuddin, A., Kaiser, L., et al · 2020
Later among the works it cites.
Zero-shot learning and its applications from autonomous vehicles to covid-19 diagnosis: A review
Rezaei, M. and Shahidi, M · 2020
Later among the works it cites.
Teaching pre-trained models to systematically reason over implicit knowledge
Talmor, A., Tafjord, O., Clark, P., Goldberg, Y., and Berant, J · 2020
Later among the works it cites.
Long range arena: A benchmark for efficient transformers
Tay, Y., Dehghani, M., Abnar, S., Shen, Y., Bahri, D., Pham, P., Rao, J., Yang, L., Ruder, S., and Metzler, D · 2020
Later among the works it cites.
Linformer: Self-attention with linear complexity
Wang, S., Li, B., Khabsa, M., Fang, H., and Ma, H · 2020
Later among the works it cites.