Fetching the paper…
Reading the bibliography…
Recently, a new recurrent neural network (RNN) named the Legendre Memory Unit (LMU) was proposed and shown to achieve state-of-the-art performance on several benchmark datasets.
Modern control theory
Brogan, W. L · 1991
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Some frequency-domain approaches to the model reduction of delay systems
Partington, J. R · 2004
Earlier work this paper cites.
Feedback systems: an introduction for scientists and engineers
Åström, K. J. and Murray, R. M · 2010
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Maas, A., Daly, R. E., Pham, P. T., Huang, D., Ng, A. Y., and Potts, C · 2011
Earlier work this paper cites.
Subword language modeling with neural networks
Mikolov, T., Sutskever, I., Deoras, A., Le, H.-S., Kombrink, S., and Cernocky, J · 2012
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Pennington, J., Socher, R., and Manning, C. D · 2014
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Bowman, S. R., Angeli, G., Potts, C., and Manning, C. D · 2015
Earlier work this paper cites.
A simple way to initialize recurrent networks of rectified linear units
Le, Q. V., Jaitly, N., and Hinton, G. E · 2015
Earlier work this paper cites.
Stanford neural machine translation systems for spoken language domains
Luong, M.-T. and Manning, C. D · 2015
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Sennrich, R., Haddow, B., and Birch, A · 2015
Earlier work this paper cites.
Training very deep networks
Srivastava, R. K., Greff, K., and Schmidhuber, J · 2015
Cited alongside, same era.
Strongly-typed recurrent neural networks
Balduzzi, D. and Ghifary, M · 2016
Cited alongside, same era.
Architectural complexity measures of recurrent neural networks
Zhang, S., Wu, Y., Che, T., Lin, Z., Memisevic, R., Salakhutdinov, R. R., and Bengio, Y · 2016
Cited alongside, same era.
Parallelizing linear recurrent neural nets over sequence length
Martin, E. and Cundy, C · 2017
Cited alongside, same era.
Learning to generate reviews and discovering sentiment
Radford, A., Jozefowicz, R., and Sutskever, I · 2017
Cited alongside, same era.
The Design and Implementation of Low-Latency Prediction Serving Systems
Crankshaw, D · 2019
Later among the works it cites.
Justifying recommendations using distantly-labeled reviews and fine-grained aspects
Ni, J., Li, J., and McAuley, J · 2019
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2019
Later among the works it cites.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Sanh, V., Debut, L., Chaumond, J., and Wolf, T · 2019
Later among the works it cites.
Natural language understanding with the quora question pairs dataset
Sharma, L., Graesser, L., Nangia, N., and Evci, U · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Cited alongside, same era.
Deep contextualized word representations
Peters, M. E., Neumann, M., Iyyer, M., Gardner, M., Clark, C., Lee, K., and Zettlemoyer, L · 2018
Cited alongside, same era.
Improving language understanding by generative pre-training
Radford, A., Narasimhan, K., Salimans, T., and Sutskever, I · 2018
Cited alongside, same era.
Baseline needs more love: On simple word-embedding-based models and associated pooling mechanisms
Shen, D., Wang, G., Wang, W., Min, M. R., Su, Q., Zhang, Y., Li, C., Henao, R., and Carin, L · 2018
Cited alongside, same era.
Improving spiking dynamical networks: Accurate delays, higher-order synapses, and time cells
Voelker, A. R. and Eliasmith, C · 2018
Cited alongside, same era.
Towards non-saturating recurrent units for modelling long-term dependencies
Chandar, S., Sankar, C., Vorontsov, E., Kahou, S. E., and Bengio, Y · 2019
Cited alongside, same era.
Later among the works it cites.
Legendre memory units: Continuous-time representation in recurrent neural networks
Voelker, A., Kajić, I., and Eliasmith, C · 2019
Later among the works it cites.
Hardware aware training for efficient keyword spotting on general purpose and specialized hardware
Blouw, P., Malik, G., Morcos, B., Voelker, A. R., and Eliasmith, C · 2020
Later among the works it cites.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Later among the works it cites.
Hippo: Recurrent memory with optimal polynomial projections
Gu, A., Dao, T., Ermon, S., Rudra, A., and Re, C · 2020
Later among the works it cites.
Transformer-based online ctc/attention end-to-end speech recognition architecture
Miao, H., Cheng, G., Gao, C., Zhang, P., and Yan, Y · 2020
Later among the works it cites.
Reducing the latency of end-to-end streaming speech recognition models with a scout network
Wang, C., Wu, Y., Liu, S., Li, J., Lu, L., Ye, G., and Zhou, M · 2020
Later among the works it cites.
Transformer transducer: A streamable speech recognition model with transformer encoders and rnn-t loss
Zhang, Q., Lu, H., Sak, H., Tripathi, A., McDermott, E., Koo, S., and Kumar, S · 2020
Later among the works it cites.