Fetching the paper…
Reading the bibliography…
Gating mechanisms are widely used in neural network models, where they allow gradients to backpropagate more easily through depth or time.
Untersuchungen zu dynamischen neuronalen netzen
Hochreiter, S · 1991
Earlier work this paper cites.
Learning long-term dependencies with gradient descent is difficult
Bengio, Y., Simard, P., Frasconi, P., et al · 1994
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Gradient flow in recurrent nets: the difficulty of learning long-term dependencies, 2001
Hochreiter, S., Bengio, Y., Frasconi, P., Schmidhuber, J., et al · 2001
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Bahdanau, D., Cho, K., and Bengio, Y · 2014
Earlier work this paper cites.
Learning phrase representations using RNN encoder-decoder for statistical machine translation
Cho, K., Van Merriënboer, B., Gulcehre, C., Bahdanau, D., Bougares, F., Schwenk, H., and Bengio, Y · 2014
Earlier work this paper cites.
Empirical evaluation of gated recurrent neural networks on sequence modeling
Chung, J., Gulcehre, C., Cho, K., and Bengio, Y · 2014
Earlier work this paper cites.
Graves, A., Wayne, G., and Danihelka, I · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Koutnik, J., Greff, K., Gomez, F., and Schmidhuber, J · 2014
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R · 2014
Earlier work this paper cites.
Weston, J., Chopra, S., and Bordes, A · 2014
Earlier work this paper cites.
Zaremba, W. and Sutskever, I · 2014
Earlier work this paper cites.
An empirical exploration of recurrent network architectures
Jozefowicz, R., Zaremba, W., and Sutskever, I · 2015
Earlier work this paper cites.
A simple way to initialize recurrent networks of rectified linear units
Le, Q. V., Jaitly, N., and Hinton, G. E · 2015
Earlier work this paper cites.
Unitary evolution recurrent neural networks
Arjovsky, M., Shah, A., and Bengio, Y · 2016
Cited alongside, same era.
Hierarchical multiscale recurrent neural networks
Chung, J., Ahn, S., and Bengio, Y · 2016
Cited alongside, same era.
LSTM: A search space odyssey
Greff, K., Srivastava, R. K., Koutník, J., Steunebrink, B. R., and Schmidhuber, J · 2016
Cited alongside, same era.
Noisy activation functions
Gulcehre, C., Moczulski, M., Denil, M., and Bengio, Y · 2016
Cited alongside, same era.
Recurrent orthogonal networks and long-memory tasks
Henaff, M., Szlam, A., and LeCun, Y · 2016
Cited alongside, same era.
Impala: Scalable distributed deep-rl with importance weighted actor-learner architectures
Espeholt, L., Soyer, H., Munos, R., Simonyan, K., Mnih, V., Ward, T., Doron, Y., Firoiu, V., Harley, T., Dunning, I., et al · 2018
Later among the works it cites.
Optimizing agent behavior over long time scales by transporting value
Hung, C.-C., Lillicrap, T., Abramson, J., Wu, Y., Mirza, M., Carnevale, F., Ahuja, A., and Wayne, G · 2018
Later among the works it cites.
Recurrent experience replay in distributed reinforcement learning
Kapturowski, S., Ostrovski, G., Quan, J., Munos, R., and Dabney, W · 2018
Later among the works it cites.
The NarrativeQA reading comprehension challenge
Kočiskỳ, T., Schwarz, J., Blunsom, P., Dyer, C., Hermann, K. M., Melis, G., and Grefenstette, E · 2018
Later among the works it cites.
Fast parametric learning with activation memorization
Rae, J. W., Dyer, C., Dayan, P., and Lillicrap, T. P · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jang, E., Gu, S., and Poole, B · 2016
Cited alongside, same era.
Zoneout: Regularizing RNNs by randomly preserving hidden activations
Krueger, D., Maharaj, T., Kramár, J., Pezeshki, M., Ballas, N., Ke, N. R., Goyal, A., Bengio, Y., Courville, A., and Pal, C · 2016
Cited alongside, same era.
The concrete distribution: A continuous relaxation of discrete random variables
Maddison, C. J., Mnih, A., and Teh, Y. W · 2016
Cited alongside, same era.
Pointer sentinel mixture models
Merity, S., Xiong, C., Bradbury, J., and Socher, R · 2016
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
Mnih, V., Badia, A. P., Mirza, M., Graves, A., Lillicrap, T., Harley, T., Silver, D., and Kavukcuoglu, K · 2016
Cited alongside, same era.
Control of memory, active perception, and action in Minecraft
Oh, J., Chockalingam, V., Singh, S., and Lee, H · 2016
Cited alongside, same era.
Dilated recurrent neural networks
Chang, S., Zhang, Y., Han, W., Yu, M., Guo, X., Tan, W., Cui, X., Witbrock, M., Hasegawa-Johnson, M. A., and Huang, T. S · 2017
Cited alongside, same era.
Later among the works it cites.
Relational recurrent neural networks
Santoro, A., Faulkner, R., Raposo, D., Rae, J., Chrzanowski, M., Weber, T., Wierstra, D., Vinyals, O., Pascanu, R., and Lillicrap, T · 2018
Later among the works it cites.
Ordered neurons: Integrating tree structures into recurrent neural networks
Shen, Y., Tan, S., Sordoni, A., and Courville, A · 2018
Later among the works it cites.
Can recurrent neural networks warp time?
Tallec, C. and Ollivier, Y · 2018
Later among the works it cites.
Learning longer-term dependencies in RNNs with auxiliary losses
Trinh, T. H., Dai, A. M., Luong, M.-T., and Le, Q. V · 2018
Later among the works it cites.
The unreasonable effectiveness of the forget gate
van der Westhuizen, J. and Lasenby, J · 2018
Later among the works it cites.
Unsupervised predictive memory in a goal-directed agent
Wayne, G., Hung, C.-C., Amos, D., Mirza, M., Ahuja, A., Grabska-Barwinska, A., Rae, J., Mirowski, P., Leibo, J. Z., Santoro, A., et al · 2018
Later among the works it cites.
Towards non-saturating recurrent units for modelling long-term dependencies
Chandar, S., Sankar, C., Vorontsov, E., Kahou, S. E., and Bengio, Y · 2019
Closest in time.
Transformer-xl: Attentive language models beyond a fixed-length context
Dai, Z., Yang, Z., Yang, Y., Cohen, W. W., Carbonell, J., Le, Q. V., and Salakhutdinov, R · 2019
Closest in time.
Single headed attention rnn: Stop thinking with your head
Merity, S · 2019
Closest in time.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I · 2019
Closest in time.