Very deep convolutional networks for natural language processing
Original
Conneau, A., Schwenk, H., Barrault, L., and LeCun, Y · 2016
Later among the works it cites.
Memory-efficient backpropagation through time
Gruslys, A., Munos, R., Danihelka, I., Lanctot, M., and Graves, A · 2016
Later among the works it cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Later among the works it cites.
Recurrent orthogonal networks and long-memory tasks
Henaff, M., Szlam, A., and LeCun, Y · 2016
Later among the works it cites.
On multiplicative integration with recurrent neural networks
Wu, Y., Zhang, S., Zhang, Y., Bengio, Y., and Salakhutdinov, R. R · 2016
Later among the works it cites.
Efficient character-level document classification by combining convolution and recurrent layers
Original
Xiao, Y. and Cho, K · 2016
Later among the works it cites.
Z-forcing: Training stochastic recurrent networks
GOYAL, A. G. A. P., Sordoni, A., Côté, M.-A., Ke, N., and Bengio, Y · 2017
Later among the works it cites.
Decoupled neural interfaces using synthetic gradients
Jaderberg, M., Czarnecki, W. M., Osindero, S., Vinyals, O., Graves, A., and Kavukcuoglu, K · 2017
Later among the works it cites.
Unsupervised pretraining for sequence to sequence learning
Ramachandran, P., Liu, P. J., and Le, Q. V · 2017
Later among the works it cites.
PixelCNN++: Improving the PixelCNN with discretized logistic mixture likelihood and other modifications
Salimans, T., Karpathy, A., Chen, X., and Kingma, D. P · 2017
Later among the works it cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Later among the works it cites.
Learning to skim text
Yu, A. W., Lee, H., and Le, Q. V · 2017
Later among the works it cites.
Recurrent highway networks
Zilly, J. G., Srivastava, R. K., Koutník, J., and Schmidhuber, J · 2017
Later among the works it cites.
Skip RNN: Learning to skip state updates in recurrent neural networks
Campos, V., Jou, B., Giró-i Nieto, X., Torres, J., and Chang, S.-F · 2018
Closest in time.
Generating Wikipedia by summarizing long sequences
Liu, P. J., Saleh, M., Pot, E., Goodrich, B., Sepassi, R., Kaiser, L., and Shazeer, N · 2018
Closest in time.
Neural speed reading via skim-rnn
Seo, M., Min, S., Farhadi, A., and Hajishirzi, H · 2018
Closest in time.