Fetching the paper…
Reading the bibliography…
Recurrent neural networks (RNNs) are known to be difficult to train due to the gradient vanishing and exploding problems and thus difficult to learn long-term patterns and construct deep networks.
B. Parlett, “Laguerre’s method applied to the matrix eigenvalue problem,” Mathematics of Computation , vol. 18, no. 87, pp. 464–485, 1964
1964
Earlier work this paper cites.
S. Hochreiter, “Untersuchungen zu dynamischen neuronalen netzen,” 1991
1991
Earlier work this paper cites.
M. I. Jordan, “Serial order: A parallel distributed processing approach,” Advances in psychology , vol. 121, pp. 471–495, 1997
1997
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural computation , vol. 9, no. 8, pp. 1735–1780, 1997
1997
Earlier work this paper cites.
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner, “Gradient-based learning applied to document recognition,” Proceedings of the IEEE , vol. 86, no. 11, pp. 2278–2324, 1998
1998
Earlier work this paper cites.
A. Graves, S. Fernández, and J. Schmidhuber, “Multi-dimensional recurrent neural networks,” in Proceedings of the 17th International Conference on Artificial Neural Networks . Springer-Verlag, 2007, pp. 549–558
2007
Earlier work this paper cites.
V. Nair and G. E. Hinton, “Rectified linear units improve restricted boltzmann machines,” in ICML , 2010
2010
Earlier work this paper cites.
I. Sutskever, J. Martens, and G. E. Hinton, “Generating text with recurrent neural networks,” in Proceedings of the 28th International Conference on Machine Learning , 2011, pp. 1017–1024
2011
Earlier work this paper cites.
T. Mikolov, I. Sutskever, A. Deoras, H.-S. Le, S. Kombrink, and J. Cernocky, “Subword language modeling with neural networks,” preprint (http://www. fit. vutbr. cz/imikolov/rnnlm/char. pdf) , 2012
2012
Earlier work this paper cites.
T. Mikolov and G. Zweig, “Context dependent recurrent neural network language model,” IEEE Spoken Language Technology Workshop (SLT) , vol. 12, pp. 234–239, 2012
2012
Earlier work this paper cites.
R. Pascanu, C. Gulcehre, K. Cho, and Y. Bengio, “How to construct deep recurrent neural networks,” Proceeding of the International Conference on Learning Representations , 2013
2013
Earlier work this paper cites.
R. Pascanu, T. Mikolov, and Y. Bengio, “On the difficulty of training recurrent neural networks,” in International Conference on Machine Learning , 2013, pp. 1310–1318
2013
Earlier work this paper cites.
K. Cho, B. Van Merriënboer, C. Gulcehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio, “Learning phrase representations using rnn encoder-decoder for statistical machine translation,” Proceedings of the Empirical Methods in Natural Language Processing , 2014
2014
Earlier work this paper cites.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” Proceeding of the International Conference on Learning Representations , 2014
2014
Earlier work this paper cites.
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
R. Jozefowicz, W. Zaremba, and I. Sutskever, “An empirical exploration of recurrent network architectures,” in Proceedings of the 32nd International Conference on Machine Learning , 2015, pp. 2342–2350
2015
Earlier work this paper cites.
M. Arjovsky, A. Shah, and Y. Bengio, “Unitary evolution recurrent neural networks,” Proceedings of the International Conference on Machine Learning , 2015
2015
Earlier work this paper cites.
N. Kalchbrenner, I. Danihelka, and A. Graves, “Grid long short-term memory,” Proceeding of the International Conference on Learning Representations , 2015
2015
Earlier work this paper cites.
X.-Y. Zhang, F. Yin, Y.-M. Zhang, C.-L. Liu, and Y. Bengio, “Drawing and recognizing chinese characters with recurrent neural network,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 40, pp. 849–862, 2016
2016
Earlier work this paper cites.
Q. Wu, C. Shen, P. Wang, A. Dick, and A. van den Hengel, “Image captioning and visual question answering based on attributes and external knowledge,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 40, pp. 1367–1381, 2016
2016
Earlier work this paper cites.
B. Neyshabur, Y. Wu, R. R. Salakhutdinov, and N. Srebro, “Path-normalized optimization of recurrent neural networks with relu activations,” in Advances in Neural Information Processing Systems , 2016, pp. 3477–3485
2016
Earlier work this paper cites.
D. Krueger and R. Memisevic, “Regularizing rnns by stabilizing activations,” in Proceeding of the International Conference on Learning Representations , 2016
2016
Earlier work this paper cites.
D. Krueger, T. Maharaj, J. Kramár, M. Pezeshki, N. Ballas, N. R. Ke, A. Goyal, Y. Bengio, H. Larochelle, A. Courville et al. , “Zoneout: Regularizing rnns by randomly preserving hidden activations,” Proceeding of the International Conference on Learning Representations , 2016
2016
Earlier work this paper cites.
A. Shahroudy, J. Liu, T.-T. Ng, and G. Wang, “Ntu rgb+ d: A large scale dataset for 3d human activity analysis,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2016, pp. 1010–1019
2016
Earlier work this paper cites.
G.-B. Zhou, J. Wu, C.-L. Zhang, and Z.-H. Zhou, “Minimal gated unit for recurrent neural networks,” International Journal of Automation and Computing , vol. 13, no. 3, pp. 226–234, 2016
2016
Earlier work this paper cites.
J. Collins, J. Sohl-Dickstein, and D. Sussillo, “Capacity and trainability in recurrent neural networks,” Proceeding of the International Conference on Learning Representations , 2016
2016
Earlier work this paper cites.
J. Bradbury, S. Merity, C. Xiong, and R. Socher, “Quasi-recurrent neural networks,” Proceeding of the International Conference on Learning Representations , 2016
2016
Earlier work this paper cites.
S. Wisdom, T. Powers, J. Hershey, J. Le Roux, and L. Atlas, “Full-capacity unitary recurrent neural networks,” in Advances in Neural Information Processing Systems , 2016, pp. 4880–4888
2016
Cited alongside, same era.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 770–778
2016
Cited alongside, same era.
2016
Cited alongside, same era.
A. Veit, M. J. Wilber, and S. Belongie, “Residual networks behave like ensembles of relatively shallow networks,” in Advances in Neural Information Processing Systems , 2016, pp. 550–558
2016
Cited alongside, same era.
Z. Yang, Z. Dai, R. Salakhutdinov, and W. W. Cohen, “Breaking the softmax bottleneck: A high-rank rnn language model,” Proceeding of the International Conference on Learning Representations , 2017
2017
Later among the works it cites.
Q. Ke, S. An, M. Bennamoun, F. Sohel, and F. Boussaid, “Skeletonnet: Mining deep part features for 3-d action recognition,” IEEE Signal Processing Letters , vol. 24, no. 6, pp. 731–735, 2017
2017
Later among the works it cites.
C. Li, Y. Hou, P. Wang, and W. Li, “Joint distance maps based action recognition with convolutional neural networks,” IEEE Signal Processing Letters , vol. 24, no. 5, pp. 624–628, 2017
2017
Later among the works it cites.
Q. Ke, M. Bennamoun, S. An, F. Sohel, and F. Boussaid, “A new representation of skeleton sequences for 3d action recognition,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2017
2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. G. Zilly, R. K. Srivastava, J. Koutník, and J. Schmidhuber, “Recurrent highway networks,” Proceedings of the International Conference on Machine Learning , 2016
2016
Cited alongside, same era.
Y. Wu, S. Zhang, Y. Zhang, Y. Bengio, and R. R. Salakhutdinov, “On multiplicative integration with recurrent neural networks,” in Advances in neural information processing systems , 2016, pp. 2856–2864
2016
Cited alongside, same era.
K. He, X. Zhang, S. Ren, and J. Sun, “Identity mappings in deep residual networks,” in European Conference on Computer Vision . Springer, 2016, pp. 630–645
2016
Cited alongside, same era.
S. Semeniuta, A. Severyn, and E. Barth, “Recurrent dropout without memory loss,” Proceedings of the International Conference on Computational Linguistics , 2016
2016
Cited alongside, same era.
T. Cooijmans, N. Ballas, C. Laurent, Ç. Gülçehre, and A. Courville, “Recurrent batch normalization,” Proceeding of the International Conference on Learning Representations , 2016
2016
Cited alongside, same era.
D. Ha, A. Dai, and Q. V. Le, “Hypernetworks,” arXiv preprint arXiv:1609.09106 , 2016
2016
Cited alongside, same era.
J. Chung, S. Ahn, and Y. Bengio, “Hierarchical multiscale recurrent neural networks,” Proceeding of the International Conference on Learning Representations , 2016
2016
Cited alongside, same era.
B. Zoph and Q. V. Le, “Neural architecture search with reinforcement learning,” Proceeding of the International Conference on Learning Representations , 2016
2016
Cited alongside, same era.
M. Liu, H. Liu, and C. Chen, “Enhanced skeleton visualization for view invariant human action recognition,” Pattern Recognition , vol. 68, pp. 346–362, 2017
2017
Later among the works it cites.
2017
Later among the works it cites.
J. Liu, G. Wang, P. Hu, L. yu Duan, and A. C. Kot, “Global context-aware attention lstm networks for 3d action recognition,” 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , pp. 3671–3680, 2017
2017
Later among the works it cites.
J. Liu, A. Shahroudy, D. Xu, A. C. Kot, and G. Wang, “Skeleton-based action recognition using spatio-temporal lstm network with trust gates,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 40, pp. 3007–3021, 2018
2018
Later among the works it cites.
P. Zhang, C. Lan, J. Xing, W. Zeng, J. Xue, and N. Zheng, “View adaptive neural networks for high performance skeleton-based human action recognition,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 41, pp. 1963–1978, 2018
2018
Later among the works it cites.
B. Shuai, Z. L. Zuo, B. Wang, and G. Wang, “Scene segmentation with dag-recurrent neural networks,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 40, pp. 1480–1493, 2018
2018
Later among the works it cites.
J. Li, C. Liu, and Y. Gong, “Layer trajectory lstm,” ArXiv , vol. abs/1808.09522, 2018
2018
Later among the works it cites.
S. Merity, N. S. Keskar, and R. Socher, “Regularizing and optimizing lstm language models,” international conference on learning representations , 2018
2018
Later among the works it cites.
S. Li, W. Li, C. Cook, C. Zhu, and Y. Gao, “Independently recurrent neural network (indrnn): Building a longer and deeper rnn,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2018, pp. 5457–5466
2018
Later among the works it cites.
J. Zhang, Y. Lin, Z. Song, and I. S. Dhillon, “Learning long term dependencies via fourier recurrent units,” Proceedings of the International Conference on Machine Learning , 2018
2018
Later among the works it cites.
B. Krause, E. Kahembwe, I. Murray, and S. Renals, “Dynamic evaluation of neural sequence models,” Proceedings of the International Conference on Machine Learning , pp. 2766–2775, 2018
2018
Later among the works it cites.
C. Li, Q. Zhong, D. Xie, and S. Pu, “Co-occurrence feature learning from skeleton data for action recognition and detection with hierarchical aggregation,” in Proceedings of the Twenty-Seventh International Joint Conference on Artificial Intelligence , 7 2018, pp. 786–792
2018
Later among the works it cites.
S. Yan, Y. Xiong, D. Lin, and xiaoou Tang, “Spatial temporal graph convolutional networks for skeleton-based action recognition,” Proceedings of the AAAI Conference on Artificial Intelligence , pp. 7444–7452, 2018
2018
Later among the works it cites.
K. C. Thakkar and P. J. Narayanan, “Part-based graph convolutional network for action recognition.” Proceedings of the British Machine Vision Conference , p. 270, 2018
2018
Later among the works it cites.
A. Shahroudy, T.-T. Ng, Y. Gong, and G. Wang, “Deep multimodal feature analysis for action recognition in rgb+d videos,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 40, pp. 1045–1058, 2018
2018
Later among the works it cites.
J. Liu, G. Wang, L. yu Duan, K. Abdiyeva, and A. C. Kot, “Skeleton-based human action recognition with global context-aware attention lstm networks,” IEEE Transactions on Image Processing , vol. 27, pp. 1586–1599, 2018
2018
Later among the works it cites.
M. Liu and J. Yuan, “Recognizing human actions as the evolution of pose estimation maps,” 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition , pp. 1159–1168, 2018
2018
Later among the works it cites.
S. Li, W. Li, C. Cook, C. Zhu, and Y. Gao, “A fully trainable network with rnn-based pooling,” Neurocomputing , 2019
2019
Closest in time.
2019
Closest in time.
L. S, Q. Wang, and P. Turaga, “Temporal transformer networks: Joint learning of invariant and discriminative time warping,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2019, pp. 12 426–12 435
2019
Closest in time.
J. Liu, A. Shahroudy, M. Perez, G. Wang, L. yu Duan, and A. C. Kot, “Ntu rgb+d 120: A large-scale benchmark for 3d human activity understanding,” IEEE transactions on pattern analysis and machine intelligence , vol. 42, no. 10, 2019
2019
Closest in time.
J. fang Hu, W.-S. Zheng, L. Ma, G. Wang, J.-H. Lai, and J. Zhang, “Early action prediction by soft regression,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 41, no. 11, pp. 2568–2583, 2019
2019
Closest in time.
2019
Closest in time.
J. Chung, C. Gulcehre, K. Cho, and Y. Bengio, “Gated feedback recurrent neural networks,” in International Conference on Machine Learning , 2015, pp. 2067–2075
2075
Closest in time.