Fetching the paper…
Reading the bibliography…
Whereas deep neural networks were first mostly used for classification tasks, they are rapidly expanding in the realm of structured output problems, where the observed target is composed of multiple random variables that have a rich joint distribution, given the input.
D. E. Rumelhart, G. E. Hinton, and R. J. Williams, “Learning representations by back-propagating errors,” Nature , vol. 323, pp. 533–536, 1986
1986
Earlier work this paper cites.
G. E. Hinton and T. J. Sejnowski, “Learning and relearning in Boltzmann machines,” in Parallel Distributed Processing: Explorations in the Microstructure of Cognition. Volume 1: Foundations , D. E. Rumelhart and J. L. McClelland, Eds. Cambridge, MA: MIT Press, 1986, pp. 282–317
1986
Earlier work this paper cites.
J. S. Garofolo, L. F. Lamel, W. M. Fisher, J. G. Fiscus, and D. S. Pallett, “Darpa timit acoustic-phonetic continous speech corpus cd-rom. nist speech disc 1-1.1,” NASA STI/Recon Technical Report N , vol. 93, p. 27403, 1993
1993
Earlier work this paper cites.
Y. Bengio, P. Simard, and P. Frasconi, “Learning long-term dependencies with gradient descent is difficult,” IEEE Transactions on Neural Nets , pp. 157–166, 1994
1994
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural Computation , vol. 9, no. 8, pp. 1735–1780, 1997
1997
Earlier work this paper cites.
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner, “Gradient-based learning applied to document recognition,” Proceedings of the IEEE , vol. 86, no. 11, pp. 2278–2324, Nov. 1998
1998
Earlier work this paper cites.
S. Hochreiter, F. F. Informatik, Y. Bengio, P. Frasconi, and J. Schmidhuber, “Gradient flow in recurrent nets: the difficulty of learning long-term dependencies,” in Field Guide to Dynamical Recurrent Networks , J. Kolen and S. Kremer, Eds. IEEE Press, 2000
2000
Earlier work this paper cites.
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu, “Bleu: a method for automatic evaluation of machine translation,” in Proceedings of the 40th annual meeting on association for computational linguistics . Association for Computational Linguistics, 2002, pp. 311–318
2002
Earlier work this paper cites.
C.-Y. Lin, “Rouge: A package for automatic evaluation of summaries,” in Text summarization branches out: Proceedings of the ACL-04 workshop , vol. 8, 2004
2004
Earlier work this paper cites.
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, “Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks,” in ICML’2006 , Pittsburgh, USA, 2006, pp. 369–376
2006
Earlier work this paper cites.
A. Mnih and G. E. Hinton, “Three new graphical models for statistical language modelling,” 2007, pp. 641–648
2007
Earlier work this paper cites.
2007
Earlier work this paper cites.
G. Taylor, R. Fergus, Y. LeCun, and C. Bregler, “Convolutional learning of spatio-temporal features,” in ECCV’10 , 2010
2010
Earlier work this paper cites.
H. Larochelle and G. E. Hinton, “Learning to combine foveal glimpses with a third-order Boltzmann machine,” in Advances in Neural Information Processing Systems 23 , 2010, pp. 1243–1251
2010
Earlier work this paper cites.
T. Mikolov, S. Kombrink, L. Burget, J. Cernocky, and S. Khudanpur, “Extensions of recurrent neural network language model,” in Proc. 2011 IEEE international conference on acoustics, speech and signal processing (ICASSP 2011) , 2011
2011
Earlier work this paper cites.
D. L. Chen and W. B. Dolan, “Collecting highly parallel data for paraphrase evaluation,” in Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics , Portland, Oregon, USA, June 2011, pp. 190–200
2011
Earlier work this paper cites.
A. Graves, “Sequence transduction with recurrent neural networks,” in Proceedings of the 29th International Conference on Machine Learning (ICML 2012) , 2012
2012
Earlier work this paper cites.
A. Krizhevsky, I. Sutskever, and G. Hinton, “ImageNet classification with deep convolutional neural networks,” in Advances in Neural Information Processing Systems 25 (NIPS’2012) , 2012
2012
Earlier work this paper cites.
G. E. Hinton, L. Deng, D. Yu, G. E. Dahl, A. Mohamed, N. Jaitly, A. Senior, V. Vanhoucke, P. Nguyen, T. N. Sainath, and B. Kingsbury, “Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups,” IEEE Signal Process. Mag. , vol. 29, no. 6, pp. 82–97, 2012
2012
Earlier work this paper cites.
M. Denil, L. Bazzani, H. Larochelle, and N. de Freitas, “Learning where to attend with deep architectures for image tracking,” Neural Computation , vol. 24, no. 8, pp. 2151–2184, 2012
2012
Earlier work this paper cites.
G. Kulkarni, V. Premraj, V. Ordonez, S. Dhar, S. Li, Y. Choi, A. C. Berg, and T. Berg, “Babytalk: Understanding and generating simple image descriptions,” Pattern Analysis and Machine Intelligence, IEEE Transactions on , vol. 35, no. 12, pp. 2891–2903, 2013
2013
Earlier work this paper cites.
N. Boulanger-Lewandowski, Y. Bengio, and P. Vincent, “Audio chord recognition with recurrent neural networks,” in ISMIR , 2013
2013
Earlier work this paper cites.
N. Kalchbrenner and P. Blunsom, “Recurrent continuous translation models,” in EMNLP’2013 , 2013
2013
Earlier work this paper cites.
A. Graves, “Generating sequences with recurrent neural networks,” arXiv:1308.0850, Tech. Rep., 2013
2013
Earlier work this paper cites.
2013
Earlier work this paper cites.
Y. Tang and R. Salakhutdinov, “Learning stochastic feedforward neural networks,” in NIPS’2013 , 2013
2013
Cited alongside, same era.
M. Hodosh, P. Young, and J. Hockenmaier, “Framing image description as a ranking task: Data, models and evaluation metrics,” Journal of Artificial Intelligence Research , pp. 853–899, 2013
2013
Cited alongside, same era.
A. Graves, A.-r. Mohamed, and G. Hinton, “Speech recognition with deep recurrent neural networks,” in ICASSP’2013 , 2013, pp. 6645–6649
2013
Cited alongside, same era.
K. Cho, B. van Merrienboer, C. Gulcehre, F. Bougares, H. Schwenk, and Y. Bengio, “Learning phrase representations using RNN encoder-decoder for statistical machine translation,” in Proceedings of the Empiricial Methods in Natural Language Processing (EMNLP 2014) , Oct. 2014
2014
Cited alongside, same era.
L. Tóth, “Combining time-and frequency-domain convolution in convolutional neural network-based phone recognition,” in ICASSP 2014 , 2014, pp. 190–194
2014
Later among the works it cites.
2014
Later among the works it cites.
V. Mnih, N. Heess, A. Graves, and k. kavukcuoglu, “Recurrent models of visual attention,” in Advances in Neural Information Processing Systems 27 , Z. Ghahramani, M. Welling, C. Cortes, N. Lawrence, and K. Weinberger, Eds. Curran Associates, Inc., 2014, pp. 2204–2212
2014
Later among the works it cites.
Y. Zheng, R. S. Zemel, Y.-J. Zhang, and H. Larochelle, “A neural autoregressive approach to attention-based recognition,” International Journal of Computer Vision , vol. 113, no. 1, pp. 67–79, 2014
2014
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Chung, C. Gulcehre, K. Cho, and Y. Bengio, “Empirical evaluation of gated recurrent neural networks on sequence modeling,” NIPS’2014 Deep Learning workshop, arXiv 1412.3555, 2014
2014
Cited alongside, same era.
M. D. Zeiler and R. Fergus, “Visualizing and understanding convolutional networks,” in ECCV’14 , 2014
2014
Cited alongside, same era.
2014
Cited alongside, same era.
P. Sermanet, D. Eigen, X. Zhang, M. Mathieu, R. Fergus, and Y. LeCun, “Overfeat: Integrated recognition, localization and detection using convolutional networks,” International Conference on Learning Representations , 2014
2014
Cited alongside, same era.
A. S. Razavian, H. Azizpour, J. Sullivan, and S. Carlsson, “Cnn features off-the-shelf: an astounding baseline for recognition,” in Computer Vision and Pattern Recognition Workshops (CVPRW), 2014 IEEE Conference on . IEEE, 2014, pp. 512–519
2014
Cited alongside, same era.
R. Kiros, R. Salakhutdinov, and R. Zemel, “Unifying visual-semantic embeddings with multimodal neural language models,” arXiv: 1411.2539 [cs.LG]
2014
Cited alongside, same era.
I. Sutskever, O. Vinyals, and Q. V. Le, “Sequence to sequence learning with neural networks,” in NIPS’2014 , 2014
2014
Cited alongside, same era.
2014
Cited alongside, same era.
2014
Later among the works it cites.
J. Weston, S. Chopra, and A. Bordes, “Memory networks,” arXiv preprint arXiv:1410.3916 , 2014
2014
Later among the works it cites.
L. Wehbe, B. Murphy, P. Talukdar, A. Fyshe, A. Ramdas, and T. Mitchell, “Simultaneously uncovering the patterns of brain regions involved in different story reading subprocesses,” PLOS ONE , vol. 9, no. 11, p. e112575, Nov. 2014
2014
Later among the works it cites.
K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,” in ICLR , 2015
2015
Closest in time.
2015
Closest in time.
2015
Closest in time.
2015
Closest in time.
2015
Closest in time.
K. Xu, J. L. Ba, R. Kiros, K. Cho, A. Courville, R. Salakhutdinov, R. S. Zemel, and Y. Bengio, “Show, attend and tell: Neural image caption generation with visual attention,” in ICML’2015 , 2015
2015
Closest in time.
L. Yao, A. Torabi, K. Cho, N. Ballas, C. Pal, H. Larochelle, and A. Courville, “Describing videos by exploiting temporal structure,” arXiv: 1502.08029 , 2015
2015
Closest in time.
J. T. Springenberg, A. Dosovitskiy, T. Brox, and M. Riedmiller, “Striving for simplicity: The all convolutional net,” in ICLR , 2015
2015
Closest in time.
2015
Closest in time.
J. Chorowski, D. Bahdanau, D. Serdyuk, K. Cho, and Y. Bengio, “Attention-based models for speech recognition,” arXiv preprint arXiv: 1506.07503 , 2015
2015
Closest in time.
T. Raiko, M. Berglund, G. Alain, and L. Dinh, “Techniques for learning binary stochastic feedforward neural networks,” in ICLR , 2015
2015
Closest in time.
2015
Closest in time.
2015
Closest in time.
A. Torabi, C. Pal, H. Larochelle, and A. Courville, “Using descriptive video services to create a large data source for video annotation research,” arXiv preprint arXiv: 1503.01070 , 2015
2015
Closest in time.
O. Vinyals, M. Fortunato, and N. Jaitly, “Pointer networks,” arXiv preprint arXiv:1506.03134 , 2015
2015
Closest in time.
2015
Closest in time.