Fetching the paper…
Reading the bibliography…
Recurrent neural networks (RNNs) have shown promising performance for language modeling.
A stochastic approximation method
H. Robbins and S. Monro. 1951 · 1951
Earlier work this paper cites.
Backpropagation through time: what it does and how to do it
P. Werbos. 1990 · 1990
Earlier work this paper cites.
A practical Bayesian framework for backpropagation networks
D. J. C. MacKay. 1992 · 1992
Earlier work this paper cites.
Building a large annotated corpus of english: The penn treebank
M. P. Marcus, M. A. Marcinkiewicz, and B. Santorini. 1993 · 1993
Earlier work this paper cites.
Bayesian learning for neural networks
R. M. Neal. 1995 · 1995
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Learning question classifiers
X. Li and D. Roth. 2002 · 2002
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
K. Papineni, S. Roukos, T. Ward, and W. Zhu. 2002 · 2002
Earlier work this paper cites.
Mining and summarizing customer reviews
M. Hu and B. Liu. 2004 · 2004
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
C. Lin. 2004 · 2004
Earlier work this paper cites.
A sentimental education: Sentiment analysis using subjectivity summarization based on minimum cuts
B. Pang and L. Lee. 2004 · 2004
Earlier work this paper cites.
Meteor: An automatic metric for mt evaluation with improved correlation with human judgments
S. Banerjee and A. Lavie. 2005 · 2005
Earlier work this paper cites.
Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales
B. Pang and L. Lee. 2005 · 2005
Earlier work this paper cites.
Annotating expressions of opinions and emotions in language
J. Wiebe, T. Wilson, and C. Cardie. 2005 · 2005
Earlier work this paper cites.
Visualizing data using t-SNE
L. Van der Maaten and G. E. Hinton. 2008 · 2008
Earlier work this paper cites.
Large-scale machine learning with stochastic gradient descent
L Bottou. 2010 · 2010
Earlier work this paper cites.
Recurrent neural network based language model
T. Mikolov, M. Karafiát, L. Burget, J. Cernockỳ, and S. Khudanpur. 2010 · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
J. Duchi, E. Hazan, and Y. Singer. 2011 · 2011
Earlier work this paper cites.
Practical variational inference for neural networks
A. Graves. 2011 · 2011
Earlier work this paper cites.
Generating text with recurrent neural networks
I. Sutskever, J. Martens, and G. E. Hinton. 2011 · 2011
Earlier work this paper cites.
Bayesian learning via stochastic gradient Langevin dynamics
M. Welling and Y. W. Teh. 2011 · 2011
Cited alongside, same era.
Improving neural networks by preventing co-adaptation of feature detectors
G. Hinton, N. Srivastava, A. Krizhevsky, I. Sutskever, and R Salakhutdinov. 2012 · 2012
Cited alongside, same era.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
T. Tieleman and G. Hinton. 2012 · 2012
Cited alongside, same era.
On fast dropout and its applicability to recurrent networks
J. Bayer, C. Osendorfer, D. Korhammer, N. Chen, S. Urban, and P. van der Smagt. 2013 · 2013
Cited alongside, same era.
Generating sequences with recurrent neural networks
A. Graves. 2013 · 2013
Cited alongside, same era.
Scalable deep poisson factor analysis for topic modeling
Z. Gan, C. Chen, R. Henao, D. Carlson, and L. Carin. 2015 · 2015
Later among the works it cites.
Probabilistic backpropagation for scalable learning of Bayesian neural networks
J. M. Hernández-Lobato and R. P. Adams. 2015 · 2015
Later among the works it cites.
Adam: A method for stochastic optimization
D. Kingma and J. Ba. 2015 · 2015
Later among the works it cites.
Variational dropout and the local reparameterization trick
D. Kingma, T. Salimans, and M. Welling. 2015 · 2015
Later among the works it cites.
Skip-thought vectors
R. Kiros, Y. Zhu, R. Salakhutdinov, R. Zemel, R. Urtasun, A. Torralba, and S. Fidler. 2015 · 2015
Later among the works it cites.
Bayesian dark knowledge
A. Korattikara, V. Rathod, K. Murphy, and M. Welling. 2015 · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Framing image description as a ranking task: Data, models and evaluation metrics
M. Hodosh, P. Young, and J. Hockenmaier. 2013 · 2013
Cited alongside, same era.
Regularization and nonlinearities for neural language models: when are they needed?
M. Pachitariu and M. Sahani. 2013 · 2013
Cited alongside, same era.
On the difficulty of training recurrent neural networks
R. Pascanu, T. Mikolov, and Y. Bengio. 2013 · 2013
Cited alongside, same era.
Regularization of neural networks using DropConnect
L. Wan, M. Zeiler, S. Zhang, Y. LeCun, and R. Fergus. 2013 · 2013
Cited alongside, same era.
Fast Dropout training
S. Wang and C. Manning. 2013 · 2013
Cited alongside, same era.
Stochastic gradient Hamiltonian Monte Carlo
T. Chen, E. B. Fox, and C. Guestrin. 2014 · 2014
Cited alongside, same era.
Learning phrase representations using RNN encoder-decoder for statistical machine translation
K. Cho, B. Van Merriënboer, C. Gulcehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio. 2014 · 2014
Cited alongside, same era.
Rnndrop: A novel dropout for rnns in asr
T. Moon, H. Choi, H. Lee, and I. Song. 2015 · 2015
Later among the works it cites.
Cider: Consensus-based image description evaluation
R. Vedantam, C. L. Zitnick, and D. Parikh. 2015 · 2015
Later among the works it cites.
Show and tell: A neural image caption generator
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan. 2015 · 2015
Later among the works it cites.
Bridging the gap between stochastic gradient MCMC and stochastic optimization
C. Chen, D. Carlson, Z. Gan, C. Li, and L. Carin. 2016 · 2016
Closest in time.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun. 2016 · 2016
Closest in time.
Visualizing and understanding recurrent networks
A. Karpathy, J. Johnson, and L. Fei-Fei. 2016 · 2016
Closest in time.
Sequence-level knowledge distillation
Y. Kim and A. M. Rush. 2016 · 2016
Closest in time.
Distilling an ensemble of greedy dependency parsers into one mst parser
A. Kuncoro, M. Ballesteros, L. Kong, C. Dyer, and N. A. Smith. 2016 · 2016
Closest in time.
Adding gradient noise improves learning for very deep networks
A. Neelakantan, L. Vilnis, Q. Le, I. Sutskever, L. Kaiser, K. Kurach, and J. Martens. 2016 · 2016
Closest in time.
Recurrent dropout without memory loss
S. Semeniuta, A. Severyn, and E. Barth. 2016 · 2016
Closest in time.
Consistency and fluctuations for stochastic gradient Langevin dynamics
Y. W. Teh, A. H. Thiéry, and S. J. Vollmer. 2016 · 2016
Closest in time.
Theano: A Python framework for fast computation of mathematical expressions
Theano Development Team. 2016 · 2016
Closest in time.
Semantic compositional networks for visual captioning
Z. Gan, C. Gan, X. He, Y. Pu, K. Tran, J. Gao, L. Carin, and L. Deng. 2017 · 2017
Closest in time.