Fetching the paper…
Reading the bibliography…
The process of designing neural architectures requires expert knowledge and extensive trial and error.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
R. J. Williams · 1992
Earlier work this paper cites.
Long Short-Term Memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Evolving Neural Networks through Augmenting Topologies
K. O. Stanley and R. Miikkulainen · 2002
Earlier work this paper cites.
Evolving memory cell structures for sequence learning
J. Bayer, D. Wierstra, J. Togelius, and J. Schmidhuber · 2009
Earlier work this paper cites.
A Hypercube-Based Encoding for Evolving Large-Scale Neural Networks
K. O. Stanley, D. B. D’Ambrosio, and J. Gauci · 2009
Earlier work this paper cites.
Algorithms for hyper-parameter optimization
J. S. Bergstra, R. Bardenet, Y. Bengio, and B. Kégl · 2011
Earlier work this paper cites.
Practical Bayesian Optimization of Machine Learning Algorithms
J. Snoek, H. Larochelle, and R. P. Adams · 2012
Earlier work this paper cites.
On the Properties of Neural Machine Translation: Encoder-Decoder Approaches
K. Cho, B. van Merrienboer, D. Bahdanau, and Y. Bengio · 2014
Earlier work this paper cites.
Recurrent neural network regularization
W. Zaremba, I. Sutskever, and O. Vinyals · 2014
Earlier work this paper cites.
An empirical exploration of recurrent network architectures
R. Jozefowicz, W. Zaremba, and I. Sutskever · 2015
Earlier work this paper cites.
Highway Networks
R. K. Srivastava, K. Greff, and J. Schmidhuber · 2015
Earlier work this paper cites.
Going deeper with convolutions
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich · 2015
Earlier work this paper cites.
Improved semantic representations from tree-structured long short-term memory networks
K. S. Tai, R. Socher, and C. D. Manning · 2015
Cited alongside, same era.
An awkward disparity between bleu/ribes scores and human judgements in machine translation
L. Tan, J. Dehdari, and J. van Genabith · 2015
Cited alongside, same era.
Layer Normalization
J. L. Ba, J. R. Kiros, and G. E. Hinton · 2016
Cited alongside, same era.
Designing Neural Network Architectures using Reinforcement Learning
B. Baker, O. Gupta, N. Naik, and R. Raskar · 2016
Cited alongside, same era.
The IWSLT 2016 Evaluation Campaign
M. Cettolo, J. Niehues, S. Stüker, L. Bentivogli, and M. Federico · 2016
Cited alongside, same era.
Multi30k: Multilingual English-German Image Descriptions
D. Elliott, S. Frank, K. Sima’an, and L. Specia · 2016
Cited alongside, same era.
Tying Word Vectors and Word Classifiers: A Loss Framework for Language Modeling
H. Inan, K. Khosravi, and R. Socher · 2016
Later among the works it cites.
Safe and efficient off-policy reinforcement learning
R. Munos, T. Stepleton, A. Harutyunyan, and M. Bellemare · 2016
Later among the works it cites.
Using the output embedding to improve language models
O. Press and L. Wolf · 2016
Later among the works it cites.
Minimal gated unit for recurrent neural networks
G.-B. Zhou, J. Wu, C.-L. Zhang, and Z.-H. Zhou · 2016
Later among the works it cites.
Recurrent Highway Networks
J. G. Zilly, R. K. Srivastava, J. Koutník, and J. Schmidhuber · 2016
Later among the works it cites.
Quasi-Recurrent Neural Networks
J. Bradbury, S. Merity, C. Xiong, and R. Socher · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Convolution by evolution: Differentiable pattern producing networks
C. Fernando, D. Banarse, M. Reynolds, F. Besse, D. Pfau, M. Jaderberg, M. Lanctot, and D. Wierstra · 2016
Cited alongside, same era.
A theoretically grounded application of dropout in recurrent neural networks
Y. Gal and Z. Ghahramani · 2016
Cited alongside, same era.
LSTM: A search space odyssey
K. Greff, R. K. Srivastava, J. Koutník, B. R. Steunebrink, and J. Schmidhuber · 2016
Cited alongside, same era.
HyperNetworks
D. Ha, A. Dai, and Q. V. Le · 2016
Cited alongside, same era.
Q( λ \lambda ) with Off-Policy Corrections
A. Harutyunyan, M. G. Bellemare, T. Stepleton, and R. Munos · 2016
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
Closest in time.
Massive exploration of neural machine translation architectures
D. Britz, A. Goldie, M.-T. Luong, and Q. V. Le · 2017
Closest in time.
Self-Normalizing Neural Networks
G. Klambauer, T. Unterthiner, A. Mayr, and S. Hochreiter · 2017
Closest in time.
On the State of the Art of Evaluation in Neural Language Models
G. Melis, C. Dyer, and P. Blunsom · 2017
Closest in time.
Attention Is All You Need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Closest in time.
Neural Architecture Search with Reinforcement Learning
B. Zoph and Q. V. Le · 2017
Closest in time.