Fetching the paper…
Reading the bibliography…
We propose Efficient Neural Architecture Search (ENAS), a fast and inexpensive approach for automatic model design.
A method for solving the convex programming problem with convergence rate o ( 1 / k 2 ) o(1/k^{2})
Nesterov, Yurii E · 1983
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, Ronald J · 1992
Earlier work this paper cites.
The penn treebank: Annotating predicate argument structure
Marcus, Mitchell, Kim, Grace, Marcinkiewicz, Mary Ann, MacIntyre, Robert, Bies, Ann, Ferguson, Mark, Katz, Karen, and Schasberger, Britta · 1994
Earlier work this paper cites.
Long short-term memory
Hochreiter, Sepp and Schmidhuber, Jürgen · 1997
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, Alex · 2009
Earlier work this paper cites.
Network in network
Lin, Min, Chen, Qiang, and Yan, Shuicheng · 2013
Earlier work this paper cites.
Cnn features off-the-shelf: an astounding baseline for recognition
Razavian, Ali Sharif, Azizpour, Hossein, Josephine, Sullivan, and Carlsson, Stefan · 2014
Earlier work this paper cites.
Recurrent neural network regularization
Zaremba, Wojciech, Sutskever, Ilya, and Vinyals, Oriol · 2014
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
He, Kaiming, Zhang, Xiangyu, Rein, Shaoqing, and Sun, Jian · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, Sergey and Szegedy, Christian · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, Diederik P. and Ba, Jimmy Lei · 2015
Earlier work this paper cites.
Gradient estimation using stochastic computation graphs
Schulman, John, Heess, Nicolas, Weber, Theophane, and Abbeel, Pieter · 2015
Earlier work this paper cites.
A theoretically grounded application of dropout in recurrent neural networks
Gal, Yarin and Ghahramani, Zoubin · 2016
Earlier work this paper cites.
Shake-shake regularization of 3-branch residual networks
Gastaldi, Xavier · 2016
Earlier work this paper cites.
Densely connected convolutional networks
Huang, Gao, Liu, Zhuang, van der Maaten, Laurens, and Weinberger, Kilian Q · 2016
Earlier work this paper cites.
Multi-task sequence to sequence learning
Luong, Minh-Thang, Le, Quoc V., Sutskever, Ilya, Vinyals, Oriol, and Kaiser, Lukasz · 2016
Cited alongside, same era.
Convolutional neural fabrics
Saxena, Shreyas and Verbeek, Jakob · 2016
Cited alongside, same era.
Transfer learning for low-resource neural machine translation
Zoph, Barret, Yuret, Deniz, May, Jonathan, and Knight, Kevin · 2016
Cited alongside, same era.
Xception: Deep learning with depthwise separable convolutions
Chollet, Francois · 2017
Cited alongside, same era.
Capacity and trainability in recurrent neural networks
Collins, Jasmine, Sohl-Dickstein, Jascha, and Sussillo, David · 2017
Cited alongside, same era.
Peephole: Predicting network performance before training
Deng, Boyang, Yan, Junjie, and Lin, Dahua · 2017
Cited alongside, same era.
On the state of the art of evaluation in neural language models
Melis, Gábor, Dyer, Chris, and Blunsom, Phil · 2017
Later among the works it cites.
Regularizing and optimizing LSTM language models
Merity, Stephen, Keskar, Nitish Shirish, and Socher, Richard · 2017
Later among the works it cites.
Deeparchitect: Automatically designing and training deep architectures
Negrinho, Renato and Gordon, Geoff · 2017
Later among the works it cites.
Large-scale evolution of image classifiers
Real, Esteban, Moore, Sherry, Selle, Andrew, Saxena, Saurabh, Leon, Yutaka Suematsu, Tan, Jie, Le, Quoc, and Kurakin, Alex · 2017
Later among the works it cites.
Learning time-efficient deep architectures with budgeted super networks
Veniat, Tom and Denoyer, Ludovic · 2017
Later among the works it cites.
Recurrent highway networks
Zilly, Julian Georg, Srivastava, Rupesh Kumar, Koutník, Jan, and Schmidhuber, Jürgen · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Improved regularization of convolutional neural networks with cutout
DeVries, Terrance and Taylor, Graham W · 2017
Cited alongside, same era.
Improving neural language models with a continuous cache
Grave, Edouard, Joulin, Armand, and Usunier, Nicolas · 2017
Cited alongside, same era.
Hypernetworks
Ha, David, Dai, Andrew, and Le, Quoc V · 2017
Cited alongside, same era.
Tying word vectors and word classifiers: a loss framework for language modeling
Inan, Hakan, Khosravi, Khashayar, and Socher, Richard · 2017
Cited alongside, same era.
Dynamic evaluation of neural sequence models
Krause, Ben, Kahembwe, Emmanuel, Murray, Iain, and Renals, Steve · 2017
Cited alongside, same era.
Fractalnet: Ultra-deep neural networks without residuals
Larsson, Gustav, Maire, Michael, and Shakhnarovich, Gregory · 2017
Cited alongside, same era.
Later among the works it cites.
Neural architecture search with reinforcement learning
Zoph, Barret and Le, Quoc V · 2017
Later among the works it cites.
SMASH: one-shot model architecture search through hypernetworks
Brock, Andrew, Lim, Theodore, Ritchie, James M., and Weston, Nick · 2018
Closest in time.
Efficient architecture search by network transformation
Cai, Han, Chen, Tianyao, Zhang, Weinan, Yu, Yong., and Wang, Jun · 2018
Closest in time.
Hierarchical representations for efficient architecture search
Liu, Hanxiao, Simonyan, Karen, Vinyals, Oriol, Fernando, Chrisantha, and Kavukcuoglu, Koray · 2018
Closest in time.
Peephole: Predicting network performance before training
Real, Esteban, Aggarwal, Alok, Huang, Yanping, and Le, Quoc V · 2018
Closest in time.
Breaking the softmax bottleneck: A high-rank rnn language model
Yang, Zhilin, Dai, Zihang, Salakhutdinov, Ruslan, and Cohen, William · 2018
Closest in time.
Practical network blocks design with q-learning
Zhong, Zhao, Yan, Junjie, and Liu, Cheng-Lin · 2018
Closest in time.
Learning transferable architectures for scalable image recognition
Zoph, Barret, Vasudevan, Vijay, Shlens, Jonathon, and Le, Quoc V · 2018
Closest in time.