Fetching the paper…
Reading the bibliography…
This paper addresses the scalability challenge of architecture search by formulating the task in a differentiable manner.
Hierarchical optimization: An introduction
G Anandalingam and TL Friesz · 1992
Earlier work this paper cites.
Acceleration of stochastic approximation by averaging
Boris T Polyak and Anatoli B Juditsky · 1992
Earlier work this paper cites.
An overview of bilevel optimization
Benoît Colson, Patrice Marcotte, and Gilles Savard · 2007
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Gradient-based hyperparameter optimization through reversible learning
Dougal Maclaurin, David Duvenaud, and Ryan Adams · 2015
Earlier work this paper cites.
Going deeper with convolutions
Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, Andrew Rabinovich, Jen-Hao Rick Chang, et al · 2015
Earlier work this paper cites.
A theoretically grounded application of dropout in recurrent neural networks
Yarin Gal and Zoubin Ghahramani · 2016
Earlier work this paper cites.
Improving neural language models with a continuous cache
Edouard Grave, Armand Joulin, and Nicolas Usunier · 2016
Earlier work this paper cites.
Sgdr: Stochastic gradient descent with warm restarts
Ilya Loshchilov and Frank Hutter · 2016
Earlier work this paper cites.
Scalable gradient-based tuning of continuous regularization hyperparameters
Jelena Luketina, Mathias Berglund, Klaus Greff, and Tapani Raiko · 2016
Earlier work this paper cites.
Hyperparameter optimization with approximate gradient
Fabian Pedregosa · 2016
Earlier work this paper cites.
Convolutional neural fabrics
Shreyas Saxena and Jakob Verbeek · 2016
Earlier work this paper cites.
Julian Georg Zilly, Rupesh Kumar Srivastava, Jan Koutník, and Jürgen Schmidhuber · 2016
Earlier work this paper cites.
Connectivity learning in multi-branch networks
Karim Ahmed and Lorenzo Torresani · 2017
Earlier work this paper cites.
Designing neural network architectures using reinforcement learning
Bowen Baker, Otkrist Gupta, Nikhil Naik, and Ramesh Raskar · 2017
Cited alongside, same era.
Improved regularization of convolutional neural networks with cutout
Terrance DeVries and Graham W Taylor · 2017
Cited alongside, same era.
Simple and efficient architecture search for convolutional neural networks
Thomas Elsken, Jan-Hendrik Metzen, and Frank Hutter · 2017
Cited alongside, same era.
Model-agnostic meta-learning for fast adaptation of deep networks
Chelsea Finn, Pieter Abbeel, and Sergey Levine · 2017
Cited alongside, same era.
Mobilenets: Efficient convolutional neural networks for mobile vision applications
Andrew G Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam · 2017
Accelerating neural architecture search using performance prediction
Bowen Baker, Otkrist Gupta, Ramesh Raskar, and Nikhil Naik · 2018
Closest in time.
Understanding and simplifying one-shot architecture search
Gabriel Bender, Pieter-Jan Kindermans, Barret Zoph, Vijay Vasudevan, and Quoc Le · 2018
Closest in time.
Smash: one-shot model architecture search through hypernetworks
Andrew Brock, Theodore Lim, James M Ritchie, and Nick Weston · 2018
Closest in time.
Efficient architecture search by network transformation
Han Cai, Tianyao Chen, Weinan Zhang, Yong Yu, and Jun Wang · 2018
Closest in time.
Bilevel programming for hyperparameter optimization and meta-learning
Luca Franceschi, Paolo Frasconi, Saverio Salzo, and Massimilano Pontil · 2018
Closest in time.
Neural architecture search with bayesian optimisation and optimal transport
Kirthevasan Kandasamy, Willie Neiswanger, Jeff Schneider, Barnabas Poczos, and Eric Xing · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Densely connected convolutional networks
Gao Huang, Zhuang Liu, Kilian Q Weinberger, and Laurens van der Maaten · 2017
Cited alongside, same era.
Tying word vectors and word classifiers: A loss framework for language modeling
Hakan Inan, Khashayar Khosravi, and Richard Socher · 2017
Cited alongside, same era.
Dynamic evaluation of neural sequence models
Ben Krause, Emmanuel Kahembwe, Iain Murray, and Steve Renals · 2017
Cited alongside, same era.
Unrolled generative adversarial networks
Luke Metz, Ben Poole, David Pfau, and Jascha Sohl-Dickstein · 2017
Cited alongside, same era.
Deeparchitect: Automatically designing and training deep architectures
Renato Negrinho and Geoff Gordon · 2017
Cited alongside, same era.
Automatic differentiation in pytorch
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer · 2017
Cited alongside, same era.
Learning time/memory-efficient deep architectures with budgeted super networks
Tom Veniat and Ludovic Denoyer · 2017
Cited alongside, same era.
Closest in time.
On the state of the art of evaluation in neural language models
Gábor Melis, Chris Dyer, and Phil Blunsom · 2018
Closest in time.
Regularizing and optimizing lstm language models
Stephen Merity, Nitish Shirish Keskar, and Richard Socher · 2018
Closest in time.
Authors’ implementation of “Efficient Neural Architecture Search via Parameter Sharing”
Hieu Pham, Melody Y Guan, Barret Zoph, Quoc V Le, and Jeff Dean · 2018
Closest in time.
Regularized evolution for image classifier architecture search
Esteban Real, Alok Aggarwal, Yanping Huang, and Quoc V Le · 2018
Closest in time.
Differentiable neural network architecture search
Richard Shin, Charles Packer, and Dawn Song · 2018
Closest in time.
Breaking the softmax bottleneck: a high-rank rnn language model
Zhilin Yang, Zihang Dai, Ruslan Salakhutdinov, and William W Cohen · 2018
Closest in time.
Practical block-wise neural network architecture generation
Zhao Zhong, Junjie Yan, Wei Wu, Jing Shao, and Cheng-Lin Liu · 2018
Closest in time.
Learning transferable architectures for scalable image recognition
Barret Zoph, Vijay Vasudevan, Jonathon Shlens, and Quoc V Le · 2018
Closest in time.