Fetching the paper…
Reading the bibliography…
Transfer learning is a popular technique for improving the performance of neural networks.
Fixup initialization: Residual learning without normalization, 2019a
Hongyi Zhang, Yann N. Dauphin, and Tengyu Ma · 1901
Earlier work this paper cites.
Eat-nas: Elastic architecture transfer for accelerating large-scale neural architecture search, 2019
Jiemin Fang, Yukang Chen, Xinbang Zhang, Qian Zhang, Chang Huang, Gaofeng Meng, Wenyu Liu, and Xinggang Wang · 1901
Earlier work this paper cites.
Efficientnet: Rethinking model scaling for convolutional neural networks
Mingxing Tan and Quoc V Le · 1905
Earlier work this paper cites.
Energy and policy considerations for deep learning in nlp, 2019
Emma Strubell, Ananya Ganesh, and Andrew McCallum · 1906
Earlier work this paper cites.
The string-to-string correction problem
Robert A. Wagner and Michael J. Fischer · 1974
Earlier work this paper cites.
Efficient algorithms for finding maximum matching in graphs
Zvi Galil · 1986
Earlier work this paper cites.
Towards the systematic reporting of the energy and carbon footprints of machine learning
Peter Henderson, Jieru Hu, Joshua Romoff, Emma Brunskill, Dan Jurafsky, and Joelle Pineau · 2002
Earlier work this paper cites.
Rexnet: Diminishing representational bottleneck on convolutional neural network
Dongyoon Han, Sangdoo Yun, Byeongho Heo, and Young Joon Yoo · 2007
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Y. Bengio · 2010
Earlier work this paper cites.
Efficient backprop
Yann A LeCun, Léon Bottou, Genevieve B Orr, and Klaus-Robert Müller · 2012
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2012
Earlier work this paper cites.
Predicting parameters in deep learning
Misha Denil, Babak Shakibi, Laurent Dinh, Marc’Aurelio Ranzato, and Nando de Freitas · 2013
Earlier work this paper cites.
Going deeper with convolutions, 2014
Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich · 2014
Cited alongside, same era.
How transferable are features in deep neural networks?
Jason Yosinski, Jeff Clune, Yoshua Bengio, and Hod Lipson · 2014
Cited alongside, same era.
Food-101 – mining discriminative components with random forests
Lukas Bossard, Matthieu Guillaumin, and Luc Van Gool · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization, 2014
Diederik P. Kingma and Jimmy Ba · 2014
Cited alongside, same era.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification, 2015
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Cited alongside, same era.
Gpipe: Efficient training of giant neural networks using pipeline parallelism, 2018
Yanping Huang, Youlong Cheng, Ankur Bapna, Orhan Firat, Mia Xu Chen, Dehao Chen, HyoukJoong Lee, Jiquan Ngiam, Quoc V. Le, Yonghui Wu, and Zhifeng Chen · 2018
Later among the works it cites.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Jonathan Frankle and Michael Carbin · 2018
Later among the works it cites.
Path-level network transformation for efficient architecture search, 2018
Han Cai, Jiacheng Yang, Weinan Zhang, Song Han, and Yong Yu · 2018
Later among the works it cites.
Efficient Neural Architecture Search via Parameter Sharing
Hieu Pham, Melody Y. Guan, Barret Zoph, Quoc V. Le, and Jeff Dean · 2018
Later among the works it cites.
Efficient multi-objective neural architecture search via lamarckian evolution, 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
All you need is a good init, 2015
Dmytro Mishkin and Jiri Matas · 2015
Cited alongside, same era.
Distilling the knowledge in a neural network
Jeff Dean Geoffrey Hinton, Oriol Vinyals · 2015
Cited alongside, same era.
Deep Residual Learning for Image Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Cited alongside, same era.
ImageNet Large Scale Visual Recognition Challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei · 2015
Cited alongside, same era.
Net2net: Accelerating learning via knowledge transfer, 2016
Tianqi Chen, Ian Goodfellow, and Jonathon Shlens · 2016
Cited alongside, same era.
Efficient architecture search by network transformation, 2017
Han Cai, Tianyao Chen, Weinan Zhang, Yong Yu, and Jun Wang · 2017
Cited alongside, same era.
David Ha, Andrew Dai, and Quoc V. Le · 2017
Cited alongside, same era.
Thomas Elsken, Jan Hendrik Metzen, and Frank Hutter · 2018
Later among the works it cites.
Mark Sandler, Andrew G. Howard, Menglong Zhu, Andrey Zhmoginov, and Liang-Chieh Chen · 2018
Later among the works it cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Later among the works it cites.
Metainit: Initializing learning by learning to initialize
Yann N Dauphin and Samuel Schoenholz · 2019
Later among the works it cites.
Pytorch image models
Ross Wightman · 2019
Later among the works it cites.
Mnasnet: Platform-aware neural architecture search for mobile, 2019
Mingxing Tan, Bo Chen, Ruoming Pang, Vijay Vasudevan, Mark Sandler, Andrew Howard, and Quoc V. Le · 2019
Later among the works it cites.
Knowledge distillation: A survey, 2020
Jianping Gou, Baosheng Yu, Stephen John Maybank, and Dacheng Tao · 2020
Later among the works it cites.
Parameter prediction for unseen deep architectures, 2021
Boris Knyazev, Michal Drozdzal, Graham W. Taylor, and Adriana Romero-Soriano · 2021
Later among the works it cites.