Fetching the paper…
Reading the bibliography…
We introduce techniques for rapidly transferring the information stored in one neural net into another neural net.
The cascade-correlation learning architecture
Fahlman, Scott E. and Lebiere, Christian · 1990
Earlier work this paper cites.
Lifelong learning: A case study
Thrun, Sebastian · 1995
Earlier work this paper cites.
Model compression
Buciluǎ, Cristian, Caruana, Rich, and Niculescu-Mizil, Alexandru · 2006
Earlier work this paper cites.
A fast learning algorithm for deep belief nets
Hinton, Geoffrey E., Osindero, Simon, and Teh, Yee Whye · 2006
Earlier work this paper cites.
Knowledge transfer in deep convolutional neural nets
Gutstein, Steven, Fuentes, Olac, and Freudenthal, Eric · 2008
Earlier work this paper cites.
What is the best multi-stage architecture for object recognition?
Jarrett, Kevin, Kavukcuoglu, Koray, Ranzato, Marc’Aurelio, and LeCun, Yann · 2009
Earlier work this paper cites.
Deep sparse rectifier neural networks
Glorot, X., Bordes, A., and Bengio, Y · 2011
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tieleman, T and Hinton, G · 2012
Cited alongside, same era.
Maxout networks
Goodfellow, Ian J., Warde-Farley, David, Mirza, Mehdi, Courville, Aaron, and Bengio, Yoshua · 2013
Cited alongside, same era.
Lifelong machine learning systems: Beyond learning algorithms
Silver, DL, Yang, Q, and Li, L · 2013
Cited alongside, same era.
Parsing with compositional vector grammars
Socher, Richard, Bauer, John, Manning, Christopher D., and Ng, Andrew Y · 2013
Cited alongside, same era.
FitNets: Hints for thin deep nets
Romero, Adriana, Ballas, Nicolas, Ebrahimi Kahou, Samira, Chassang, Antoine, Gatta, Carlo, and Bengio, Yoshua · 2014
Cited alongside, same era.
Dropout: A simple way to prevent neural networks from overfitting
Srivastava, Nitish, Hinton, Geoffrey, Krizhevsky, Alex, Sutskever, Ilya, and Salakhutdinov, Ruslan · 2014
TensorFlow: Large-scale machine learning on heterogeneous systems, 2015
Abadi, Martín, Agarwal, Ashish, Barham, Paul, Brevdo, Eugene, Chen, Zhifeng, Citro, Craig, Corrado, Greg S., Davis, Andy, Dean, Jeffrey, Devin, Matthieu, Ghemawat, Sanjay, Goodfellow, Ian, Harp, Andrew, Irving, Geoffrey, Isard, Michael, Jia, Yangqing, Jozefowicz, Rafal, Kaiser, Lukasz, Kudlur, Manjunath, Levenberg, Josh, Mané, Dan, Monga, Rajat, Moore, Sherry, Murray, Derek, Olah, Chris, Schuster, Mike, Shlens, Jonathon, Steiner, Benoit, Sutskever, Ilya, Talwar, Kunal, Tucker, Paul, Vanhoucke, Vincent, Vasudevan, Vijay, Viégas, Fernanda, Vinyals, Oriol, Warden, Pete, Wattenberg, Martin, Wicke, Martin, Yu, Yuan, and Zheng, Xiaoqiang · 2015
Closest in time.
Distilling the knowledge in a neural network
Hinton, Geoffrey, Vinyals, Oriol, and Dean, Jeff · 2015
Closest in time.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, Sergey and Szegedy, Christian · 2015
Closest in time.
FitNets and batch normalization
Mahayri, Amjad, Ballas, Nicolas, and Courville, Aaron · 2015
Closest in time.
Never-ending learning
Mitchell, T., Cohen, W., Hruschka, E., Talukdar, P., Betteridge, J., Carlson, A., Dalvi, B., Gardner, M., Kisiel, B., Krishnamurthy, J., Lao, N., Mazaitis, K., Mohamed, T., Nakashole, N., Platanios, E., Ritter, A., Samadi, M., Settles, B., Wang, R., Wijaya, D., Gupta, A., Chen, X., Saparov, A., Greaves, M., and Welling, J · 2015
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Going deeper with convolutions
Szegedy, Christian, Liu, Wei, Jia, Yangqing, Sermanet, Pierre, Reed, Scott, Anguelov, Dragomir, Erhan, Dumitru, Vanhoucke, Vincent, and Rabinovich, Andrew · 2014
Cited alongside, same era.
Closest in time.
Very deep convolutional networks for large-scale image recognition
Simonyan, Karen and Zisserman, Andrew · 2015
Closest in time.