Fetching the paper…
Reading the bibliography…
In many real-world deployments of machine learning systems, data arrive piecemeal.
A simple weight decay can improve generalization
Anders Krogh and John A Hertz · 1992
Earlier work this paper cites.
Flat minima
Sepp Hochreiter and J Schmidhuber · 1997
Earlier work this paper cites.
Alpha seeding for support vector machines
Dennis DeCoste and Kiri Wagstaff · 2000
Earlier work this paper cites.
Greedy layer-wise training of deep networks
Yoshua Bengio, Pascal Lamblin, Dan Popovici, and Hugo Larochelle · 2007
Earlier work this paper cites.
Classifier ensembles for detecting concept change in streaming data: Overview and perspectives
Ludmila I Kuncheva · 2008
Earlier work this paper cites.
Why does unsupervised pre-training help deep learning?
Dumitru Erhan, Yoshua Bengio, Aaron Courville, Pierre-Antoine Manzagol, Pascal Vincent, and Samy Bengio · 2010
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio · 2010
Earlier work this paper cites.
StreamRec: a real-time recommender system
Badrish Chandramouli, Justin J Levandoski, Ahmed Eldawy, and Mohamed F Mokbel · 2011
Earlier work this paper cites.
Deep learning of representations for unsupervised and transfer learning
Yoshua Bengio · 2011
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
Ilya Sutskever, James Martens, George Dahl, and Geoffrey Hinton · 2013
Earlier work this paper cites.
Practical lessons from predicting clicks on ads at Facebook
Xinran He, Junfeng Pan, Ou Jin, Tianbing Xu, Bo Liu, Tao Xu, Yanxin Shi, Antoine Atallah, Ralf Herbrich, Stuart Bowers, et al · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Intriguing properties of neural networks
Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus · 2014
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Warm start for parameter selection of linear classifiers
Bo-Yu Chu, Chia-Hua Ho, Cheng-Hao Tsai, Chieh-Yen Lin, and Chih-Jen Lin · 2015
Earlier work this paper cites.
Training very deep networks
Rupesh K Srivastava, Klaus Greff, and Jürgen Schmidhuber · 2015
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on ImageNet classification
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Cited alongside, same era.
Matching networks for one shot learning
Oriol Vinyals, Charles Blundell, Timothy Lillicrap, and Daan Wierstra · 2016
Cited alongside, same era.
Unsupervised domain adaptation using approximate label matching
Jordan T Ash, Robert E Schapire, and Barbara E Engelhardt · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2016
Cited alongside, same era.
Active learning for convolutional neural networks: A core-set approach
Ozan Sener and Silvio Savarese · 2018
Later among the works it cites.
Algorithmic regularization in over-parameterized matrix sensing and neural networks with quadratic activations
Yuanzhi Li, Tengyu Ma, and Hongyang Zhang · 2018
Later among the works it cites.
Critical learning periods in deep networks
Alessandro Achille, Matteo Rovere, and Stefano Soatto · 2018
Later among the works it cites.
Reptile: a scalable metalearning algorithm
Alex Nichol and John Schulman · 2018
Later among the works it cites.
Model-based active exploration
Pranav Shyam, Wojciech Jaśkowski, and Faustino Gomez · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lili Mou, Zhao Meng, Rui Yan, Ge Li, Yan Xu, Lu Zhang, and Zhi Jin · 2016
Cited alongside, same era.
Sgdr: Stochastic gradient descent with warm restarts
Ilya Loshchilov and Frank Hutter · 2016
Cited alongside, same era.
Prototypical networks for few-shot learning
Jake Snell, Kevin Swersky, and Richard Zemel · 2017
Cited alongside, same era.
Model-agnostic meta-learning for fast adaptation of deep networks
Chelsea Finn, Pieter Abbeel, and Sergey Levine · 2017
Cited alongside, same era.
Implicit regularization in matrix factorization
Suriya Gunasekar, Blake E Woodworth, Srinadh Bhojanapalli, Behnam Neyshabur, and Nati Srebro · 2017
Cited alongside, same era.
Regularizing neural networks by penalizing confident output distributions
Gabriel Pereyra, George Tucker, Jan Chorowski, Łukasz Kaiser, and Geoffrey Hinton · 2017
Cited alongside, same era.
Improving efficiency of SVM k k -fold cross-validation by alpha seeding
Zeyi Wen, Bin Li, Ramamohanarao Kotagiri, Jian Chen, Yawen Chen, and Rui Zhang · 2017
Cited alongside, same era.
Colin Wei, Jason D Lee, Qiang Liu, and Tengyu Ma · 2018
Later among the works it cites.
Visualizing the loss landscape of neural nets
Hao Li, Zheng Xu, Gavin Taylor, Christoph Studer, and Tom Goldstein · 2018
Later among the works it cites.
Sensitivity and generalization in neural networks: an empirical study
Roman Novak, Yasaman Bahri, Daniel A Abolafia, Jeffrey Pennington, and Jascha Sohl-Dickstein · 2018
Later among the works it cites.
Energy and policy considerations for deep learning in NLP
Emma Strubell, Ananya Ganesh, and Andrew McCallum · 2019
Closest in time.
Deep batch active learning by diverse, uncertain gradient lower bounds
Jordan T Ash, Chicheng Zhang, Akshay Krishnamurthy, John Langford, and Alekh Agarwal · 2019
Closest in time.
https://github.com/ej0cl6/deep-active-learning , 2017–2019
Rostamiz · 2019
Closest in time.
https://github.com/ej0cl6/deep-active-learning , 2018–2019
Kuan-Hao Huang · 2019
Closest in time.
Non-vacuous generalization bounds at the ImageNet scale: a PAC-Bayesian compression approach
Wenda Zhou, Victor Veitch, Morgane Austern, Ryan P Adams, and Peter Orbanz · 2019
Closest in time.
An analysis of pre-training on object detection
Hengduo Li, Bharat Singh, Mahyar Najibi, Zuxuan Wu, and Larry S Davis · 2019
Closest in time.
Gradient surgery for multi-task learning
Tianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine, Karol Hausman, and Chelsea Finn · 2020
Closest in time.