Fetching the paper…
Reading the bibliography…
In real-world systems, models are frequently updated as more data becomes available, and in addition to achieving high accuracy, the goal is to also maintain a low difference in predictions compared to the base model (i.e.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
Strictly proper scoring rules, prediction, and estimation
Tilmann Gneiting and Adrian E Raftery · 2007
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio · 2010
Earlier work this paper cites.
Do deep nets really need to be deep?
Lei Jimmy Ba and Rich Caruana · 2013
Earlier work this paper cites.
A comparative analysis of offline and online evaluations and discussion of research paper recommender system evaluation
Joeran Beel, Marcel Genzmehr, Stefan Langer, Andreas Nürnberger, and Bela Gipp · 2013
Earlier work this paper cites.
Improving the sensitivity of online controlled experiments by utilizing pre-experiment data
Alex Deng, Ya Xu, Ron Kohavi, and Toby Walker · 2013
Earlier work this paper cites.
Surrogate regret bounds for bipartite ranking via strongly proper losses
Shivani Agarwal · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2015
Earlier work this paper cites.
Unifying distillation and privileged information
David Lopez-Paz, Léon Bottou, Bernhard Schölkopf, and Vladimir Vapnik · 2015
Earlier work this paper cites.
Online batch selection for faster training of neural networks
Ilya Loshchilov and Frank Hutter · 2015
Earlier work this paper cites.
Andrei A Rusu, Sergio Gomez Colmenarejo, Caglar Gulcehre, Guillaume Desjardins, James Kirkpatrick, Razvan Pascanu, Volodymyr Mnih, Koray Kavukcuoglu, and Raia Hadsell · 2015
Earlier work this paper cites.
Ad recommendation systems for life-time value optimization
Georgios Theocharous, Philip S Thomas, and Mohammad Ghavamzadeh · 2015
Earlier work this paper cites.
Estimating numerical error in neural network simulations on graphics processing units
James P Turner and Thomas Nowotny · 2015
Earlier work this paper cites.
Learning using privileged information: similarity control and knowledge transfer
Vladimir Vapnik and Rauf Izmailov · 2015
Cited alongside, same era.
Data-driven metric development for online controlled experiments: Seven lessons learned
Alex Deng and Xiaolin Shi · 2016
Cited alongside, same era.
Launch and iterate: Reducing prediction churn
Mahdi Milani Fard, Quentin Cormier, Kevin Canini, and Maya Gupta · 2016
Cited alongside, same era.
Satisfying real-world goals with dataset constraints
Gabriel Goh, Andrew Cotter, Maya Gupta, and Michael P Friedlander · 2016
Cited alongside, same era.
Harnessing deep neural networks with logic rules
Zhiting Hu, Xuezhe Ma, Zhengzhong Liu, Eduard Hovy, and Eric Xing · 2016
Cited alongside, same era.
Distillation as a defense to adversarial perturbations against deep neural networks
Model compression via distillation and quantization
Antonio Polino, Razvan Pascanu, and Dan Alistarh · 2018
Later among the works it cites.
How does batch normalization help optimization?
Shibani Santurkar, Dimitris Tsipras, Andrew Ilyas, and Aleksander Madry · 2018
Later among the works it cites.
On warm-starting neural network training
Jordan T Ash and Ryan P Adams · 2019
Later among the works it cites.
Bin Dong, Jikai Hou, Yiping Lu, and Zhihua Zhang · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Nicolas Papernot, Patrick McDaniel, Xi Wu, Somesh Jha, and Ananthram Swami · 2016
Cited alongside, same era.
Rethinking the inception architecture for computer vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna · 2016
Cited alongside, same era.
Composite multiclass losses
Robert C Williamson, Elodie Vernet, and Mark D Reid · 2016
Cited alongside, same era.
Domain adaptation of dnn acoustic models using knowledge distillation
Taichi Asami, Ryo Masumura, Yoshikazu Yamaguchi, Hirokazu Masataki, and Yushi Aono · 2017
Cited alongside, same era.
Learning from noisy labels with distillation
Yuncheng Li, Jianchao Yang, Yale Song, Liangliang Cao, Jiebo Luo, and Li-Jia Li · 2017
Cited alongside, same era.
Visual relationship detection with internal and external linguistic knowledge distillation
Ruichi Yu, Ang Li, Vlad I Morariu, and Larry S Davis · 2017
Cited alongside, same era.
mixup: Beyond empirical risk minimization
Hongyi Zhang, Moustapha Cisse, Yann N Dauphin, and David Lopez-Paz · 2017
Cited alongside, same era.
Dylan J Foster, Spencer Greenberg, Satyen Kale, Haipeng Luo, Mehryar Mohri, and Karthik Sridharan · 2019
Later among the works it cites.
Towards understanding knowledge distillation
Mary Phuong and Christoph Lampert · 2019
Later among the works it cites.
A survey on image data augmentation for deep learning
Connor Shorten and Taghi M Khoshgoftaar · 2019
Later among the works it cites.
Keras documentation: Text classification with transformer, 2020
Keras · 2020
Later among the works it cites.
Why distillation helps: a statistical perspective
Aditya Krishna Menon, Ankit Singh Rawat, Sashank J Reddi, Seungyeon Kim, and Sanjiv Kumar · 2020
Later among the works it cites.
Self-distillation amplifies regularization in hilbert space
Hossein Mobahi, Mehrdad Farajtabar, and Peter L Bartlett · 2020
Later among the works it cites.
Locally adaptive label smoothing for predictive churn
Dara Bahri and Heinrich Jiang · 2021
Closest in time.
On the reproducibility of neural network predictions
Srinadh Bhojanapalli, Kimberly Wilber, Andreas Veit, Ankit Singh Rawat, Seungyeon Kim, Aditya Menon, and Sanjiv Kumar · 2021
Closest in time.