Fetching the paper…
Reading the bibliography…
For industrial-scale advertising systems, prediction of ad click-through rate (CTR) is a central problem.
Deep Ensembles: A Loss Landscape Perspective
Stanislav Fort, Huiyi Hu, and Balaji Lakshminarayanan. 2020 · 1912
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams. 1992 · 1992
Earlier work this paper cites.
Multitask Learning
Rich Caruana. 1997 · 1997
Earlier work this paper cites.
Beating the hold-out: Bounds for k-fold and progressive cross-validation. In COLT
Avrim Blum, Adam Kalai, and John Langford. 1999 · 1999
Earlier work this paper cites.
Ensemble methods in machine learning
T. G. Dietterich. 2000 · 2000
Earlier work this paper cites.
Learning to Rank Using Gradient Descent. In ICML
Chris Burges, Tal Shaked, Erin Renshaw, Ari Lazier, Matt Deeds, Nicole Hamilton, and Greg Hullender. 2005 · 2005
Earlier work this paper cites.
A Schur–Newton Method for the Matrix \ \backslash boldmath p th Root and its Inverse
Chun-Hua Guo and Nicholas J Higham. 2006 · 2006
Earlier work this paper cites.
Numerical Optimization
Jorge Nocedal and Stephen J. Wright. 2006 · 2006
Earlier work this paper cites.
Simple, Robust, Scalable Semi-supervised Learning via Expectation Regularization. In ICML
Gideon S Mann and Andrew McCallum. 2007 · 2007
Earlier work this paper cites.
Exploration scavenging. In ICML
John Langford, Alexander Strehl, and Jennifer Wortman. 2008 · 2008
Earlier work this paper cites.
Curriculum Learning. In ICML
Yoshua Bengio, Jérôme Louradour, Ronan Collobert, and Jason Weston. 2009 · 2009
Earlier work this paper cites.
From Ranknet to Lambdarank to LambdaMart: An overview
Christopher JC Burges. 2010 · 2010
Earlier work this paper cites.
Adaptive bound optimization for online convex optimization
H Brendan McMahan and Matthew Streeter. 2010 · 2010
Earlier work this paper cites.
Combined regression and ranking. In In KDD’10
D. Sculley. 2010 · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer. 2011 · 2011
Earlier work this paper cites.
Counterfactual Reasoning and Learning Systems: The Example of Computational Advertising
Léon Bottou, Jonas Peters, Joaquin Quiñonero-Candela, Denis X. Charles, D. Max Chickering, Elon Portugaly, Dipankar Ray, Patrice Simard, and Ed Snelson. 2013 · 2013
Earlier work this paper cites.
Ad click prediction: a view from the trenches. In SIGKDD
H Brendan McMahan, Gary Holt, D. Sculley, Michael Young, Dietmar Ebner, Julian Grady, Lan Nie, Todd Phillips, Eugene Davydov, Daniel Golovin, et al · 2013
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning. In ICML
Ilya Sutskever, James Martens, George Dahl, and Geoffrey Hinton. 2013 · 2013
Earlier work this paper cites.
Local case-control sampling: Efficient subsampling in imbalanced data sets
William Fithian and Trevor Hastie. 2014 · 2014
Earlier work this paper cites.
Practical Lessons from Predicting Clicks on Ads at Facebook
Xinran He, Junfeng Pan, Ou Jin, Tianbing Xu, Bo Liu, Tao Xu, Yanxin Shi, Antoine Atallah, Ralf Herbrich, Stuart Bowers, and Joaquin Quiñonero Candela. 2014 · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba. 2014 · 2014
Earlier work this paper cites.
Machine learning: The high interest credit card of technical debt
David Sculley, Gary Holt, Daniel Golovin, Eugene Davydov, Todd Phillips, Dietmar Ebner, Vinay Chaudhary, and Michael Young. 2014 · 2014
Earlier work this paper cites.
The VCG auction in theory and practice
Hal R Varian and Christopher Harris. 2014 · 2014
Earlier work this paper cites.
Weight uncertainty in neural network. In International conference on machine learning . PMLR, 1613–1622
Charles Blundell, Julien Cornebise, Koray Kavukcuoglu, and Daan Wierstra. 2015 · 2015
Earlier work this paper cites.
Net2net: Accelerating learning via knowledge transfer
Tianqi Chen, Ian Goodfellow, and Jonathon Shlens. 2015 · 2015
Earlier work this paper cites.
Distilling the Knowledge in a Neural Network. In NIPS Deep Learning and Representation Learning Workshop
Geoffrey Hinton, Oriol Vinyals, and Jeffrey Dean. 2015 · 2015
Cited alongside, same era.
Batch Learning from Logged Bandit Feedback through Counterfactual Risk Minimization
Adith Swaminathan and Thorsten Joachims. 2015 · 2015
Cited alongside, same era.
Improving deep neural networks using softplus units. In IJCNN
Hao Zheng, Zhanlei Yang, Wenju Liu, Jizhong Liang, and Yanpeng Li. 2015 · 2015
Cited alongside, same era.
Gaussian error linear units (gelus)
Dan Hendrycks and Kevin Gimpel. 2016 · 2016
Cited alongside, same era.
Large-scale Validation of Counterfactual Learning Methods: A Test-Bed
Damien Lefortier, Adith Swaminathan, Xiaotao Gu, Thorsten Joachims, and M. de Rijke. 2016 · 2016
Cited alongside, same era.
Regularized evolution for image classifier architecture search. In AAAI
Esteban Real, Alok Aggarwal, Yanping Huang, and Quoc V Le. 2019 · 2019
Later among the works it cites.
Large batch optimization for deep learning: Training bert in 76 minutes
Y. You, J. Li, S. Reddi, J. Hseu, S. Kumar, S. Bhojanapalli, X. Song, J. Demmel, K. Keutzer, and C. Hsieh. 2019 · 2019
Later among the works it cites.
Disentangling adaptive gradient methods from learning rates
Naman Agarwal, Rohan Anil, Elad Hazan, Tomer Koren, and Cyril Zhang. 2020 · 2020
Later among the works it cites.
Towards understanding ensemble, knowledge distillation and self-distillation in deep learning
Zeyuan Allen-Zhu and Yuanzhi Li. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning to Rank with Selection Bias in Personal Search. In ACM SIGIR
Xuanhui Wang, Michael Bendersky, Donald Metzler, and Marc Najork. 2016 · 2016
Cited alongside, same era.
Critical learning periods in deep neural networks
Alessandro Achille, Matteo Rovere, and Stefano Soatto. 2017 · 2017
Cited alongside, same era.
Continuously differentiable exponential linear units
Jonathan T Barron. 2017 · 2017
Cited alongside, same era.
Sponsored Search Auctions with Rich Ads
Ruggiero Cavallo, Prabhakar Krishnamurthy, Maxim Sviridenko, and Christopher A. Wilkens. 2017 · 2017
Cited alongside, same era.
Scalable Learning of Non-Decomposable Objectives. In AIStats
Elad Eban, Mariano Schain, Alan Mackey, Ariel Gordon, Rif A Saurous, and Gal Elidan. 2017 · 2017
Cited alongside, same era.
Unbiased Learning-to-Rank with Biased Feedback. In WSDM
Thorsten Joachims, Adith Swaminathan, and Tobias Schnabel. 2017 · 2017
Cited alongside, same era.
Model ensemble for click prediction in bing search ads. In WWW
Xiaoliang Ling, Weiwei Deng, Chen Gu, Hucheng Zhou, Cui Li, and Feng Sun. 2017 · 2017
Cited alongside, same era.
Rohan Anil, Vineet Gupta, Tomer Koren, Kevin Regan, and Yoram Singer. 2020 · 2020
Later among the works it cites.
Can weight sharing outperform random architecture search? an investigation with tunas. In CVPR
Gabriel Bender, Hanxiao Liu, Bo Chen, Grace Chu, Shuyang Cheng, Pieter-Jan Kindermans, and Quoc V Le. 2020 · 2020
Later among the works it cites.
Zhe Chen, Yuyan Wang, Dong Lin, Derek Cheng, Lichan Hong, Ed Chi, and Claire Cui. 2020 · 2020
Later among the works it cites.
Underspecification Presents Challenges for Credibility in Modern Machine Learning
A. D’Amour, K. Heller, D. Moldovan, B. Adlam, B. Alipanahi, A. Beutel, C. Chen, J. Deaton, J. Eisenstein, M. D. Hoffman, F. Hormozdiari, N. Houlsby, S. Hou, G. Jerfel, A. Karthikesalingam, M. Lucic, Y. Ma, C. McLean, D. Mincu, A. Mitani, A. Montanari, Z. Nado, V. Natarajan, C. Nielson, T. F. Osborne, R. Raman, K. Ramasamy, R. Sayres, J. Schrouff, M. Seneviratne, S. Sequeira, H. Suresh, V. Veitch, M. Vladymyrov, X. Wang, K. Webster, S. Yadlowsky, T. Yun, X. Zhai, and D. Sculley. 2020 · 2020
Later among the works it cites.
Analyzing the role of model uncertainty for electronic health records. In CHIL
Michael W Dusenberry, Dustin Tran, Edward Choi, Jonas Kemp, Jeremy Nixon, Ghassen Jerfel, Katherine Heller, and Andrew M Dai. 2020 · 2020
Later among the works it cites.
Linear mode connectivity and the lottery ticket hypothesis. In International Conference on Machine Learning
Jonathan Frankle, Gintare Karolina Dziugaite, Daniel Roy, and Michael Carbin. 2020 · 2020
Later among the works it cites.
A domain-specific supercomputer for training deep neural networks
Norman P Jouppi, Doe Hyun Yoon, George Kurian, Sheng Li, Nishant Patil, James Laudon, Cliff Young, and David Patterson. 2020 · 2020
Later among the works it cites.
When ensembling smaller models is more efficient than single large models
Dan Kondratyuk, Mingxing Tan, Matthew Brown, and Boqing Gong. 2020 · 2020
Later among the works it cites.
On power laws in deep ensembles
Ekaterina Lobacheva, Nadezhda Chirkova, Maxim Kodryan, and Dmitry Vetrov. 2020 · 2020
Later among the works it cites.
Anti-Distillation: Improving reproducibility of deep networks
Gil I Shamir and Lorenzo Coviello. 2020a · 2020
Later among the works it cites.
Smooth activations and reproducibility in deep networks
Gil I Shamir, Dong Lin, and Lorenzo Coviello. 2020 · 2020
Later among the works it cites.
Automatic cross-replica sharding of weight update in data-parallel training
Yuanzhong Xu, HyoukJoong Lee, Dehao Chen, Hongjun Choi, Blake Hechtman, and Shibo Wang. 2020 · 2020
Later among the works it cites.
Rezero is all you need: Fast convergence at large depth. In UAI
Thomas Bachlechner, Bodhisattwa Prasad Majumder, Henry Mao, Gary Cottrell, and Julian McAuley. 2021 · 2021
Later among the works it cites.
On the Reproducibility of Neural Network Predictions
Srinadh Bhojanapalli, Kimberly Jenney Wilber, Andreas Veit, Ankit Singh Rawat, Seungyeon Kim, Aditya Krishna Menon, and Sanjiv Kumar. 2021 · 2021
Later among the works it cites.
Synthesizing Irreproducibility in Deep Networks
Robert R Snapp and Gil I Shamir. 2021 · 2021
Later among the works it cites.
On Nondeterminism and Instability in Neural Network Optimization
Cecilia Summers and Michael J. Dinneen. 2021 · 2021
Later among the works it cites.
Wisdom of Committees: An Overlooked Approach To Faster and More Accurate Models
Xiaofang Wang, Dan Kondratyuk, Eric Christiansen, Kris M. Kitani, Yair Alon, and Elad Eban. 2021a · 2021
Later among the works it cites.
Dropout Prediction Variation Estimation Using Neuron Activation Strength
Haichao Yu, Zhe Chen, Dong Lin, Gil Shamir, and Jie Han. 2021 · 2021
Later among the works it cites.
Randomness in neural network training: Characterizing the impact of tooling
Donglin Zhuang, Xingyao Zhang, Shuaiwen Leon Song, and Sara Hooker. 2021 · 2021
Later among the works it cites.
Reproducibility in optimization: Theoretical framework and Limits
Kwangjun Ahn, Prateek Jain, Ziwei Ji, Satyen Kale, Praneeth Netrapalli, and Gil I. Shamir. 2022 · 2022
Closest in time.