Fetching the paper…
Reading the bibliography…
Cascades are a classical strategy to enable inference cost to vary adaptively across samples, wherein a sequence of classifiers are invoked in turn.
Roy Schwartz, Jesse Dodge, Noah A. Smith, and Oren Etzioni · 1907
Earlier work this paper cites.
Ix. on the problem of the most efficient tests of statistical hypotheses
Jerzy Neyman and Egon Sharpe Pearson · 1933
Earlier work this paper cites.
On optimum recognition error and reject tradeoff
C. Chow · 1970
Earlier work this paper cites.
Optimal brain damage
Yann LeCun, John Denker, and Sara Solla · 1989
Earlier work this paper cites.
Using anytime algorithms in intelligent systems
Shlomo Zilberstein · 1996
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
The foundations of cost-sensitive learning
Charles Elkan · 2001
Earlier work this paper cites.
Rapid object detection using a boosted cascade of simple features
P. Viola and M. Jones · 2001
Earlier work this paper cites.
Parallel support vector machines: The cascade svm
Hans Graf, Eric Cosatto, Leon Bottou, Igor Dourdanovic, and Vladimir Vapnik · 2004
Earlier work this paper cites.
Cascade ensembles
N. García-Pedrajas, D. Ortiz-Boyer, R. del Castillo-Gomariz, and C. Hervás-Martínez · 2005
Earlier work this paper cites.
Predicting good probabilities with supervised learning
Alexandru Niculescu-Mizil and Rich Caruana · 2005
Earlier work this paper cites.
Model compression
Cristian Bucilǎ, Rich Caruana, and Alexandru Niculescu-Mizil · 2006
Earlier work this paper cites.
Cost curves: An improved method for visualizing classifier performance
Chris Drummond and Robert C Holte · 2006
Earlier work this paper cites.
On the design of cascades of boosted ensembles for face detection
S. Charles Brubaker, Jianxin Wu, Jie Sun, Matthew D. Mullin, and James M. Rehg · 2007
Earlier work this paper cites.
Classification with a reject option using a hinge loss
Peter L. Bartlett and Marten H. Wegkamp · 2008
Earlier work this paper cites.
A tutorial on conformal prediction
Glenn Shafer and Vladimir Vovk · 2008
Earlier work this paper cites.
Multiple-instance pruning for learning efficient cascade detectors
Cha Zhang and Paul Viola · 2008
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Dataset shift in machine learning
Joaquin Quinonero-Candela, Masashi Sugiyama, Anton Schwaighofer, and Neil D. Lawrence · 2009
Earlier work this paper cites.
Boosting classifier cascades
Mohammad J. Saberian and Nuno Vasconcelos · 2010
Earlier work this paper cites.
Classifier cascade for minimizing feature evaluation cost
Minmin Chen, Zhixiang Xu, Kilian Weinberger, Olivier Chapelle, and Dor Kedem · 2012
Earlier work this paper cites.
Class probability estimates are unreliable for imbalanced data (and how to fix them)
Byron C. Wallace and Issa J. Dahabreh · 2012
Earlier work this paper cites.
Supervised sequential classification under budget constraints
Kirill Trapeznikov and Venkatesh Saligrama · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Boosting algorithms for detector cascade learning
Mohammad Saberian and Nuno Vasconcelos · 2014
Earlier work this paper cites.
Learning with deep cascades
Giulia DeSalvo, Mehryar Mohri, and Umar Syed · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey E. Hinton, Oriol Vinyals, and Jeffrey Dean · 2015
Earlier work this paper cites.
Sparse convolutional neural networks
Baoyuan Liu, Min Wang, Hassan Foroosh, Marshall Tappen, and Marianna Penksy · 2015
Earlier work this paper cites.
Deep neural networks are easily fooled: High confidence predictions for unrecognizable images
Anh Nguyen, Jason Yosinski, and Jeff Clune · 2015
Earlier work this paper cites.
An analysis of deep neural network models for practical applications
Alfredo Canziani, Adam Paszke, and Eugenio Culurciello · 2016
Cited alongside, same era.
Boosting with abstention
Corinna Cortes, Giulia DeSalvo, and Mehryar Mohri · 2016
Cited alongside, same era.
Deep networks with stochastic depth
Gao Huang, Yu Sun, Zhuang Liu, Daniel Sedra, and Kilian Q. Weinberger · 2016
Cited alongside, same era.
Conditional deep learning for energy-efficient and enhanced pattern recognition
Priyadarshini Panda, Abhronil Sengupta, and Kaushik Roy · 2016
Cited alongside, same era.
Branchynet: Fast inference via early exiting from deep neural networks
Surat Teerapittayanon, Bradley McDanel, and H. T. Kung · 2016
Cited alongside, same era.
Adaptive neural networks for fast test-time prediction
Tolga Bolukbasi, Joseph Wang, Ofer Dekel, and Venkatesh Saligrama · 2017
Beyond synthetic noise: Deep learning on controlled noisy labels
Lu Jiang, Di Huang, Mason Liu, and Weilong Yang · 2020
Later among the works it cites.
FastBERT: a self-distilling bert with adaptive inference time
Weijie Liu, Peng Zhou, Zhe Zhao, Zhiruo Wang, Haotang Deng, and Qi Ju · 2020
Later among the works it cites.
Consistent estimators for learning to defer to an expert
Hussein Mozannar and David Sontag · 2020
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu · 2020
Later among the works it cites.
Why should we add early exits to neural networks?
Simone Scardapane, Michele Scarpiniti, Enzo Baccarelli, and Aurelio Uncini · 2020
Later among the works it cites.
The right tool for the job: Matching model and instance complexities
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
On calibration of modern neural networks
Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q. Weinberger · 2017
Cited alongside, same era.
A baseline for detecting misclassified and out-of-distribution examples in neural networks
Dan Hendrycks and Kevin Gimpel · 2017
Cited alongside, same era.
Quantized neural networks: Training neural networks with low precision weights and activations
Itay Hubara, Matthieu Courbariaux, Daniel Soudry, Ran El-Yaniv, and Yoshua Bengio · 2017
Cited alongside, same era.
What uncertainties do we need in bayesian deep learning for computer vision?
Alex Kendall and Yarin Gal · 2017
Cited alongside, same era.
Fractalnet: Ultra-deep neural networks without residuals
Gustav Larsson, Michael Maire, and Gregory Shakhnarovich · 2017
Cited alongside, same era.
Multi-scale dense networks for resource efficient image classification
Gao Huang, Danlu Chen, Tianhong Li, Felix Wu, Laurens van der Maaten, and Kilian Weinberger · 2018
Cited alongside, same era.
Roy Schwartz, Gabriel Stanovsky, Swabha Swayamdipta, Jesse Dodge, and Noah A. Smith · 2020
Later among the works it cites.
Any-width networks
T. Vu, M. Eder, T. Price, and J. Frahm · 2020
Later among the works it cites.
DeeBERT: Dynamic early exiting for accelerating BERT inference
Ji Xin, Raphael Tang, Jaejun Lee, Yaoliang Yu, and Jimmy Lin · 2020
Later among the works it cites.
BERT loses patience: Fast and robust inference with early exit
Wangchunshu Zhou, Canwen Xu, Tao Ge, Julian McAuley, Ke Xu, and Furu Wei · 2020
Later among the works it cites.
Classification with rejection based on cost-sensitive classification
Nontawat Charoenphakdee, Zhenghang Cui, Yivan Zhang, and Masashi Sugiyama · 2021
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Later among the works it cites.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
William Fedus, Barret Zoph, and Noam Shazeer · 2021
Later among the works it cites.
Selective classification via one-sided prediction
Aditya Gangrade, Anil Kag, and Venkatesh Saligrama · 2021
Later among the works it cites.
Dynamic neural networks: A survey
Y. Han, G. Huang, S. Song, L. Yang, H. Wang, and Y. Wang · 2021
Later among the works it cites.
Machine learning with a reject option: A survey
Kilian Hendrickx, Lorenzo Perini, Dries Van der Plas, Wannes Meert, and Jesse Davis · 2021
Later among the works it cites.
Aleatoric and epistemic uncertainty in machine learning: an introduction to concepts and methods
Eyke Hüllermeier and Willem Waegeman · 2021
Later among the works it cites.
When in doubt, summon the titans: Efficient inference with large models
Ankit Singh Rawat, Manzil Zaheer, Aditya Krishna Menon, Amr Ahmed, and Sanjiv Kumar · 2021
Later among the works it cites.
From generalized zero-shot learning to long-tail with class descriptors
Dvir Samuel, Yuval Atzmon, and Gal Chechik · 2021
Later among the works it cites.
Generalizing consistent multi-class classification with rejection to be compatible with arbitrary losses
Yuzhou Cao, Tianchi Cai, Lei Feng, Lihong Gu, Jinjie GU, Bo An, Gang Niu, and Masashi Sugiyama · 2022
Later among the works it cites.
Sample efficient learning of predictors that complement humans
Mohammad-Amin Charusaie, Hussein Mozannar, David Sontag, and Samira Samadi · 2022
Later among the works it cites.
David Dohan, Winnie Xu, Aitor Lewkowycz, Jacob Austin, David Bieber, Raphael Gontijo Lopes, Yuhuai Wu, Henryk Michalewski, Rif A. Saurous, Jascha Sohl-dickstein, Kevin Murphy, and Charles Sutton · 2022
Later among the works it cites.
Babybear: Cheap inference triage for expensive language models, 2022
Leila Khalili, Yao You, and John Bohannon · 2022
Later among the works it cites.
TangoBERT: Reducing inference cost by using cascaded architecture, 2022
Jonathan Mamou, Oren Pereg, Moshe Wasserblat, and Roy Schwartz · 2022
Later among the works it cites.
Post-hoc estimators for learning to defer to an expert
Harikrishna Narasimhan, Wittawat Jitkrittum, Aditya Krishna Menon, Ankit Singh Rawat, and Sanjiv Kumar · 2022
Later among the works it cites.
Predicting on the edge: Identifying where a larger model does better, 2022
Taman Narayan, Heinrich Jiang, Sen Zhao, and Sanjiv Kumar · 2022
Later among the works it cites.
Model cascading: Towards jointly improving efficiency and accuracy of nlp systems
Neeraj Varshney and Chitta Baral · 2022
Later among the works it cites.
Wisdom of committees: An overlooked approach to faster and more accurate models
Xiaofang Wang, Dan Kondratyuk, Eric Christiansen, Kris M. Kitani, Yair Movshovitz-Attias, and Elad Eban · 2022
Later among the works it cites.
Efficient edge inference by selective query
Anil Kag, Igor Fedorov, Aditya Gangrade, Paul Whatmough, and Venkatesh Saligrama · 2023
Closest in time.