Fetching the paper…
Reading the bibliography…
Learning curves provide insight into the dependence of a learner's generalization performance on the training set size.
E. J. Swift, “Studies in the psychology and physiology of learning,” Am. J. Psychol. , vol. 14, no. 2, pp. 201–251, 1903
1903
Earlier work this paper cites.
O. Lipmann, “Der einfluss der einzelnen wiederholungen auf verschieden starke und verschieden alte associationen,” Z. Psychol. Physiol. Si. , vol. 35, pp. 195–233, 1904
1904
Earlier work this paper cites.
K. W. Spence, “The differential response in animals to stimuli varying within a single dimension.” Psychol. Rev. , vol. 44, no. 5, p. 430, 1937
1937
Earlier work this paper cites.
E. E. Ghiselli, “A comparison of methods of scoring maze and discrimination learning,” J.Gen.Psychol. , vol. 17, no. 1, p. 15, 1937
1937
Earlier work this paper cites.
L. J. Savage, The foundations of statistics . John Wiley, Inc., 1954
1954
Earlier work this paper cites.
A. Ritchie, “Thinking and machines,” Philos. , vol. 32, no. 122, p. 258, 1957
1957
Earlier work this paper cites.
F. Rosenblatt, “The perceptron: a probabilistic model for information storage and organization in the brain.” Psychol. Rev. , vol. 65, no. 6, p. 386, 1958
1958
Earlier work this paper cites.
T. M. Cover, “Geometrical and statistical properties of systems of linear inequalities with applications in pattern recognition,” IEEE Trans. Electron. , vol. 14, no. 3, pp. 326–334, 1965
1965
Earlier work this paper cites.
T. Cover and P. Hart, “Nearest neighbor pattern classification,” IEEE Trans. IT , vol. 13, no. 1, pp. 21–27, 1967
1967
Earlier work this paper cites.
I. J. Good, “On the principle of total evidence,” Br. J. Philos. Sci. , vol. 17, no. 4, pp. 319–321, 1967
1967
Earlier work this paper cites.
G. Hughes, “On the mean accuracy of statistical pattern recognizers,” IEEE Trans. IT , vol. 14, no. 1, pp. 55–63, 1968
1968
Earlier work this paper cites.
T. M. Cover, “Rates of convergence for nearest neighbor procedures,” in Hawaii Int. Conf. on Systems Sci. , vol. 415, 1968
1968
Earlier work this paper cites.
S. Raudys, “On the problems of sample size in pattern recognition (in Russian),” in 2nd All-Union Conf. on Stat. Meth. in Contr. Theory . Nauka, 1970
1970
Earlier work this paper cites.
D. Peterson, “Some convergence properties of a nearest neighbor decision rule,” IEEE Trans. IT , vol. 16, no. 1, pp. 26–31, 1970
1970
Earlier work this paper cites.
L. Kanal and B. Chandrasekaran, “On dimensionality and sample size in statistical pattern classification,” Pattern Recognit. , vol. 3, no. 3, pp. 225–234, 1971
1971
Earlier work this paper cites.
D. H. Foley, “Considerations of Sample and Feature Size,” IEEE T. Inform. Theory , vol. 18, no. 5, pp. 618–626, 1972
1972
Earlier work this paper cites.
M. Osborne, “A modification of veto logic for a committee of threshold logic units and the use of 2-class classifiers for function estimation,” Ph.D. dissertation, Oregon State University, 1975
1975
Earlier work this paper cites.
R. Duin, “On the choice of smoothing parameters for parzen estimators of probability density functions,” IEEE T Comput. , vol. 25, no. 11, pp. 1175–1179, 1976
1976
Earlier work this paper cites.
J. M. Van Campenhout, “On the peaking of the hughes mean recognition accuracy: the resolution of an apparent paradox,” IEEE T Syst. Man Cy. , vol. 8, no. 5, pp. 390–395, 1978
1978
Earlier work this paper cites.
A. K. Jain and W. G. Waller, “On the optimal number of features in the classification of multivariate gaussian data,” Pattern Recognit. , vol. 10, no. 5-6, pp. 365–374, 1978
1978
Earlier work this paper cites.
R. P. W. Duin, “On the accuracy of statistical pattern recognizers,” Ph.D. dissertation, Technische Hogeschool Delft, 1978
1978
Earlier work this paper cites.
L. E. Yelle, “The learning curve: Historical review and comprehensive survey,” Decision sciences , vol. 10, no. 2, pp. 302–328, 1979
1979
Earlier work this paper cites.
S. Raudys and V. Pikelis, “On dimensionality, sample size, classification error, and complexity of classification algorithm in pattern recognition,” IEEE Trans Pattern Anal. Mach. Intell. , no. 3, pp. 242–252, 1980
1980
Earlier work this paper cites.
C. A. Michelli and G. Wahba, “Design problems for optimal surface interpolation,” Approx. Theory Appl. , pp. 329–348, 1981
1981
Earlier work this paper cites.
A. Jain and B. Chandrasekaran, “Dimensionality and Sample Size Considerations in Pattern Recognition Practice,” Handbook of Statistics , vol. 2, pp. 835–855, 1982
1982
Earlier work this paper cites.
A. K. Jain and B. Chandrasekaran, “Dimensionality and sample size considerations in pattern recognition practice,” in Handbook of statistics , P. Krishnaiah and L. Kanal, Eds. Elsevier, 1982, vol. 2, ch. 39, pp. 835–855
1982
Earlier work this paper cites.
V. Vapnik, Estimation of Dependences Based on Empirical Data Berlin . Springer, 1982
1982
Earlier work this paper cites.
B. Efron, “Estimating the error rate of a prediction rule: improvement on cross-validation,” JASA , vol. 78, no. 382, p. 316, 1983
1983
Earlier work this paper cites.
A. K. Jain et al. , “Bootstrap techniques for error estimation,” IEEE Trans Pattern Anal. Mach. Intell. , no. 5, pp. 628–633, 1987
1987
Earlier work this paper cites.
S. Patarnello and P. Carnevali, “Learning networks of neurons with boolean logic,” Europhys. Lett. , vol. 4, no. 4, p. 503, 1987
1987
Earlier work this paper cites.
P. Langley, “Machine Learning as an Experimental Science,” Mach. Learn. , vol. 3, no. 1, pp. 5–8, 1988
1988
Earlier work this paper cites.
S. Ahmad and G. Tesauro, “Study of scaling and generalization in neutral networks,” Neural Netw. , vol. 1, no. 1, p. 3, 1988
1988
Earlier work this paper cites.
F. Vallet et al. , “Linear and nonlinear extension of the pseudo-inverse solution for learning boolean functions,” Europhys. Lett. , vol. 9, no. 4, p. 315, 1989
1989
Earlier work this paper cites.
W. L. Buntine, “A critique of the valiant model.” in IJCAI , 1989, pp. 837–842
1989
Earlier work this paper cites.
W. E. Sarrett and M. J. Pazzani, “Average case analysis of empirical and explanation-based learning algorithms,” Dept. of Information & CS, UC Irvine, Tech. Rep. 89-35, 1989
1989
Earlier work this paper cites.
N. Tishby et al. , “Consistent inference of probabilities in layered networks: Predictions and generalization,” in IJCNN , vol. 2, 1989, pp. 403–409
1989
Earlier work this paper cites.
E. Levin et al. , “A statistical approach to learning and generalization in layered neural networks,” COLT , vol. 78, p. 245, 1989
1989
Earlier work this paper cites.
P. R. Graves, “The total evidence theorem for probability kinematics,” Philos. Sci. , vol. 56, no. 2, pp. 317–324, 1989
1989
Earlier work this paper cites.
L. E. Atlas et al. , “Training connectionist networks with queries and selective sampling,” in NeurIPS , 1990, pp. 566–573
1990
Earlier work this paper cites.
H. Sompolinsky et al. , “Learning from examples in large neural networks,” Phys. Rev. Lett. , vol. 65, no. 13, p. 1683, 1990
1990
Earlier work this paper cites.
M. Opper et al. , “On the ability of the optimal perceptron to generalise,” J. Phys. A , vol. 23, no. 11, p. L581, 1990
1990
Earlier work this paper cites.
D. B. Schwartz et al. , “Exhaustive Learning,” Neural Comput. , vol. 2, no. 3, pp. 374–385, 1990
1990
Earlier work this paper cites.
G. Györgyi, “First-order transition to perfect generalization in a neural network with binary synapses,” Phys. Rev. A , vol. 41, no. 12, pp. 7097–7100, 1990
1990
Earlier work this paper cites.
S. Raudys and A. Jain, “Small sample size effects in statistical pattern recognition: recommendations for practitioners,” IEEE Trans Pattern Anal. Mach. Intell. , vol. 13, no. 3, pp. 252–264, 1991
1991
Earlier work this paper cites.
J. W. Shavlik et al. , “Symbolic and neural learning algorithms: An experimental comparison,” Mach. Learn. , vol. 6, no. 2, pp. 111–143, 1991
1991
Earlier work this paper cites.
D. Cohn and G. Tesauro, “Can neural networks do better than the vapnik-chervonenkis bounds?” in NeurIPS , 1991, pp. 911–917
1991
Earlier work this paper cites.
M. Opper and D. Haussler, “Calculation of the learning curve of bayes optimal classification algorithm for learning a perceptron with noise,” in COLT , vol. 91, 1991, pp. 75–87
1991
Earlier work this paper cites.
S. Amari, “Universal property of learning curves under entropy loss,” in IJCNN , vol. 2, 1992, pp. 368–373 vol.2
1992
Earlier work this paper cites.
S.-i. Amari et al. , “Four types of learning curves,” Neural Comput. , vol. 4, no. 4, pp. 605–618, 1992
1992
Earlier work this paper cites.
H. S. Seung et al. , “Statistical mechanics of learning from examples,” Phys. Rev. A , vol. 45, no. 8, p. 6056, 1992
1992
Earlier work this paper cites.
D. Hansel et al. , “Memorization without generalization in a multilayered neural network,” Europhys. Lett. , vol. 20, no. 5, pp. 471–476, 1992
1992
Earlier work this paper cites.
T. L. Watkin et al. , “The statistical mechanics of learning a rule,” Rev. Mod. Phys. , vol. 65, no. 2, p. 499, 1993
1993
Earlier work this paper cites.
D. Haussler and M. Warmuth, “The probably approximately correct (pac) and other learning models,” in Foundations of Knowledge Acquisition . Springer, 1993, pp. 291–312
1993
Earlier work this paper cites.
S.-i. Amari, “A universal theorem on learning curves,” Neural Netw. , vol. 6, no. 2, pp. 161–166, jan 1993
1993
Earlier work this paper cites.
S.-i. Amari and N. Murata, “Statistical Theory of Learning Curves under Entropic Loss Criterion,” Neural Comput. , vol. 5, no. 1, pp. 140–153, 1993
1993
Earlier work this paper cites.
T. L. H. Watkin et al. , “The Statistical Mechanics of Learning a Rule,” Rev. Mod. Phys. , vol. 65, no. 2, pp. 499–556, 1993
1993
Earlier work this paper cites.
K. Kang et al. , “Generalization in a two-layer neural network,” Phys. Rev. E , vol. 48, no. 6, p. 4805, 1993
1993
Earlier work this paper cites.
H. Schwarze and J. A. Hertz, “Statistical Mechanics of Learning in a Large Committee Machine,” NeurIPS , pp. 523–530, 1993
1993
Earlier work this paper cites.
H. Sompolinsky, “Theoretical issues in learning from examples,” in NEC Research Symposium , 1993, pp. 217–237
1993
Earlier work this paper cites.
M. Biehl and A. Mietzner, “Statistical mechanics of unsupervised learning,” Europhys. Lett. , vol. 24, no. 5, p. 421, 1993
1993
Earlier work this paper cites.
C. Cortes et al. , “Learning curves: Asymptotic values and rate of convergence,” in NeurIPS , 1994, pp. 327–334
1994
Earlier work this paper cites.
E. Perez and L. A. Rendell, “Using multidimensional projection to find relations,” in ML Proc. Elsevier, 1995, pp. 447–455
1995
Earlier work this paper cites.
D. Schuurmans, “Characterizing rational versus exponential learning curves,” in EuroCOLT . Springer, 1995, pp. 272–286
1995
Earlier work this paper cites.
K. Ritter et al. , “Multivariate integration and approximation for random fields satisfying sacks-ylvisaker conditions,” Ann. Appl. Prob. , vol. 5, no. 2, pp. 518–540, 1995
1995
Earlier work this paper cites.
M. Opper, “Statistical Mechanics of Learning : Generalization,” The Handbook of Brain Theory and Neural Networks , p. 20, 1995
1995
Earlier work this paper cites.
H. Seung, “Annealed theories of learning,” CTP-PRSRI Joint Workshop on Theoretical Physics. Singapore, World Scientific , 1995
1995
Earlier work this paper cites.
R. P. W. Duin, “Small sample size generalization,” in SCIA . Springer, 1995, pp. 1–8
1995
Earlier work this paper cites.
M. Skurichina and R. P. Duin, “Stabilizing classifiers for very small sample sizes,” in ICPR , vol. 2. IEEE, 1996, pp. 891–896
1996
Earlier work this paper cites.
R. Kohavi, “Scaling up the accuracy of naive-bayes classifiers: A decision-tree hybrid.” in Kdd , vol. 96, 1996, pp. 202–207
1996
Earlier work this paper cites.
G. H. John and P. Langley, “Static versus dynamic sampling for data mining.” in KDD , vol. 96, 1996, pp. 367–370
1996
Earlier work this paper cites.
L. Devroye et al. , A Probabilistic Theory of Pattern Recognition . New York, NY, USA: Springer, 1996, vol. 31
1996
Cited alongside, same era.
D. Haussler et al. , “Rigorous learning curve bounds from statistical mechanics [longer version],” Mach. Learn. , vol. 25, no. 2-3, pp. 195–236, 1996
1996
Cited alongside, same era.
L. Plaskota, Noisy information and computational complexity . Cambridge University Press, 1996, vol. 95, no. 55
1996
Cited alongside, same era.
K. Ritter, “Almost optimal differentiation using noisy data,” J. of Approx. Theory , vol. 86, no. 3, pp. 293–309, 1996
1996
Cited alongside, same era.
R. P. Duin et al. , “Experiments with a featureless approach to pattern recognition,” Pattern Recognit. Lett. , vol. 18, no. 11-13, pp. 1159–1166, 1997
1997
Cited alongside, same era.
P. D. Grünwald and W. Kotłowski, “Bounds on individual risk for log-loss predictors,” JMLR , vol. 19, pp. 813–816, 2011
2011
Later among the works it cites.
R. O. Duda et al. , Pattern classification . John Wiley & Sons, 2012
2012
Later among the works it cites.
K. P. Murphy, Machine learning: a probabilistic perspective . MIT Press, 2012
2012
Later among the works it cites.
R. L. Figueroa et al. , “Predicting sample size required for classification performance,” BMC Med.Inform.Decis , vol. 12, no. 1, p. 8, 2012
2012
Later among the works it cites.
N. Bertoldi et al. , “Evaluating the learning curve of domain adaptive statistical machine translation systems,” in Workshop on Statistical Machine Translation , 2012, pp. 433–441
2012
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
L. Torgo, “Kernel regression trees,” in ECML , 1997, pp. 118–127
1997
Cited alongside, same era.
N. Mørch et al. , “Nonlinear versus linear models in functional neuroimaging: Learning curves and generalization crossover,” in IPMI . Springer, 1997, pp. 259–270
1997
Cited alongside, same era.
P. Domingos and M. Pazzani, “On the optimality of the simple bayesian classifier under zero-one loss,” Mach. Learn. , vol. 29, no. 2-3, pp. 103–130, 1997
1997
Cited alongside, same era.
C. Harris-Jones and T. L. Haines, “Sample size and misclassification: Is more always better,” AMS Center for Advanced Technologies , 1997
1997
Cited alongside, same era.
——, “Characterizing rational versus exponential learning curves,” journal of computer and system sciences , vol. 55, no. 1, pp. 140–160, 1997
1997
Cited alongside, same era.
M. Opper, “Regression with Gaussian processes: Average case performance,” Theoretical aspects of neural computation: A multidisciplinary perspective , pp. 17–23, 1997
1997
Cited alongside, same era.
G. Vetter et al. , “Phase transitions in learning,” J Mind and Behav. , pp. 335–350, 1997
1997
Cited alongside, same era.
P. Kolachina et al. , “Prediction of Learning Curves in Machine Translation,” in ACL , Jeju Island, Korea, 2012, pp. 22–30
2012
Later among the works it cites.
M. Loog and R. P. W. Duin, “The dipping phenomenon,” in S+SSPR , Hiroshima, Japan, 2012, pp. 310–317
2012
Later among the works it cites.
S. Ben-david et al. , “Minimizing the misclassification error rate using a surrogate convex loss,” ICML , pp. 1863–1870, 2012
2012
Later among the works it cites.
B. Brumen et al. , “Best-fit learning curve model for the c4. 5 algorithm,” Informatica , vol. 25, no. 3, pp. 385–399, 2014
2014
Later among the works it cites.
S. Shalev-Shwartz and S. Ben-David, Understanding machine learning: From theory to algorithms . Cambridge university press, 2014
2014
Later among the works it cites.
X. L. Meng and X. Xie, “I Got More Data, My Model is More Refined, but My Estimator is Getting Worse! Am I Just Dumb?” Econometric Reviews , vol. 33, no. 1-4, pp. 218–250, 2014
2014
Later among the works it cites.
G. M. Weiss and A. Battistin, “Generating well-behaved learning curves: An empirical study,” in ICDATA , 2014
2014
Later among the works it cites.
R. A. Bhat et al. , “Adapting predicate frames for urdu propbanking,” in Workshop on Language Technology for Closely Related Languages and Language Variants , 2014, pp. 47–55
2014
Later among the works it cites.
T. Domhan et al. , “Speeding up automatic hyperparameter optimization of deep neural networks by extrapolation of learning curves,” in IJCAI , 2015
2015
Later among the works it cites.
J. N. van Rijn et al. , “Fast algorithm selection using learning curves,” in LNCS , vol. 9385. Springer, oct 2015, pp. 298–309
2015
Later among the works it cites.
L. Le Gratiet and J. Garnier, “Asymptotic analysis of the learning curve for Gaussian process regression,” Mach. Learn. , vol. 98, no. 3, pp. 407–433, 2015
2015
Later among the works it cites.
K. Konyushkova et al. , “Introducing geometry in active learning for image segmentation,” in CVPR , 2015, pp. 2974–2982
2015
Later among the works it cites.
R. Duin and E. Pekalska, Pattern Recognition: Introduction and Terminology . 37 Steps, 2016
2016
Later among the works it cites.
A. Joulin et al. , “Learning visual features from large weakly supervised data,” in ECCV . Springer, 2016, pp. 67–84
2016
Later among the works it cites.
J. H. Krijthe and M. Loog, “The peaking phenomenon in semi-supervised learning,” in S+SSPR . Springer, 2016, pp. 299–309
2016
Later among the works it cites.
M. Loog et al. , “On measuring and quantifying performance: Error rates, surrogate loss, and an example in semi-supervised learning,” in Handbook of PR and CV . World Scientific, 2016
2016
Later among the works it cites.
M. Loog and Y. Yang, “An empirical investigation into the inconsistency of sequential active learning,” in ICPR , 2016, pp. 210–215
2016
Later among the works it cites.
K. Weiss et al. , “A survey of transfer learning,” J. Big Data , vol. 3, no. 1, p. 9, 2016
2016
Later among the works it cites.
2016
Later among the works it cites.
J. Hestness et al. , “Deep Learning Scaling is Predictable, Empirically,” arXiv:1712.00409 , 2017
2017
Later among the works it cites.
C. Sun et al. , “Revisiting Unreasonable Effectiveness of Data in Deep Learning Era,” in ICCV , 2017, pp. 843–852
2017
Later among the works it cites.
J. Hestness et al. , “Deep learning scaling is predictable, empirically,” arXiv:1712.00409 , 2017
2017
Later among the works it cites.
K. M. Ting et al. , “Defying the gravity of learning curve: a characteristic of nearest neighbour anomaly detectors,” Mach. Learn. , vol. 106, no. 1, pp. 55–91, 2017
2017
Later among the works it cites.
2017
Later among the works it cites.
P. Grünwald and T. van Ommen, “Inconsistency of bayesian inference for misspecified linear models, and a proposal for repairing it,” Bayesian Anal. , vol. 12, no. 4, pp. 1069–1103, 2017
2017
Later among the works it cites.
M. Loog, “Supervised classification: Quite a brief overview,” in Machine Learning Techniques for Space Weather . Elsevier, 2018, pp. 113–145
2018
Later among the works it cites.
D. Sculley et al. , “Winner’s curse? on pace, progress, and empirical rigor,” in ICLR , 2018
2018
Later among the works it cites.
B. Strang et al. , “Don’t rule out simple models prematurely: a large scale benchmark comparing linear and non-linear classifiers in openml,” in IDA . Springer, 2018, pp. 303–315
2018
Later among the works it cites.
D. Mahajan et al. , “Exploring the limits of weakly supervised pretraining,” in ECCV , 2018, pp. 181–196
2018
Later among the works it cites.
A. N. Richter and T. M. Khoshgoftaar, “Learning curve estimation with large imbalanced datasets,” in ICMLA , 2019, pp. 763–768
2019
Later among the works it cites.
P. Nakkiran et al. , “Deep double descent: Where bigger models and more data hurt,” in ICLR , 2019
2019
Later among the works it cites.
A. Zollanvari et al. , “A theoretical analysis of the peaking phenomenon in classification,” Journal of Classification , pp. 1–14, 2019
2019
Later among the works it cites.
M. Belkin et al. , “Reconciling modern machine-learning practice and the classical bias–variance trade-off,” PNAS , vol. 116, no. 32, pp. 15 849–15 854, 2019
2019
Later among the works it cites.
2019
Later among the works it cites.
T. Viering et al. , “Open problem: Monotonicity of learning,” in Conference on Learning Theory , 2019, pp. 3198–3201
2019
Later among the works it cites.
2019
Later among the works it cites.
N. Ipsen and L. K. Hansen, “Phase transition in PCA with missing data: Reduced signal-to-noise ratio, not sample size!” in ICML , 2019, pp. 2951–2960
2019
Later among the works it cites.
2019
Later among the works it cites.
S. Spigler et al. , “A jamming transition from under-to over-parametrization affects generalization in deep learning,” J Phys. A , vol. 52, no. 47, p. 474001, 2019
2019
Later among the works it cites.
2019
Later among the works it cites.
Z. Wang et al. , “Characterizing and avoiding negative transfer,” in CVPR , 2019, pp. 11 293–11 302
2019
Later among the works it cites.
W. M. Kouw and M. Loog, “A review of domain adaptation without target labels,” IEEE Trans Pattern Anal. Mach. Intell. , 2019
2019
Later among the works it cites.
M. Loog et al. , “Minimizers of the empirical risk and risk monotonicity,” in NeurISP , 2019, pp. 7478–7487
2019
Later among the works it cites.
L. Chen et al. , “Multiple descent: Design your own generalization curve,” arXiv:2008.01036 , 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
2020
Later among the works it cites.
O. Bousquet et al. , “A theory of universal learning,” arXiv:2011.04483 , 2020
2020
Later among the works it cites.
J. Kaplan et al. , “Scaling laws for neural language models,” arXiv:2001.08361 , 2020
2020
Later among the works it cites.
M. Loog et al. , “A brief prehistory of double descent,” PNAS , vol. 117, no. 20, pp. 10 625–10 626, 2020
2020
Later among the works it cites.
P. Nakkiran et al. , “Optimal regularization can mitigate double descent,” arXiv:2003.01897 , 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
T. J. Viering et al. , “Making learners (more) monotone,” in IDA . Springer, 2020, pp. 535–547
2020
Later among the works it cites.
Z. Mhammedi and H. Husain, “Risk-monotonicity in statistical learning,” arXiv:2011.14126 , 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
D. Hoiem et al. , “Learning curves for analysis of deep networks,” in ICML . PMLR, 2021, pp. 4287–4296
2021
Closest in time.
2021
Closest in time.
M. Hutter, “Learning curve theory,” arXiv:2102.04074 , 2021
2021
Closest in time.
M. A. Gordon et al. , “Data and parameter scaling laws for neural machine translation,” in EMNLP , 2021, pp. 5915–5922
2021
Closest in time.
2022
Closest in time.
F. Mohr et al. , “LCDB 1.0: An extensive learning curves database for classification tasks,” in ECML , 2022, p. accepted
2022
Closest in time.
X. Zhai et al. , “Scaling vision transformers,” in CVPR , 2022, pp. 12 104–12 113
2022
Closest in time.
2022
Closest in time.
O. J. Bousquet et al. , “Monotone learning,” in COLT , 2022, pp. 842–866
2022
Closest in time.
V. Pestov, “A universally consistent learning rule with a universally monotone error,” JMLR , vol. 23, no. 157, pp. 1–27, 2022
2022
Closest in time.