Fetching the paper…
Reading the bibliography…
Learning curves are a concept from social sciences that has been adopted in the context of machine learning to assess the performance of a learning algorithm with respect to a certain resource, e.g., the number of training examples or the number of training iterations.
1903
Earlier work this paper cites.
Waltz M, Fu K (1965) A heuristic approach to reinforcement learning control systems. IEEE Transactions on Automatic Control 10(4):390–398
1965
Earlier work this paper cites.
Hughes GF (1968) On the mean accuracy of statistical pattern recognizers. IEEE Trans Inf Theory 14(1):55–63
1968
Earlier work this paper cites.
Vallet F, Cailton JG, Refregier P (1989) Linear and nonlinear extension of the pseudo-inverse solution for learning boolean functions. EPL (Europhysics Letters) 9(4):315
1989
Earlier work this paper cites.
Murata N, Yoshizawa S, Amari S (1992) Learning curves, model selection and complexity of neural networks. In: Advances in Neural Information Processing Systems 5. pp 607–614
1992
Earlier work this paper cites.
Seung HS, Sompolinsky H, Tishby N (1992) Statistical mechanics of learning from examples. Physical Review A 45(8):6056
1992
Earlier work this paper cites.
Amari S, Murata N (1993) Statistical theory of learning curves under entropic loss criterion. Neural Computation 5(1):140–153
1993
Earlier work this paper cites.
Cortes C, Jackel LD, Solla SA, et al (1993) Learning curves: Asymptotic values and rate of convergence. In: Advances in Neural Information Processing Systems 6. pp 327–334
1993
Earlier work this paper cites.
Cortes C, Jackel LD, Chiang W (1994) Limits in learning machine accuracy imposed by data quality. In: Advances in Neural Information Processing Systems 7. pp 239–246
1994
Earlier work this paper cites.
Bishop C (1995) Regularization and complexity control in feed-forward networks. In: Proceedings International Conference on Artificial Neural Networks ICANN’95, pp 141–148
1995
Earlier work this paper cites.
John GH, Langley P (1996) Static versus dynamic sampling for data mining. In: Proceedings of the Second International Conference on Knowledge Discovery and Data Mining (KDD-96). pp 367–370
1996
Earlier work this paper cites.
Mørch NJS, Hansen LK, Strother SC, et al (1997) Nonlinear versus linear models in functional neuroimaging: Learning curves and generalization crossover. In: Information Processing in Medical Imaging, 15th International Conference, IPMI’97. pp 259–270
1997
Earlier work this paper cites.
Fine T, Mukherjee S (1999) Parameter convergence and learning curves for neural networks. Neural Comput 11(3):747–769
1999
Earlier work this paper cites.
Frey LJ, Fisher DH (1999) Modeling decision tree performance with the power law. In: Proceedings of the Seventh International Workshop on Artificial Intelligence and Statistics, AISTATS 1999
1999
Earlier work this paper cites.
Provost FJ, Jensen DD, Oates T (1999) Efficient progressive sampling. In: Proceedings of the Fifth ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. pp 23–32
1999
Earlier work this paper cites.
Domingos P, Hulten G (2000) Mining High-Speed Data Streams. In: Proceedings of the sixth ACM SIGKDD international conference on Knowledge discovery and data mining, pp 71–80
2000
Earlier work this paper cites.
Fürnkranz J, Petrak J (2001) An evaluation of landmarking variants. In: Working Notes of the ECML/PKDD 2000 Workshop on Integrating Aspects of Data Mining, Decision Support and Meta-Learning, pp 57–68
2000
Earlier work this paper cites.
Petrak J (2000) Fast subsampling performance estimates for classification algorithm selection. In: Proceedings of the ECML-00 Workshop on Meta-Learning: Building Automatic Advice Strategies for Model Selection and Method Combination, pp 3–14
2000
Earlier work this paper cites.
Gu B, Hu F, Liu H (2001) Modelling classification performance for large data sets. In: Advances in Web-Age Information Management, Second International Conference, WAIM 2001. pp 317–328
2001
Earlier work this paper cites.
Ng AY, Jordan MI (2001) On discriminative vs. generative classifiers: A comparison of logistic regression and naive bayes. In: Advances in Neural Information Processing Systems 14. pp 841–848
2001
Earlier work this paper cites.
Meek C, Thiesson B, Heckerman D (2002) The learning-curve sampling method applied to model-based clustering. Journal of Machine Learning Research 2:397–418
2002
Earlier work this paper cites.
Leite R, Brazdil P (2003) Improving progressive sampling via meta-learning. In: Progress in Artificial Intelligence, 11th Protuguese Conference on Artificial Intelligence, EPIA 2003. pp 313–323
2003
Earlier work this paper cites.
Mukherjee S, Tamayo P, Rogers S, et al (2003) Estimating dataset size requirements for classifying DNA microarray data. Journal of Computational Biology 10(2):119–142
2003
Earlier work this paper cites.
Perlich C, Provost FJ, Simonoff JS (2003) Tree induction vs. logistic regression: A learning-curve analysis. Journal of Machine Learning Research 4:211–255
2003
Earlier work this paper cites.
Weiss GM, Provost FJ (2003) Learning when training data are costly: The effect of class distribution on tree induction. Journal of Artificial Intelligence Research 19:315–354
2003
Earlier work this paper cites.
Boonyanunta N, Zeephongsekul P (2004) Predicting the relationship between the size of training sample and the predictive power of classifiers. In: Knowledge-Based Intelligent Information and Engineering Systems, 8th International Conference, KES 2004. pp 529–535
2004
Earlier work this paper cites.
Van den Bosch A (2004) Wrapped progressive sampling search for optimizing learning algorithm parameters. In: Proceedings of the 16th Belgian-Dutch Conference on Artificial Intelligence, pp 219–226
2004
Earlier work this paper cites.
Forman G, Cohen I (2004) Learning from little: Comparison of classifiers given little training. In: Knowledge Discovery in Databases: PKDD 2004, 8th European Conference on Principles and Practice of Knowledge Discovery in Databases. pp 161–172
2004
Earlier work this paper cites.
Leite R, Brazdil P (2004) Improving progressive sampling via meta-learning on learning curves. In: Machine Learning: ECML 2004, 15th European Conference on Machine Learning. pp 250–261
2004
Earlier work this paper cites.
Leite R, Brazdil P (2005) Predicting relative performance of classifiers from samples. In: Machine Learning, Proceedings of the Twenty-Second International Conference (ICML 2005). pp 497–503
2005
Earlier work this paper cites.
Singh S (2005) Modeling performance of different classification methods: deviation from the power law. Project Report, Department of Computer Science, Vanderbilt University, USA
2005
Earlier work this paper cites.
Ng W, Dash M (2006) An evaluation of progressive sampling for imbalanced data sets. In: Workshops Proceedings of the 6th IEEE International Conference on Data Mining (ICDM 2006). pp 657–661
2006
Earlier work this paper cites.
Weiss GM, Tian Y (2006) Maximizing classifier utility when training data is costly. SIGKDD Explorations 8(2):31–38
2006
Cited alongside, same era.
Last M (2007) Predicting and optimizing classifier utility with the power law. In: Workshops Proceedings of the 7th IEEE International Conference on Data Mining (ICDM 2007). pp 219–224
2007
Cited alongside, same era.
Leite R, Brazdil P (2007) An iterative process for building learning curves and predicting relative performance of classifiers. In: Progress in Artificial Intelligence, 13th Portuguese Conference on Aritficial Intelligence, EPIA 2007. pp 87–98
2007
Cited alongside, same era.
Leite R, Brazdil P (2008) Selecting classifiers using metalearning with sampling landmarks and data characterization. In: Proceedings of the 2nd Planning to Learn Workshop (PlanLearn) at ICML/COLT/UAI, pp 35–41
2008
Cited alongside, same era.
Alwosheel A, van Cranenburgh S, Chorus CG (2018) Is your dataset big enough? sample size requirements when using artificial neural networks for discrete choice analysis. Journal of choice modelling 28:167–182
2018
Later among the works it cites.
Baker B, Gupta O, Raskar R, et al (2018) Accelerating neural architecture search using performance prediction. In: 6th International Conference on Learning Representations, ICLR’18
2018
Later among the works it cites.
Bifet A, Gavaldà R, Holmes G, et al (2018) Machine learning for data streams: with practical examples in MOA. MIT press
2018
Later among the works it cites.
Eggensperger K, Lindauer M, Hoos HH, et al (2018) Efficient benchmarking of algorithm configurators via model-based surrogates. Machine Learning 107(1):15–41
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2008
Cited alongside, same era.
Weiss GM, Tian Y (2008) Maximizing classifier utility when there are data acquisition and modeling costs. Data Mining and Knowledge Discovery 17(2):253–282
2008
Cited alongside, same era.
Last M (2009) Improving data mining utility with projective sampling. In: Proceedings of the 15th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. pp 487–496
2009
Cited alongside, same era.
Settles B (2009) Active learning literature survey. Tech. rep., University of Wisconsin
2009
Cited alongside, same era.
Hess KR, Wei C (2010) Learning curves in classification with microarray data. Seminars in oncology 37(1):65–68
2010
Cited alongside, same era.
Leite R, Brazdil P (2010) Active testing strategy to predict the best classification algorithm via sampling and metalearning. In: ECAI 2010 - 19th European Conference on Artificial Intelligence. pp 309–314
2010
Cited alongside, same era.
Tomanek K (2010) Resource-aware annotation through active learning. PhD thesis, Dortmund University of Technology
2010
Cited alongside, same era.
Figueroa RL, Zeng-Treitler Q, Kandula S, et al (2012) Predicting sample size required for classification performance. BMC Medical Informatics Decis Mak 12:8
2012
Cited alongside, same era.
Strang B, van der Putten P, van Rijn JN, et al (2018) Don’t rule out simple models prematurely: A large scale benchmark comparing linear and non-linear classifiers in openml. In: Advances in Intelligent Data Analysis XVII. pp 303–315
2018
Later among the works it cites.
Loog M, Viering TJ, Mey A (2019) Minimizers of the empirical risk and risk monotonicity. In: Advances in Neural Information Processing Systems 32, pp 7476–7485
2019
Later among the works it cites.
Oyedare T, Park JJ (2019) Estimating the required training dataset size for transmitter classification using deep learning. In: 2019 IEEE International Symposium on Dynamic Spectrum Access Networks, DySPAN 2019. pp 1–10
2019
Later among the works it cites.
Richter AN, Khoshgoftaar TM (2019) Approximating learning curves for imbalanced big data with limited labels. In: 31st IEEE International Conference on Tools with Artificial Intelligence, ICTAI 2019. pp 237–242
2019
Later among the works it cites.
Bornschein J, Visin F, Osindero S (2020) Small data, big decisions: Model selection in the small-data regime. In: Proceedings of the 37th International Conference on Machine Learning. pp 1035–1044
2020
Later among the works it cites.
Dong X, Yang Y (2020) Nas-bench-201: Extending the scope of reproducible neural architecture search. In: 8th International Conference on Learning Representations, ICLR 2020
2020
Later among the works it cites.
Long D, Zhang S, Zhang Y (2020) Performance prediction based on neural architecture features. Cognitive Computation and Systems 2(2):80–83
2020
Later among the works it cites.
Nakkiran P, Kaplun G, Bansal Y, et al (2020) Deep double descent: Where bigger models and more data hurt. In: 8th International Conference on Learning Representations, ICLR’20
2020
Later among the works it cites.
Viering TJ, Mey A, Loog M (2020) Making learners (more) monotone. In: Advances in Intelligent Data Analysis XVIII. pp 535–547
2020
Later among the works it cites.
Eggensperger K, Müller P, Mallik N, et al (2021) HPOBench: A collection of reproducible multi-fidelity benchmark problems for HPO. In: Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks
2021
Later among the works it cites.
Hüllermeier E, Waegeman W (2021) Aleatoric and epistemic uncertainty in machine learning: an introduction to concepts and methods. Machine Learning 110(3):457–506
2021
Later among the works it cites.
2021
Later among the works it cites.
Mhammedi Z, Husain H (2021) Risk-monotonicity in statistical learning. In: Advances in Neural Information Processing Systems 34, pp 10,732–10,744
2021
Later among the works it cites.
Mohr F, van Rijn JN (2021) Towards model selection using learning curve cross-validation. In: 8th ICML Workshop on Automated Machine Learning (AutoML)
2021
Later among the works it cites.
Nakkiran P, Venkat P, Kakade SM, et al (2021) Optimal regularization can mitigate double descent. In: 9th International Conference on Learning Representations, ICLR 2021
2021
Later among the works it cites.
Brazdil P, van Rijn JN, Soares C, et al (2022) Metalearning: Applications to Automated Machine Learning and Data Mining, 2nd edn. Springer
2022
Closest in time.
Mohr F, Viering TJ, Loog M, et al (2022) LCDB 1.0: An extensive learning curves database for classification tasks. In: Machine Learning and Knowledge Discovery in Databases - European Conference, ECML PKDD. pp 3–19
2022
Closest in time.
Pfisterer F, Schneider L, Moosbauer J, et al (2022) YAHPO gym - an efficient multi-objective multi-fidelity benchmark for hyperparameter optimization. In: International Conference on Automated Machine Learning, AutoML. pp 3/1–39
2022
Closest in time.
Wang X, Chen Y, Zhu W (2022) A survey on curriculum learning. IEEE Transactions on Pattern Analysis and Machine Intelligence 44(9):4555–4576
2022
Closest in time.
Adriaensen S, Rakotoarison H, Müller S, et al (2023) Efficient bayesian learning curve extrapolation using prior-data fitted networks. In: Advances in Neural Information Processing Systems 36, pp 19,858–19,886
2023
Closest in time.
Hollmann N, Müller S, Eggensperger K, et al (2023) Tabpfn: A transformer that solves small tabular classification problems in a second. In: The Eleventh International Conference on Learning Representations, ICLR 2023
2023
Closest in time.
Mohr F, van Rijn JN (2023) Fast and informative model selection using learning curve cross-validation. IEEE Transactions on Pattern Analysis and Machine Intelligence 45(8):9669–9680
2023
Closest in time.
Ruhkopf T, Mohan A, Deng D, et al (2023) Masif: Meta-learned algorithm selection using implicit fidelity information. Transactions on Machine Learning Research 2023
2023
Closest in time.
Viering TJ, Loog M (2023) The shape of learning curves: A review. IEEE Transactions on Pattern Analysis and Machine Intelligence 45(6):7799–7819
2023
Closest in time.
2023
Closest in time.
Egele R, Mohr F, Viering T, et al (2024) The unreasonable effectiveness of early discarding after one epoch in neural network hyperparameter optimization. Neurocomputing 597:127,964
2024
Closest in time.
Kielhöfer L, Mohr F, van Rijn JN (2024) Learning curve extrapolation methods across extrapolation settings. In: Advances in Intelligent Data Analysis XXII. pp 145–157
2024
Closest in time.