Fetching the paper…
Reading the bibliography…
The experiments covered by Machine Learning (ML) must consider two important aspects to assess the performance of a model: datasets and algorithms.
Nemenyi, P.: Distribution-free multiple comparisons. In: Biometrics, vol. 18, p. 263 (1962). International Biometric Soc 1441 I ST, NW, SUITE 700, WASHINGTON, DC 20005-2210
1962
Earlier work this paper cites.
Birnbaum, A., Lord, F., Novick, M.: Statistical theories of mental test scores. Some latent trait models and their use in inferring an examinee’s ability. Addison-Wesley, Reading, MA (1968)
1968
Earlier work this paper cites.
Elo, A.E.: The Rating of Chessplayers, Past and Present, (1978)
1978
Earlier work this paper cites.
Lord, F.M., Wingersky, M.S.: Comparison of irt true-score and equipercentile observed-score “equatings”. Applied Psychological Measurement 8
1984
Earlier work this paper cites.
de Andrade, D.F., Tavares, H.R., da Cunha Valle, R.: Teoria da resposta ao item: conceitos e aplicações. ABE, Sao Paulo (2000)
2000
Earlier work this paper cites.
Monard, M.C., Baranauskas, J.A.: Conceitos sobre aprendizado de máquina. Sistemas inteligentes-Fundamentos e aplicações 1
2003
Earlier work this paper cites.
Rizopoulos, D.: ltm: An r package for latent variable modeling and item response theory analyses. Journal of statistical software 17
2006
Earlier work this paper cites.
Gautier, L.: rpy2: A simple and efficient access to r from python. URL http://rpy. sourceforge. net/rpy2. html 3
2008
Earlier work this paper cites.
Ferri, C., Hernández-Orallo, J., Modroiu, R.: An experimental comparison of performance measures for classification. Pattern Recognition Letters 30
2009
Earlier work this paper cites.
Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., et al
2011
Earlier work this paper cites.
Domingos, P.: A few useful things to know about machine learning. Communications of the ACM 55
2012
Earlier work this paper cites.
Glickman, M.E.: Example of the glicko-2 system. Boston University, 1–6 (2012)
2012
Cited alongside, same era.
Bellemare, M.G., Naddaf, Y., Veness, J., Bowling, M.: The arcade learning environment: An evaluation platform for general agents. Journal of Artificial Intelligence Research 47
2013
Cited alongside, same era.
Adedoyin, O., Mokobi, T., et al
2013
Cited alongside, same era.
Vanschoren, J., Van Rijn, J.N., Bischl, B., Torgo, L.: Openml: networked science in machine learning. ACM SIGKDD Explorations Newsletter 15
2014
Cited alongside, same era.
Samothrakis, S., Perez, D., Lucas, S.M., Rohlfshagen, P.: Predicting dominance rankings for score-based games. IEEE Transactions on Computational Intelligence and AI in Games 8
2014
Cited alongside, same era.
Martínez-Plumed, F., Prudêncio, R.B., Martínez-Usó, A., Hernández-Orallo, J.: Making sense of item response theory in machine learning. In: Proceedings of the Twenty-second European Conference on Artificial Intelligence, pp. 1140–1148 (2016)
2016
Later among the works it cites.
Dua, D., Graff, C.: UCI Machine Learning Repository (2017). http://archive.ics.uci.edu/ml
2017
Later among the works it cites.
Bischl, B., Casalicchio, G., Feurer, M., Hutter, F., Lang, M., Mantovani, R.G., van Rijn, J.N., Vanschoren, J.: Openml benchmarking suites and the openml100. stat 1050
2017
Later among the works it cites.
Kubat, M.: An Introduction to Machine Learning. Springer, ??? (2017)
2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Veček, N., Mernik, M., Črepinšek, M.: A chess rating system for evolutionary algorithms: A new method for the comparison and ranking of evolutionary algorithms. Information Sciences 277
2014
Cited alongside, same era.
Smith, M.R., Martinez, T.: Reducing the effects of detrimental instances. In: 2014 13th International Conference on Machine Learning and Applications, pp. 183–188 (2014). IEEE
2014
Cited alongside, same era.
Prudêncio, R.B., Hernández-Orallo, J., Martınez-Usó, A.: Analysis of instance hardness in machine learning using item response theory. In: Second International Workshop on Learning over Multiple Contexts in ECML 2015. Porto, Portugal, 11 September 2015, vol. 1 (2015)
2015
Cited alongside, same era.
Perez-Liebana, D., Samothrakis, S., Togelius, J., Schaul, T., Lucas, S.M., Couëtoux, A., Lee, J., Lim, C.-U., Thompson, T.: The 2014 general video game playing competition. IEEE Transactions on Computational Intelligence and AI in Games 8
2015
Cited alongside, same era.
Pereira, D.G., Afonso, A., Medeiros, F.M.: Overview of friedman’s test and post-hoc analysis. Communications in Statistics-Simulation and Computation 44
2015
Cited alongside, same era.
OpenML: An open, collaborative, frictionless, automated machine learning environment. https://docs.openml.org/#studies-under-construction
Cited in the paper.
Facebook: Rethinking AI Benchmarking. Facebook. https://dynabench.org/about
Cited in the paper.
2017
Later among the works it cites.
Martinez-Plumed, F., Hernandez-Orallo, J.: Dual indicators to analyze ai benchmarks: Difficulty, discrimination, ability, and generality. IEEE Transactions on Games 12
2018
Later among the works it cites.
Martínez-Plumed, F., Prudêncio, R.B., Martínez-Usó, A., Hernández-Orallo, J.: Item response theory in ai: Analysing machine learning classifiers at the instance level. Artificial Intelligence 271
2019
Later among the works it cites.
2019
Later among the works it cites.
Cardoso, L.F., Santos, V.C., Francês, R.S.K., Prudêncio, R.B., Alves, R.C.: Decoding machine learning benchmarks. In: Brazilian Conference on Intelligent Systems, pp. 412–425 (2020). Springer
2020
Later among the works it cites.
Song, H., Flach, P.: Efficient and robust model benchmarks with item response theory and adaptive testing. International Journal of Interactive Multimedia & Artificial Intelligence 6
2021
Closest in time.