Fetching the paper…
Reading the bibliography…
Standardized datasets and benchmarks have spurred innovations in computer vision, natural language processing, multi-modal and tabular settings.
Bagging predictors
Leo Breiman · 1996
Earlier work this paper cites.
The foundations of cost-sensitive learning
Charles Elkan · 2001
Earlier work this paper cites.
Smote: Synthetic minority over-sampling technique
Nitesh V Chawla, Kevin W Bowyer, Lawrence O Hall, and W Philip Kegelmeyer · 2002
Earlier work this paper cites.
Using random forest to learn imbalanced data
Chao Chen, Andy Liaw, and Leo Breiman · 2004
Earlier work this paper cites.
Semi-supervised learning literature survey
Xiaojin Jerry Zhu · 2005
Earlier work this paper cites.
Tri-training: Exploiting unlabeled data using three classifiers
Zhi-Hua Zhou and Ming Li · 2005
Earlier work this paper cites.
Casting out demons: Sanitizing training data for anomaly sensors
Gabriela F Cretu, Angelos Stavrou, Michael E Locasto, Salvatore J Stolfo, and Angelos D Keromytis · 2008
Earlier work this paper cites.
A survey of semi-supervised learning methods
Nitin Namdeo Pise and Parag Kulkarni · 2008
Earlier work this paper cites.
ImageNet: a Large-Scale Hierarchical Image Database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Fei-Fei Li · 2009
Earlier work this paper cites.
Learning Multiple Layers of Features from Tiny Images
Alex Krizhevsky · 2009
Earlier work this paper cites.
Semi-supervised learning (chapelle, o. et al., eds.; 2006)[book reviews]
Olivier Chapelle, Bernhard Scholkopf, and Alexander Zien · 2009
Earlier work this paper cites.
Rusboost: A hybrid approach to alleviating class imbalance
Chris Seiffert, Taghi M Khoshgoftaar, Jason Van Hulse, and Amri Napolitano · 2010
Earlier work this paper cites.
UCF101: A dataset of 101 human actions classes from videos in the wild
Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah · 2012
Earlier work this paper cites.
Scikit-clean
Shihab Shahriar Khan · 2013
Earlier work this paper cites.
Microsoft COCO: Common Objects in Context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick · 2014
Earlier work this paper cites.
SNAP Datasets: Stanford Large Network Dataset Collection
Jure Leskovec and Andrej Krevl · 2014
Cited alongside, same era.
Self-labeled techniques for semi-supervised learning: taxonomy, software and empirical study
Isaac Triguero, Salvador García, and Francisco Herrera · 2015
Cited alongside, same era.
Efficient and robust automated machine learning
Matthias Feurer, Aaron Klein, Katharina Eggensperger, Jost Springenberg, Manuel Blum, and Frank Hutter · 2015
Cited alongside, same era.
SQuAD: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang · 2016
Cited alongside, same era.
The kinetics human action video dataset
Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, et al · 2017
Cited alongside, same era.
Open Graph Benchmark: Datasets for Machine Learning on Graphs
Weihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong, Hongyu Ren, Bowen Liu, Michele Catasta, and Jure Leskovec · 2020
Later among the works it cites.
AutoML@NeurIPS 2018 Challenge: Design and Results
Hugo Jair Escalante, Wei-Wei Tu, Isabelle Guyon, Daniel L. Silver, Evelyne Viegas, Yuqiang Chen, Wenyuan Dai, and Qiang Yang · 2020
Later among the works it cites.
AutoGluon-Tabular: Robust and Accurate AutoML for Structured Data
Nick Erickson, Jonas Mueller, Alexander Shirkov, Hang Zhang, Pedro Larroy, Mu Li, and Alexander Smola · 2020
Later among the works it cites.
H2O AutoML: Scalable Automatic Machine Learning
Erin LeDell and Sebastien Poirier · 2020
Later among the works it cites.
Auto-sklearn 2.0: Hands-free automl via meta-learning
Matthias Feurer, Katharina Eggensperger, Stefan Falkner, Marius Lindauer, and Frank Hutter · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman · 2018
Cited alongside, same era.
The UCR Time Series Classification Archive, October 2018
Hoang Anh Dau, Eamonn Keogh, Kaveh Kamgar, Chin-Chia Michael Yeh, Yan Zhu, Shaghayegh Gharghabi, Chotirat Ann Ratanamahatana, Yanping, Bing Hu, Nurjahan Begum, Anthony Bagnall, Abdullah Mueen, Gustavo Batista, and Hexagon-ML · 2018
Cited alongside, same era.
Credit card fraud detection: A novel approach using aggregation strategy and feedback mechanism
Changjun Jiang, Jiahui Song, Guanjun Liu, Lutao Zheng, and Wenjing Luan · 2018
Cited alongside, same era.
Learning from binary labels with instance-dependent noise
Aditya Krishna Menon, Brendan Van Rooyen, and Nagarajan Natarajan · 2018
Cited alongside, same era.
Catboost: unbiased boosting with categorical features
Liudmila Prokhorenkova, Gleb Gusev, Aleksandr Vorobev, Anna Veronika Dorogush, and Andrey Gulin · 2018
Cited alongside, same era.
A two-stage ensemble method for the detection of class-label noise
Maryam Sabzevari, Gonzalo Martínez-Muñoz, and Alberto Suárez · 2018
Cited alongside, same era.
Realistic evaluation of deep semi-supervised learning algorithms
Avital Oliver, Augustus Odena, Colin A Raffel, Ekin Dogus Cubuk, and Ian Goodfellow · 2018
Cited alongside, same era.
Later among the works it cites.
A mixed solution-based high agreement filtering method for class noise detection in binary classification
Maryam Samami, Ebrahim Akbari, Moloud Abdar, Pawel Plawiak, Hossein Nematzadeh, Mohammad Ehsan Basiri, and Vladimir Makarenkov · 2020
Later among the works it cites.
A survey on semi-supervised learning
Jesper E Van Engelen and Holger H Hoos · 2020
Later among the works it cites.
Vime: Extending the success of self-and semi-supervised learning to tabular domain
Jinsung Yoon, Yao Zhang, James Jordon, and Mihaela van der Schaar · 2020
Later among the works it cites.
Benchmarking Multimodal AutoML for Tabular Data with Text Fields
Xingjian Shi, Jonas Mueller, Nick Erickson, Mu Li, and Alexander J Smola · 2021
Later among the works it cites.
Detecting New Account Fraud and Transaction Fraud with Amazon Fraud Detector
Amazon Fraud Detector · 2021
Later among the works it cites.
Confident learning: Estimating uncertainty in dataset labels
Curtis Northcutt, Lu Jiang, and Isaac Chuang · 2021
Later among the works it cites.
Why do tree-based models still outperform deep learning on tabular data?
Léo Grinsztajn, Edouard Oyallon, and Gaël Varoquaux · 2022
Closest in time.
Fraud detection and prevention in e-commerce: A systematic literature review
Vinicius Facco Rodrigues, Lucas Micol Policarpo, Diórgenes Eugênio da Silveira, Rodrigo da Rosa Righi, Cristiano André da Costa, Jorge Luis Victória Barbosa, Rodolfo Stoffel Antunes, Rodrigo Scorsatto, and Tanuj Arcot · 2022
Closest in time.
Massih-Reza Amini, Vasilii Feofanov, Loic Pauletto, Emilie Devijver, and Yury Maximov · 2022
Closest in time.
A supervised machine learning algorithm for detecting and predicting fraud in credit card transactions
Jonathan Kwaku Afriyie, Kassim Tawiah, Wilhemina Adoma Pels, Sandra Addai-Henne, Harriet Achiaa Dwamena, Emmanuel Odame Owiredu, Samuel Amening Ayeh, and John Eshun · 2023
Closest in time.