R. A. Fisher, “The use of multiple measurements in taxonomic problems,” Annals of eugenics , vol. 7, no. 2, pp. 179–188, 1936
1936
Earlier work this paper cites.
W. S. McCulloch and W. Pitts, “A logical calculus of the ideas immanent in nervous activity,” The bulletin of mathematical biophysics , vol. 5, no. 4, pp. 115–133, 1943
1943
Earlier work this paper cites.
E. Fix, Discriminatory analysis: nonparametric discrimination, consistency properties . USAF school of Aviation Medicine, 1951
1951
Earlier work this paper cites.
F. J. Massey Jr, “The kolmogorov-smirnov test for goodness of fit,” Journal of the American statistical Association , vol. 46, no. 253, pp. 68–78, 1951
1951
Earlier work this paper cites.
C. Manapragada, G. I. Webb, and M. Salehi, “Extremely fast decision tree,” in Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining , 2018, pp. 1953–1962
1962
Earlier work this paper cites.
C. Chow and C. Liu, “Approximating discrete probability distributions with dependence trees,” IEEE transactions on Information Theory , vol. 14, no. 3, pp. 462–467, 1968
1968
Earlier work this paper cites.
C. L. Giles, C. B. Miller, D. Chen, H.-H. Chen, G.-Z. Sun, and Y.-C. Lee, “Learning and extracting finite state automata with second-order recurrent neural networks,” Neural Computation , vol. 4, no. 3, pp. 393–405, 1992
1992
Earlier work this paper cites.
P. Langley and S. Sage, “Oblivious decision trees and abstract cases,” in Working notes of the AAAI-94 workshop on case-based reasoning . Seattle, WA, 1994, pp. 113–117
1994
Earlier work this paper cites.
B. G. Horne and C. L. Giles, “An experimental comparison of recurrent neural networks,” Advances in neural information processing systems , pp. 697–704, 1995
1995
Earlier work this paper cites.
G. Van Rossum and F. L. Drake Jr, Python reference manual . Centrum voor Wiskunde en Informatica Amsterdam, 1995
1995
Earlier work this paper cites.
L. Willenborg and T. De Waal, Statistical disclosure control in practice . Springer Science & Business Media, 1996, vol. 111
1996
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural computation , vol. 9, no. 8, pp. 1735–1780, 1997
1997
Earlier work this paper cites.
C. Z. Mooney, Monte carlo simulation . Sage, 1997, no. 116
1997
Earlier work this paper cites.
R. K. Pace and R. Barry, “Sparse spatial autoregressions,” Statistics & Probability Letters , vol. 33, no. 3, pp. 291–297, 1997
1997
Earlier work this paper cites.
P. Domingos and G. Hulten, “Mining high-speed data streams,” in Proceedings of the sixth ACM SIGKDD international conference on Knowledge discovery and data mining , 2000, pp. 71–80
2000
Earlier work this paper cites.
D. Micci-Barreca, “A preprocessing scheme for high-cardinality categorical attributes in classification and prediction problems,” SIGKDD Explor. , vol. 3, pp. 27–32, 2001
2001
Earlier work this paper cites.
L. Breiman, “Random forests,” Machine learning , vol. 45, no. 1, pp. 5–32, 2001
2001
Earlier work this paper cites.
J. H. Friedman, “Stochastic gradient boosting,” Computational statistics & data analysis , vol. 38, no. 4, pp. 367–378, 2002
2002
Earlier work this paper cites.
N. V. Chawla, K. W. Bowyer, L. O. Hall, and W. P. Kegelmeyer, “Smote: synthetic minority over-sampling technique,” Journal of artificial intelligence research , vol. 16, pp. 321–357, 2002
2002
Earlier work this paper cites.
R. U. David M. Lane, Introduction to Statistics . David Lane, 2003
2003
Earlier work this paper cites.
A. F. Karr, A. P. Sanil, and D. L. Banks, “Data quality: A statistical perspective,” Statistical Methodology , vol. 3, no. 2, pp. 137–173, 2006
2006
Earlier work this paper cites.
F. Moosmann, B. Triggs, and F. Jurie, “Fast discriminative visual codebooks using randomized clustering forests,” in Twentieth Annual Conference on Neural Information Processing Systems (NIPS’06) . MIT Press, 2006, pp. 985–992
2006
Earlier work this paper cites.
M. Richardson, E. Dominowska, and R. Ragno, “Predicting clicks: estimating the click-through rate for new ads,” in Proceedings of the 16th international conference on World Wide Web , 2007, pp. 521–530
2007
Earlier work this paper cites.
F. Scarselli, M. Gori, A. C. Tsoi, M. Hagenbuchner, and G. Monfardini, “The graph neural network model,” IEEE transactions on neural networks , vol. 20, no. 1, pp. 61–80, 2008
2008
Earlier work this paper cites.
A. Krizhevsky, “Learning multiple layers of features from tiny images,” 2009
2009
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in 2009 IEEE conference on computer vision and pattern recognition . Ieee, 2009, pp. 248–255
2009
Earlier work this paper cites.
Y. LeCun and C. Cortes, “MNIST handwritten digit database,” 2010. [Online]. Available: http://yann.lecun.com/exdb/mnist/
2010
Earlier work this paper cites.
S. Rendle, “Factorization machines,” in 2010 IEEE International conference on data mining . IEEE, 2010, pp. 995–1000
2010
Earlier work this paper cites.
E. Fitkov-Norris, S. Vahid, and C. Hand, “Evaluating the impact of categorical data encoding and scaling on neural network classification performance: the case of repeat consumption of identical cultural goods,” in International Conference on Engineering Applications of Neural Networks . Springer, 2012, pp. 343–352
2012
Earlier work this paper cites.
Y. Lou, R. Caruana, and J. Gehrke, “Intelligible models for classification and regression,” in Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD) , 2012
2012
Earlier work this paper cites.
X. He, J. Pan, O. Jin, T. Xu, B. Liu, T. Xu, Y. Shi, A. Atallah, R. Herbrich, S. Bowers et al. , “Practical lessons from predicting clicks on ads at facebook,” in Proceedings of the Eighth International Workshop on Data Mining for Online Advertising , 2014, pp. 1–9
2014
Earlier work this paper cites.
I. J. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial networks,” arXiv preprint arXiv:1406.2661 , 2014
Original
2014
Earlier work this paper cites.
D. P. Kingma and M. Welling, “Auto-encoding variational bayes,” 2nd International Conference on Learning Representations, ICLR 2014 - Conference Track Proceedings , no. Ml, pp. 1–14, 2014
2014
Earlier work this paper cites.
P. Baldi, P. Sadowski, and D. Whiteson, “Searching for exotic particles in high-energy physics with deep learning,” Nature communications , vol. 5, no. 1, pp. 1–9, 2014
2014
Earlier work this paper cites.
D. Merkel, “Docker: lightweight linux containers for consistent development and deployment,” Linux journal , vol. 2014, no. 239, p. 2, 2014
2014
Earlier work this paper cites.
J. Schmidhuber, “Deep learning in neural networks: An overview,” Neural networks , vol. 61, pp. 85–117, 2015
2015
Earlier work this paper cites.
A. L. Buczak and E. Guven, “A survey of data mining and machine learning methods for cyber security intrusion detection,” IEEE Communications surveys & tutorials , vol. 18, no. 2, pp. 1153–1176, 2015
2015
Earlier work this paper cites.
S. Bach, A. Binder, G. Montavon, F. Klauschen, K.-R. Müller, and W. Samek, “On pixel-wise explanations for non-linear classifier decisions by layer-wise relevance propagation,” PloS one , vol. 10, no. 7, p. e0130140, 2015
2015
Earlier work this paper cites.
I. Goodfellow, Y. Bengio, and A. Courville, Deep learning . MIT press, 2016
2016
Earlier work this paper cites.
GDPR, “Regulation (eu) 2016/679 of the european parliament and of the council,” Official Journal of the European Union , 2016. [Online]. Available: http://www.privacyregulation.eu/en/13.htm
2016
Earlier work this paper cites.
T. Chen and C. Guestrin, “Xgboost: A scalable tree boosting system,” in Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD) , 2016, pp. 785–794
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 770–778
2016
Earlier work this paper cites.
I. Goodfellow, Y. Bengio, and A. Courville, Deep Learning . MIT Press, 2016
2016
Earlier work this paper cites.
H.-T. Cheng, L. Koc, J. Harmsen, T. Shaked, T. Chandra, H. Aradhye, G. Anderson, G. Corrado, W. Chai, M. Ispir et al. , “Wide & deep learning for recommender systems,” in Proceedings of the 1st workshop on deep learning for recommender systems , 2016, pp. 7–10
2016
Earlier work this paper cites.
C. Guo and F. Berkhahn, “Entity embeddings of categorical variables,” arXiv preprint arXiv:1604.06737 , 2016
Original
2016
Earlier work this paper cites.
A. F. T. Martins and R. F. Astudillo, “From softmax to sparsemax: A sparse model of attention and multi-label classification,” arxiv:1602.02068 , 2016
Original
2016
Earlier work this paper cites.
N. Patki, R. Wedge, and K. Veeramachaneni, “The synthetic data vault,” in 2016 IEEE International Conference on Data Science and Advanced Analytics (DSAA) . IEEE, 2016, pp. 399–410
2016
Earlier work this paper cites.
A. Radford, L. Metz, and S. Chintala, “Unsupervised representation learning with deep convolutional generative adversarial networks,” in 4th International Conference on Learning Representations, ICLR 2016 , 2016
2016
Earlier work this paper cites.
G. Kasneci and T. Gottron, “Licon: A linear weighting scheme for the contribution ofinput variables in deep artificial neural networks,” in CIKM , 2016
2016
Earlier work this paper cites.
M. T. Ribeiro, S. Singh, and C. Guestrin, “” why should i trust you?” explaining the predictions of any classifier,” in Proceedings of the 22nd ACM SIGKDD international conference on knowledge discovery and data mining , 2016, pp. 1135–1144
2016
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in Advances in neural information processing systems , 2017, pp. 5998–6008
2017
Earlier work this paper cites.
H. Guo, R. Tang, Y. Ye, Z. Li, and X. He, “Deepfm: a factorization-machine based neural network for ctr prediction,” arXiv preprint arXiv:1703.04247 , 2017
Original
2017
Earlier work this paper cites.
M. Ahmed, H. Afzal, A. Majeed, and B. Khan, “A survey of evolution in predictive models and impacting factors in customer churn,” Advances in Data Science and Adaptive Analysis , vol. 9, no. 03, p. 1750007, 2017
2017
Earlier work this paper cites.
D. Sahoo, Q. Pham, J. Lu, and S. C. Hoi, “Online deep learning: Learning deep neural networks on the fly,” arXiv preprint arXiv:1711.03705 , 2017
Original
2017
Earlier work this paper cites.
E. Choi, S. Biswal, B. Malin, J. Duke, W. F. Stewart, and J. Sun, “Generating Multi-label Discrete Patient Records using Generative Adversarial Networks,” Machine learning for healthcare conference , pp. 286–305, 2017
2017
Earlier work this paper cites.
P. Voigt and A. Von dem Bussche, “The eu general data protection regulation (gdpr),” A Practical Guide, 1st Ed., Cham: Springer International Publishing , vol. 10, p. 3152676, 2017
2017
Earlier work this paper cites.
D. Dua and C. Graff, “UCI machine learning repository,” 2017. [Online]. Available: http://archive.ics.uci.edu/ml
2017
Earlier work this paper cites.
G. Ke, Q. Meng, T. Finley, T. Wang, W. Chen, W. Ma, Q. Ye, and T.-Y. Liu, “Lightgbm: A highly efficient gradient boosting decision tree,” in Advances in neural information processing systems , 2017, pp. 3146–3154
2017
Earlier work this paper cites.
B. R. Mitchell et al. , “The spatial inductive bias of deep learning,” Ph.D. dissertation, Johns Hopkins University, 2017
2017
Earlier work this paper cites.
D. Baylor, E. Breck, H.-T. Cheng, N. Fiedel, C. Y. Foo, Z. Haque, S. Haykal, M. Ispir, V. Jain, L. Koc et al. , “ TFX
2017
Earlier work this paper cites.
N. Frosst and G. Hinton, “Distilling a neural network into a soft decision tree,” arXiv preprint arXiv:1711.09784 , 2017
Original
2017
Earlier work this paper cites.
H. Zhang, M. Cisse, Y. N. Dauphin, and D. Lopez-Paz, “mixup: Beyond empirical risk minimization,” arXiv preprint arXiv:1710.09412 , 2017
Original
2017
Earlier work this paper cites.
S. Lundberg and S.-I. Lee, “A unified approach to interpreting model predictions,” NeurIPS , 2017
2017
Earlier work this paper cites.
K. Lin, D. Li, X. He, Z. Zhang, and M.-T. Sun, “Adversarial ranking for language generation,” Advances in Neural Information Processing Systems , 2017
2017
Earlier work this paper cites.
S. Subramanian, S. Rajeswar, F. Dutil, C. Pal, and A. Courville, “Adversarial generation of natural language,” in Proceedings of the 2nd Workshop on Representation Learning for NLP , 2017, pp. 241–251
2017
Earlier work this paper cites.
J. Zhang, G. Cormode, C. M. Procopiuc, D. Srivastava, and X. Xiao, “Privbayes: Private data release via bayesian networks,” ACM Transactions on Database Systems (TODS) , vol. 42, no. 4, pp. 1–41, 2017
2017
Earlier work this paper cites.
M. Arjovsky, S. Chintala, and L. Bottou, “Wasserstein generative adversarial networks,” in International conference on machine learning . PMLR, 2017, pp. 214–223
2017
Earlier work this paper cites.
I. Gulrajani, F. Ahmed, M. Arjovsky, V. Dumoulin, and A. Courville, “Improved training of wasserstein gans,” in Proceedings of the 31st International Conference on Neural Information Processing Systems , 2017, pp. 5769–5779
2017
Earlier work this paper cites.
M. G. Bellemare, I. Danihelka, W. Dabney, S. Mohamed, B. Lakshminarayanan, S. Hoyer, and R. Munos, “The cramer distance as a solution to biased wasserstein gradients,” 2017
2017
Earlier work this paper cites.
A. Srivastava, L. Valkov, C. Russell, M. U. Gutmann, and C. Sutton, “Veegan: Reducing mode collapse in gans using implicit variational learning,” in Proceedings of the 31st International Conference on Neural Information Processing Systems , 2017, pp. 3310–3320
2017
Earlier work this paper cites.
E. Jang, S. Gu, and B. Poole, “Categorical reparameterization with gumbel-softmax,” in International Conference on Learning Representations , 2017
2017
Earlier work this paper cites.
M. Sundararajan, A. Taly, and Q. Yan, “Axiomatic attribution for deep networks,” in International Conference on Machine Learning . PMLR, 2017, pp. 3319–3328
2017
Earlier work this paper cites.
L. Breiman, J. H. Friedman, R. A. Olshen, and C. J. Stone, Classification and regression trees . Routledge, 2017
2017
Earlier work this paper cites.