Fetching the paper…
Reading the bibliography…
Machine learning is the science of discovering statistical dependencies in data, and the use of those dependencies to perform predictions.
On lines and planes of closest fit to systems of points in space
Pearson, K · 1901
Earlier work this paper cites.
Functions of positive and negative type, and their connection with the theory of integral equations
Mercer, J · 1909
Earlier work this paper cites.
Correlation and Causation
Wright, S · 1921
Earlier work this paper cites.
Analysis of a complex of statistical variables into principal components
Hotelling, H · 1933
Earlier work this paper cites.
Relations between two sets of variates
Hotelling, H · 1936
Earlier work this paper cites.
Das statistische problem der korrelation als variations-und eigenwertproblem und sein zusammenhang mit der ausgleichsrechnung
Gebelein, H · 1941
Earlier work this paper cites.
Remarks on a multivariate transformation
Rosenblatt, M · 1952
Earlier work this paper cites.
Dynamic programming and lagrange multipliers
Bellman, R · 1956
Earlier work this paper cites.
The direction of time
Reichenbach, H · 1956
Earlier work this paper cites.
On measures of dependence
Rényi, A · 1959
Earlier work this paper cites.
Fonctions de répartition à n dimensions et leurs marges
Sklar, A · 1959
Earlier work this paper cites.
On estimation of a probability density function and mode
Parzen, E · 1962
Earlier work this paper cites.
Fourier Analysis on Groups
Rudin, W · 1962
Earlier work this paper cites.
Investigating causal relations by econometric models and cross-spectral methods
Granger, C. W · 1969
Earlier work this paper cites.
Partial canonical correlations
Rao, B. R · 1969
Earlier work this paper cites.
On the uniform convergence of relative frequencies of events to their probabilities
Vapnik, V. and Chervonenkis, A · 1971
Earlier work this paper cites.
Functional analysis, volume 1 of methods of modern mathematical physics, 1972
Reed, M. and Simon, B · 1972
Earlier work this paper cites.
Counterfactuals
Lewis, D · 1974
Earlier work this paper cites.
Maximum likelihood from incomplete data via the em algorithm
Dempster, A. P., Laird, N. M., and Rubin, D. B · 1977
Earlier work this paper cites.
Bootstrap methods: another look at the jackknife
Efron, B · 1979
Earlier work this paper cites.
Multivariate analysis
Mardia, K. V., Kent, J. T., and Bibby, J. M · 1979
Earlier work this paper cites.
On nonparametric measures of dependence for random variables
Schweizer, B. and Wolff, E. F · 1981
Earlier work this paper cites.
Estimation of dependences based on empirical data , volume 40
Vapnik, V · 1982
Earlier work this paper cites.
The central role of the propensity score in observational studies for causal effects
Rosenbaum, P. R. and Rubin, D. B · 1983
Earlier work this paper cites.
Probabilistic metric spaces
Schweizer, B. and Sklar, A · 1983
Earlier work this paper cites.
Estimating optimal transformations for multiple regression and correlation
Breiman, L. and Friedman, J. H · 1985
Earlier work this paper cites.
Bayesian networks: A model of self-activated memory for evidential reasoning
Pearl, J · 1985
Earlier work this paper cites.
Distributed representations
Hinton, G., McClelland, J., and Rumelhart, D · 1986
Earlier work this paper cites.
Learning representations by back-propagating errors
Rumelhart, D. E., Hinton, G. E., and Williams, R. J · 1986
Earlier work this paper cites.
Neural networks and principal component analysis: Learning from examples without local minima
Baldi, P. and Hornik, K · 1989
Earlier work this paper cites.
The tight constant in the Dvoretzky-Kiefer-Wolfowitz inequality
Massart, P · 1990
Earlier work this paper cites.
Nonlinear principal component analysis using autoassociative neural networks
Kramer, M. A · 1991
Earlier work this paper cites.
Equivalence and synthesis of causal models
Verma, T. and Pearl, J · 1991
Earlier work this paper cites.
Scale-invariant correlation theory
Hoeffding, W · 1994
Earlier work this paper cites.
Families of
Joe, H · 1996
Earlier work this paper cites.
A Bayesian approach to causal discovery
Heckerman, D., Meek, C., and Cooper, G · 1997
Earlier work this paper cites.
Multivariate models and multivariate dependence concepts
Joe, H · 1997
Earlier work this paper cites.
Kernel principal component analysis
Schölkopf, B., Smola, A., and Müller, K.-R · 1997
Earlier work this paper cites.
No free lunch theorems for optimization
Wolpert, D. H. and Macready, W. G · 1997
Earlier work this paper cites.
Learning in Graphical Models , volume 89
Jordan, M. I · 1998
Earlier work this paper cites.
Statistical learning theory
Vapnik, V · 1998
Earlier work this paper cites.
Myopia and ambient lighting at night
Quinn, G. E., Shin, C. H., Maguire, M. G., and Stone, R. A · 1999
Earlier work this paper cites.
Gaussian identities
Roweis, S · 1999
Earlier work this paper cites.
Linear heteroencoders
Roweis, S. and Brody, C · 1999
Earlier work this paper cites.
Causal inference without counterfactuals
Dawid, A. P · 2000
Earlier work this paper cites.
Rademacher processes and bounding the risk of function learning
Koltchinskii, V. and Panchenko, D · 2000
Earlier work this paper cites.
Kernel and nonlinear canonical correlation analysis
Lai, P. L. and Fyfe, C · 2000
Earlier work this paper cites.
Some applications of concentration inequalities to statistics
Massart, P · 2000
Earlier work this paper cites.
Gaussian mixtures and their applications to signal processing
Plataniotis, K · 2000
Earlier work this paper cites.
Causation, prediction, and search , volume 81
Spirtes, P., Glymour, C. N., and Scheines, R · 2000
Earlier work this paper cites.
Probability density decomposition for conditionally dependent random variables modeled by vines
Bedford, T. and Cooke, R. M · 2001
Earlier work this paper cites.
Random forests
Breiman, L · 2001
Earlier work this paper cites.
Gaussianization
Chen, S. S. and Gopinath, R. A · 2001
Earlier work this paper cites.
Greedy function approximation: a gradient boosting machine
Friedman, J. H · 2001
Earlier work this paper cites.
Multivariate survival modelling: a unified approach with copulas
Georges, P., Lamy, A.-G., Nicolas, E., Quibel, G., and Roncalli, T · 2001
Earlier work this paper cites.
Rademacher penalties and structural risk minimization
Koltchinskii, V · 2001
Earlier work this paper cites.
Expectation propagation for approximate Bayesian inference
Minka, T. P · 2001
Earlier work this paper cites.
Learning with Kernels: Support Vector Machines, Regularization, Optimization, and Beyond
Schölkopf, B. and Smola, A. J · 2001
Earlier work this paper cites.
Venetian sea levels, british bread prices, and the principle of the common cause
Sober, E · 2001
Earlier work this paper cites.
Using the Nyström method to speed up kernel machines
Williams, C. and Seeger, M · 2001
Earlier work this paper cites.
Sampling techniques for kernel methods
Achlioptas, D., McSherry, F., and Schölkopf, B · 2002
Earlier work this paper cites.
Kernel independent component analysis
Bach, F. R. and Jordan, M. I · 2002
Earlier work this paper cites.
Vines: A new graphical model for dependent random variables
Bedford, T. and Cooke, R. M · 2002
Earlier work this paper cites.
Principal component analysis
Jolliffe, I · 2002
Earlier work this paper cites.
Applications of copula theory in financial econometrics
Patton, A. J · 2002
Earlier work this paper cites.
Inferring a semantic representation of text via cross-language correlation analysis
Vinokourov, A., Cristianini, N., and Shawe-taylor, J · 2002
Earlier work this paper cites.
Rademacher and Gaussian complexities: risk bounds and structural results
Bartlett, P. L. and Mendelson, S · 2003
Earlier work this paper cites.
Topics in optimal transportation
Villani, C · 2003
Earlier work this paper cites.
Introduction to statistical learning theory
Bousquet, O., Boucheron, S., and Lugosi, G · 2004
Earlier work this paper cites.
Convex optimization
Boyd, S. and Vandenberghe, L · 2004
Earlier work this paper cites.
Copula methods in finance
Cherubini, U., Luciano, E., and Vecchiato, W · 2004
Earlier work this paper cites.
Canonical correlation analysis: An overview with application to learning methods
Hardoon, D. R., Szedmak, S., and Shawe-Taylor, J · 2004
Earlier work this paper cites.
Probability product kernels
Jebara, T., Kondor, R., and Howard, A · 2004
Earlier work this paper cites.
Introductory lectures on convex optimization: a basic course
Nesterov, Y · 2004
Earlier work this paper cites.
Local Rademacher complexities
Bartlett, P. L., Bousquet, O., and Mendelson, S · 2005
Earlier work this paper cites.
BCI Competition III data, experiment 4a, subject 3, 1000Hz, 2005
Blankertz, B · 2005
Earlier work this paper cites.
Theory of classification: A survey of some recent advances
Boucheron, S., Bousquet, O., and Lugosi, G · 2005
Earlier work this paper cites.
Semigroup kernels on measures
Cuturi, M., Fukumizu, K., and Vert, J.-P · 2005
Earlier work this paper cites.
Eigenproblems in pattern recognition
De Bie, T., Cristianini, N., and Rosipal, R · 2005
Earlier work this paper cites.
The t copula and related copulas
Demarta, S. and McNeil, A. J · 2005
Earlier work this paper cites.
On the Nyström method for approximating a Gram matrix for improved kernel-based learning
Drineas, P. and Mahoney, M. W · 2005
Earlier work this paper cites.
Hilbertian metrics and positive definite kernels on probability measures
Hein, M. and Bousquet, O · 2005
Earlier work this paper cites.
Expectation propagation for exponential families
Seeger, M · 2005
Earlier work this paper cites.
Sparse Gaussian processes using pseudo-inputs
Snelson, E. and Ghahramani, Z · 2005
Earlier work this paper cites.
Empirical minimization
Bartlett, P. L. and Mendelson, S · 2006
Earlier work this paper cites.
Convexity, classification, and risk bounds
Bartlett, P. L., Jordan, M. I., and McAuliffe, J. D · 2006
Earlier work this paper cites.
Pattern recognition and machine learning
Bishop, C. M · 2006
Earlier work this paper cites.
Model compression
Buciluǎ, C., Caruana, R., and Niculescu-Mizil, A · 2006
Cited alongside, same era.
Reducing the dimensionality of data with neural networks
Hinton, G. E. and Salakhutdinov, R. R · 2006
Cited alongside, same era.
Correcting sample selection bias by unlabeled data
Huang, J., Gretton, A., Borgwardt, K. M., Schölkopf, B., and Smola, A. J · 2006
Cited alongside, same era.
Causal models as minimal descriptions of multivariate systems, 2006
Lemeire, J. and Dirkx, E · 2006
Cited alongside, same era.
The Rademacher complexity of linear transformation classes
Maurer, A · 2006
Cited alongside, same era.
An introduction to copulas , volume 139
Nelsen, R. B · 2006
Cited alongside, same era.
Discovering cyclic causal models by independent components analysis
Lacerda, G., Spirtes, P. L., Ramsey, J., and Hoyer, P. O · 2012
Later among the works it cites.
Semi-supervised domain adaptation with non-parametric copulas
Lopez-Paz, D., Hernández-Lobato, J. M., and Schölkopf, B · 2012
Later among the works it cites.
Chocolate consumption, cognitive function, and nobel laureates
Messerli, F. H · 2012
Later among the works it cites.
Foundations of machine learning
Mohri, M., Rostamizadeh, A., and Talwalkar, A · 2012
Later among the works it cites.
Learning from distributions via support measure machines
Muandet, K., Fukumizu, K., Dinuzzo, F., and Schölkopf, B · 2012
Later among the works it cites.
Machine learning: a probabilistic perspective
Murphy, K. P · 2012
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Patton, A. J · 2006
Cited alongside, same era.
Gaussian Processes for Machine Learning
Rasmussen, C. E. and Williams, C. K. I · 2006
Cited alongside, same era.
A linear non-Gaussian acyclic model for causal discovery
Shimizu, S., Hoyer, P. O., Hyvärinen, A., and Kerminen, A · 2006
Cited alongside, same era.
Inference with the Universum
Weston, J., Collobert, R., Sinz, F., Bottou, L., and Vapnik, V · 2006
Cited alongside, same era.
UCI machine learning repository, 2007
Asuncion, A. and Newman, D · 2007
Cited alongside, same era.
Statistical consistency of kernel canonical correlation analysis
Fukumizu, K., Bach, F. R., and Gretton, A · 2007
Cited alongside, same era.
Later among the works it cites.
Pair copula constructions for multivariate discrete data
Panagiotelis, A., Czado, C., and Joe, H · 2012
Later among the works it cites.
Restricted structural equation models for causal inference
Peters, J. M · 2012
Later among the works it cites.
The matrix cookbook, 2012
Petersen, K. B. and Pedersen, M. S · 2012
Later among the works it cites.
Copula-based kernel dependency measures
Póczos, B., Ghahramani, Z., and Schneider, J. G · 2012
Later among the works it cites.
Copula mixture model for dependency-seeking clustering
Rey, M. and Roth, V · 2012
Later among the works it cites.
On causal and anticausal learning
Schölkopf, B., Janzing, D., Peters, J., Sgouritsa, E., Zhang, K., and Mooij, J. M · 2012
Later among the works it cites.
Feature selection via dependence maximization
Song, L., Smola, A., Gretton, A., Bedo, J., and Borgwardt, K · 2012
Later among the works it cites.
Density ratio estimation in machine learning
Sugiyama, M., Suzuki, T., and Kanamori, T · 2012
Later among the works it cites.
Lecture 6.5-RMSProp: Divide the gradient by a running average of its recent magnitude
Tieleman, T. and Hinton, G · 2012
Later among the works it cites.
Nyström method vs random Fourier features: A theoretical and empirical comparison
Yang, T., Li, Y.-F., Mahdavi, M., Jin, R., and Zhou, Z.-H · 2012
Later among the works it cites.
Deep canonical correlation analysis
Andrew, G., Arora, R., Bilmes, J., and Livescu, K · 2013
Later among the works it cites.
Concentration inequalities: A nonasymptotic theory of independence
Boucheron, S., Lugosi, G., and Massart, P · 2013
Later among the works it cites.
Predicting parameters in deep learning
Denil, M., Shakibi, B., Dinh, L., de Freitas, N., et al · 2013
Later among the works it cites.
Selecting and estimating regular vine copulae and application to financial returns
Dissmann, J., Brechmann, E. C., Czado, C., and Kurowicka, D · 2013
Later among the works it cites.
Structure discovery in nonparametric regression through compositional kernel search
Duvenaud, D., Lloyd, J. R., Grosse, R., Tenenbaum, J. B., and Ghahramani, Z · 2013
Later among the works it cites.
Copulas in machine learning
Elidan, G · 2013
Later among the works it cites.
Incorporating privileged information through metric learning
Fouad, S., Tino, P., Raychaudhury, S., and Schneider, P · 2013
Later among the works it cites.
Cause-effect pairs kaggle competition, 2013
Guyon, I · 2013
Later among the works it cites.
Gaussian process conditional copulas with applications to financial time series
Hernández-Lobato, J. M., Lloyd, J. R., and Hernández-Lobato, D · 2013
Later among the works it cites.
Auto-encoding variational Bayes
Kingma, D. P. and Welling, M · 2013
Later among the works it cites.
Fastfood: computing Hilbert space expansions in loglinear time
Le, Q., Sarlos, T., and Smola, A · 2013
Later among the works it cites.
Probability in Banach Spaces: isoperimetry and processes , volume 23
Ledoux, M. and Talagrand, M · 2013
Later among the works it cites.
Correlated random features for fast semi-supervised learning
McWilliams, B., Balduzzi, D., and Buhmann, J. M · 2013
Later among the works it cites.
Efficient estimation of word representations in vector space
Mikolov, T., Chen, K., Corrado, G., and Dean, J · 2013
Later among the works it cites.
Causation: A Very Short Introduction
Mumford, S. and Anjum, R. L · 2013
Later among the works it cites.
Distribution-free distribution regression
Póczos, B., Rinaldo, A., Singh, A., and Wasserman, L · 2013
Later among the works it cites.
Learning to rank using privileged information
Sharmanska, V., Quadrianto, N., and Lampert, C. H · 2013
Later among the works it cites.
Training recurrent neural networks
Sutskever, I · 2013
Later among the works it cites.
Two numerical models designed to reproduce saturn ring temperatures as measured by Cassini-CIRS
Altobelli, N., Lopez-Paz, D., Pilorz, S., Spilker, L. J., Morishima, R., Brooks, S., Leyrat, C., Deau, E., Edgington, S., and Flandes, A · 2014
Later among the works it cites.
Efficient dimensionality reduction for canonical correlation analysis
Avron, H., Boutsidis, C., Toledo, S., and Zouzias, A · 2014
Later among the works it cites.
Do deep nets really need to be deep?
Ba, J. and Caruana, R · 2014
Later among the works it cites.
From machine learning to machine reasoning
Bottou, L · 2014
Later among the works it cites.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Dauphin, Y. N., Pascanu, R., Gulcehre, C., Cho, K., Ganguli, S., and Bengio, Y · 2014
Later among the works it cites.
Do we need hundreds of classifiers to solve real world classification problems?
Fernández-Delgado, M., Cernadas, E., Barro, S., and Amorim, D · 2014
Later among the works it cites.
Generative adversarial nets
Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y · 2014
Later among the works it cites.
Chalearn fast causation coefficient challenge, 2014
Guyon, I · 2014
Later among the works it cites.
Mind the nuisance: Gaussian process classification using privileged noise
Hernández-Lobato, D., Sharmanska, V., Kersting, K., Lampert, C. H., and Quadrianto, N · 2014
Later among the works it cites.
Kernel methods match deep neural networks on timit
Huang, P.-S., Avron, H., Sainath, T. N., Sindhwani, V., and Ramabhadran, B · 2014
Later among the works it cites.
Semi-supervised learning with deep generative models
Kingma, D. P., Mohamed, S., Rezende, D. J., and Welling, M · 2014
Later among the works it cites.
A scalable bootstrap for massive data
Kleiner, A., Talwalkar, A., Sarkar, P., and Jordan, M. I · 2014
Later among the works it cites.
Consistency of causal inference under the additive noise model
Kpotufe, S., Sgouritsa, E., Janzing, D., and Schölkopf, B · 2014
Later among the works it cites.
Learning using privileged information: SVM+ and weighted SVM
Lapin, M., Hein, M., and Schiele, B · 2014
Later among the works it cites.
Microsoft COCO: Common objects in context
Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L · 2014
Later among the works it cites.
Randomized nonlinear component analysis
Lopez-Paz, D., Sra, S., Smola, A. J., Ghahramani, Z., and Schölkopf, B · 2014
Later among the works it cites.
Distinguishing cause from effect using observational data: methods and benchmarks
Mooij, J. M., Peters, J., Janzing, D., Zscheischler, J., and Schölkopf, B · 2014
Later among the works it cites.
Learning and transferring mid-level image representations using convolutional neural networks
Oquab, M., Bottou, L., Laptev, I., and Sivic, J · 2014
Later among the works it cites.
Causal discovery with continuous additive noise models
Peters, J., Mooij, J. M., Janzing, D., and Schölkopf, B · 2014
Later among the works it cites.
Seeing the arrow of time
Pickup, L. C., Pan, Z., Wei, D., Shih, Y., Zhang, C., Zisserman, A., Schölkopf, B., and Freeman, W. T · 2014
Later among the works it cites.
Stochastic backpropagation and approximate inference in deep generative models
Rezende, D. J., Mohamed, S., and Wierstra, D · 2014
Later among the works it cites.
Derivatives and fisher information of bivariate copulas
Schepsmeier, U. and Stöber, J · 2014
Later among the works it cites.
Understanding Machine Learning: From Theory to Algorithms
Shalev-Shwartz, S. and Ben-David, S · 2014
Later among the works it cites.
Learning to transfer privileged information
Sharmanska, V., Quadrianto, N., and Lampert, C. H · 2014
Later among the works it cites.
Striving for simplicity: The all convolutional net
Springenberg, J. T., Dosovitskiy, A., Brox, T., and Riedmiller, M · 2014
Later among the works it cites.
Dropout: A simple way to prevent neural networks from overfitting
Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R · 2014
Later among the works it cites.
Two-stage sampled learning theory on distributions
Szabó, Z., Gretton, A., Póczos, B., and Sriperumbudur, B · 2014
Later among the works it cites.
Covariance kernels for fast automatic pattern discovery and extrapolation with Gaussian processes
Wilson, A. G · 2014
Later among the works it cites.
Random Laplace feature maps for semigroup kernels on histograms
Yang, J., Sindhwani, V., Fan, Q., Avron, H., and Mahoney, M · 2014
Later among the works it cites.
Deep learning
Bengio, Y., Goodfellow, I. J., and Courville, A · 2015
Later among the works it cites.
Convex optimization: Algorithms and complexity
Bubeck, S · 2015
Later among the works it cites.
The loss surfaces of multilayer networks
Choromanska, A., Henaff, M., Mathieu, M., Arous, G. B., and LeCun, Y · 2015
Later among the works it cites.
Linear dimensionality reduction: Survey, insights, and generalizations
Cunningham, J. P. and Ghahramani, Z · 2015
Later among the works it cites.
Train faster, generalize better: Stability of stochastic gradient descent
Hardt, M., Recht, B., and Singer, Y · 2015
Later among the works it cites.
Distilling the knowledge in a neural network
Hinton, G., Vinyals, O., and Dean, J · 2015
Later among the works it cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C · 2015
Later among the works it cites.
Convolutional neural networks for visual recognition, 2015
Karpathy, A · 2015
Later among the works it cites.
Deep learning
LeCun, Y., Bengio, Y., and Hinton, G · 2015
Later among the works it cites.
Towards a learning theory of cause-effect inference
Lopez-Paz, D., Muandet, K., Schölkopf, B., and Tolstikhin, I · 2015
Later among the works it cites.
A theoretical analysis of optimization by Gaussian continuation
Mobahi, H. and Fisher III, J. W · 2015
Later among the works it cites.
From Points to Probability Measures: A Statistical Learning on Distributions with Kernel Mean Embedding
Muandet, K · 2015
Later among the works it cites.
Causality
Peters, J · 2015
Later among the works it cites.
Semi-supervised learning with ladder network
Rasmus, A., Valpola, H., Honkala, M., Berglund, M., and Raiko, T · 2015
Later among the works it cites.
Optimal rates for random Fourier features
Sriperumbudur, B. K. and Szabó, Z · 2015
Later among the works it cites.
A note on the evaluation of generative models
Theis, L., Oord, A. v. d., and Bethge, M · 2015
Later among the works it cites.
An introduction to matrix concentration inequalities
Tropp, J. A · 2015
Later among the works it cites.
The probability and statistics cookbook, 2015
Vallentin, M · 2015
Later among the works it cites.
Learning using privileged information: Similarity control and knowledge transfer
Vapnik, V. and Izmailov, R · 2015
Later among the works it cites.
ResNet training in Torch, 2016
Gross, S · 2016
Closest in time.
Non-linear Causal Inference using Gaussianity Measures
Hernández-Lobato, D., Morales-Mombiela, P., Lopez-Paz, D., and Suárez, A · 2016
Closest in time.
No regret bound for extreme bandits
Nishihara, R., Lopez-Paz, D., and Bottou, L · 2016
Closest in time.
Minimax Estimation of Kernel Mean Embeddings
Tolstikhin, I., Sriperumbudur, B., and Muandet, K · 2016
Closest in time.
Lower bounds for realizable transductive learning
Tolstikhin, I. and Lopez-Paz, D · 2016
Closest in time.