The relaxation method of finding the common point of convex sets and its application to the solution of problems in convex programming
Lev M Bregman · 1967
Earlier work this paper cites.
Asymptotic evaluation of certain markov process expectations for large time. iv
Monroe D Donsker and SR Srinivasa Varadhan · 1983
Earlier work this paper cites.
Multilayer feedforward networks are universal approximators
K Hornik, M Stinchcombe, and H White · 1989
Earlier work this paper cites.
Word association norms, mutual information, and lexicography
Kenneth Ward Church and Patrick Hanks · 1990
Earlier work this paper cites.
The sample complexity of pattern classification with neural networks: the size of the weights is more important than the size of the network
Peter L Bartlett · 1998
Earlier work this paper cites.
Robust and efficient estimation by minimising a density power divergence
Ayanendranath Basu, Ian R Harris, Nils L Hjort, and MC Jones · 1998
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
Asymptotic statistics , volume 3
Aad W Van der Vaart · 2000
Earlier work this paper cites.
The im algorithm: a variational approach to information maximization
David Barber and Felix V Agakov · 2003
Earlier work this paper cites.
Estimating mutual information
Alexander Kraskov, Harald Stögbauer, and Peter Grassberger · 2004
Earlier work this paper cites.
Feature selection based on mutual information criteria of max-dependency, max-relevance, and min-redundancy
Hanchuan Peng, Fuhui Long, and Chris Ding · 2005
Earlier work this paper cites.
Direct importance estimation for covariate shift adaptation
Masashi Sugiyama, Taiji Suzuki, Shinichi Nakajima, Hisashi Kashima, Paul von Bünau, and Motoaki Kawanabe · 2008
Earlier work this paper cites.
Neural network learning: Theoretical foundations
Martin Anthony and Peter L Bartlett · 2009
Earlier work this paper cites.
Normalized (pointwise) mutual information in collocation extraction
Gerlof Bouma · 2009
Earlier work this paper cites.
A least-squares approach to direct importance estimation
Takafumi Kanamori, Shohei Hido, and Masashi Sugiyama · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky et al · 2009
Earlier work this paper cites.
Estimating divergence functionals and the likelihood ratio by convex risk minimization
XuanLong Nguyen, Martin J Wainwright, and Michael I Jordan · 2010
Earlier work this paper cites.