Fetching the paper…
Reading the bibliography…
Maximum likelihood is the most widely used statistical estimation technique.
R. A. Fisher, “On the mathematical foundations of theoretical statistics,” Philosophical Transactions of the Royal Society of London. Series A, Containing Papers of a Mathematical or Physical Character , pp. 309–368, 1922
1922
Earlier work this paper cites.
——, “Theory of statistical estimation,” in Mathematical Proceedings of the Cambridge Philosophical Society , vol. 22, no. 05. Cambridge Univ Press, 1925, pp. 700–725
1925
Earlier work this paper cites.
——, “Two new properties of mathematical likelihood,” Proceedings of the Royal Society of London. Series A , vol. 144, no. 852, pp. 285–307, 1934
1934
Earlier work this paper cites.
C. E. Shannon, “A mathematical theory of communication,” The Bell System Technical Journal , vol. 27, pp. 379–423, 623–656, 1948
1948
Earlier work this paper cites.
C. Stein, “Inadmissibility of the usual estimator for the mean of a multivariate normal distribution,” in Proceedings of the Third Berkeley symposium on mathematical statistics and probability , vol. 1, no. 399, 1956, pp. 197–206
1956
Earlier work this paper cites.
W. James and C. Stein, “Estimation with quadratic loss,” in Proceedings of the fourth Berkeley symposium on mathematical statistics and probability , vol. 1, no. 1961, 1961, pp. 361–379
1961
Earlier work this paper cites.
C. Chow and C. Liu, “Approximating discrete probability distributions with dependence trees,” Information Theory, IEEE Transactions on , vol. 14, no. 3, pp. 462–467, 1968
1968
Earlier work this paper cites.
C. Chow and T. Wagner, “Consistency of an estimate of tree-dependent probability distributions (corresp.),” Information Theory, IEEE Transactions on , vol. 19, no. 3, pp. 369–371, 1973
1973
Earlier work this paper cites.
L. Le Cam, “Maximum likelihood: an introduction,” Statistics Branch, Department of Mathematics, University of Maryland , 1979
1979
Earlier work this paper cites.
J. Berkson, “Minimum chi-square, not maximum likelihood!” The Annals of Statistics , pp. 457–487, 1980
1980
Earlier work this paper cites.
B. Efron, “Maximum likelihood and decision theory,” The Annals of Statistics , pp. 340–356, 1982
1982
Earlier work this paper cites.
I. Ibragimov, A. Nemirovskii, and R. Khas’ minskii, “Some problems on nonparametric estimation in Gaussian white noise,” Theory of Probability & Its Applications , vol. 31, no. 3, pp. 391–406, 1987
1987
Earlier work this paper cites.
J. L. Massey, “Causality, feedback, and directed information,” in Proc. Int. Symp. Inf. Theory Appl. , Honolulu, HI, Nov. 1990, pp. 303–305
1990
Earlier work this paper cites.
D. L. Donoho and J. M. Johnstone, “Ideal spatial adaptation by wavelet shrinkage,” Biometrika , vol. 81, no. 3, pp. 425–455, 1994
1994
Earlier work this paper cites.
P. Murphy and D. W. Aha, “UCI repository of machine learning databases–a machine-readable repository,” 1995
1995
Cited alongside, same era.
N. Friedman, D. Geiger, and M. Goldszmidt, “Bayesian network classifiers,” Machine learning , vol. 29, no. 2-3, pp. 131–163, 1997
1997
Cited alongside, same era.
E. L. Lehmann and G. Casella, Theory of point estimation . Springer, 1998, vol. 31
1998
Cited alongside, same era.
O. Lepski, A. Nemirovski, and V. Spokoiny, “On estimation of the L r {L}_{r} norm of a regression function,” Probability theory and related fields , vol. 113, no. 2, pp. 221–253, 1999
1999
Cited alongside, same era.
A. W. Van der Vaart, Asymptotic statistics . Cambridge university press, 2000, vol. 3
2000
Cited alongside, same era.
M. J. Wainwright and M. I. Jordan, “Graphical models, exponential families, and variational inference,” Foundations and Trends® in Machine Learning , vol. 1, no. 1-2, pp. 1–305, 2008
2008
Later among the works it cites.
D. Koller and N. Friedman, Probabilistic graphical models: principles and techniques . MIT press, 2009
2009
Later among the works it cites.
T. Hastie, R. Tibshirani, and J. Friedman, The elements of statistical learning . Springer, 2009, vol. 2, no. 1
2009
Later among the works it cites.
Y. Polyanskiy, H. V. Poor, and S. Verdú, “Channel coding rate in the finite blocklength regime,” Information Theory, IEEE Transactions on , vol. 56, no. 5, pp. 2307–2359, 2010
2010
Later among the works it cites.
T. T. Cai and M. G. Low, “Testing composite hypotheses, Hermite polynomials and optimal estimation of a nonsmooth functional,” The Annals of Statistics , vol. 39, no. 2, pp. 1012–1041, 2011
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
I. Nemenman, F. Shafee, and W. Bialek, “Entropy and inference, revisited,” NIPS , 2002
2002
Cited alongside, same era.
L. Paninski, “Estimation of entropy and mutual information,” Neural Computation , vol. 15, no. 6, pp. 1191–1253, 2003
2003
Cited alongside, same era.
L. Paninski, “Variational minimax estimation of discrete distributions under KL loss,” in Advances in Neural Information Processing Systems , 2004, pp. 1033–1040
2004
Cited alongside, same era.
E. J. Candes and T. Tao, “Near-optimal signal recovery from random projections: Universal encoding strategies?” Information Theory, IEEE Transactions on , vol. 52, no. 12, pp. 5406–5425, 2006
2006
Cited alongside, same era.
D. L. Donoho, “Compressed sensing,” Information Theory, IEEE Transactions on , vol. 52, no. 4, pp. 1289–1306, 2006
2006
Cited alongside, same era.
T. M. Cover and J. A. Thomas, Elements of Information Theory , 2nd ed. New York: Wiley, 2006
2006
Cited alongside, same era.
N. Cesa-Bianchi and G. Lugosi, Prediction, learning, and games . Cambridge University Press, 2006
2006
Cited alongside, same era.
2011
Later among the works it cites.
G. Valiant and P. Valiant, “Estimating the unseen: an n / log n n/\log n -sample estimator for entropy and support size, shown optimal via new CLTs,” in Proceedings of the 43rd annual ACM symposium on Theory of computing . ACM, 2011, pp. 685–694
2011
Later among the works it cites.
——, “The power of linear estimators,” in Foundations of Computer Science (FOCS), 2011 IEEE 52nd Annual Symposium on . IEEE, 2011, pp. 403–412
2011
Later among the works it cites.
2011
Later among the works it cites.
V. Y. Tan, A. Anandkumar, L. Tong, and A. S. Willsky, “A large-deviation analysis of the maximum-likelihood learning of markov tree structures,” Information Theory, IEEE Transactions on , vol. 57, no. 3, pp. 1714–1735, 2011
2011
Later among the works it cites.
2013
Later among the works it cites.
J. Jiao, K. Venkat, Y. Han, and T. Weissman, “Minimax estimation of functionals of discrete distributions,” available on arXiv , 2014
2014
Closest in time.
2014
Closest in time.