Fetching the paper…
Reading the bibliography…
One of the most surprising and exciting discoveries in supervised learning was the benefit of overparameterization (i.e.
Probabilistic reasoning in intelligent systems: Networks of plausible inference
Pearl, J · 1988
Earlier work this paper cites.
Building a Large Annotated Corpus of English: The Penn Treebank
Marcus, M. P., Marcinkiewicz, M. A., and Santorini, B · 1993
Earlier work this paper cites.
A Generative Constituent-Context Model for Improved Grammar Induction
Klein, D. and Manning, C · 2002
Earlier work this paper cites.
Noisy-or component analysis and its application to link analysis
Šingliar, T. and Hauskrecht, M · 2006
Earlier work this paper cites.
A probabilistic analysis of em for mixtures of separated, spherical gaussians
Dasgupta, S. and Schulman, L · 2007
Earlier work this paper cites.
Graphical models, exponential families, and variational inference
Wainwright, M. J., Jordan, M. I., et al · 2008
Earlier work this paper cites.
Nonnegative dictionary learning in the exponential noise model for adaptive music signal representation
Dikmen, O. and Févotte, C · 2011
Earlier work this paper cites.
Unsupervised learning of noisy-or bayesian networks
Halpern, Y. and Sontag, D · 2013
Earlier work this paper cites.
Discovering hidden variables in noisy-or networks using quartet tests
Jernite, Y., Halpern, Y., and Sontag, D · 2013
Cited alongside, same era.
Uci machine learning repository, 2013
Lichman, M. et al · 2013
Cited alongside, same era.
Tensor decompositions for learning latent variable models
Anandkumar, A., Ge, R., Hsu, D., Kakade, S. M., and Telgarsky, M · 2014
Cited alongside, same era.
Auto-Encoding Variational Bayes
Kingma, D. P. and Welling, M · 2014
Cited alongside, same era.
Neural variational inference and learning in belief networks
Mnih, A. and Gregor, K · 2014
Cited alongside, same era.
Stochastic Backpropagation and Approximate Inference in Deep Generative Models
Rezende, D. J., Mohamed, S., and Wierstra, D · 2014
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Zhang, C., Bengio, S., Hardt, M., Recht, B., and Vinyals, O · 2016
Later among the works it cites.
Provable learning of noisy-or networks
Arora, S., Ge, R., Ma, T., and Risteski, A · 2017
Later among the works it cites.
Tackling over-pruning in variational autoencoders
Yeung, S., Kannan, A., Dauphin, Y., and Fei-Fei, L · 2017
Later among the works it cites.
Learning and generalization in overparameterized neural networks, going beyond two layers
Allen-Zhu, Z., Li, Y., and Liang, Y · 2018
Later among the works it cites.
Algorithmic regularization in over-parameterized matrix sensing and neural networks with quadratic activations
Li, Y., Ma, T., and Zhang, H · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Reliable and scalable variational inference for the hierarchical dirichlet process
Hughes, M., Kim, D. I., and Sudderth, E · 2015
Cited alongside, same era.
Recovery guarantee of non-negative matrix factorization via alternating updates
Li, Y., Liang, Y., and Risteski, A · 2016
Cited alongside, same era.
Benefits of over-parameterization with em
Xu, J., Hsu, D. J., and Maleki, A · 2018
Later among the works it cites.
Can sgd learn recurrent neural networks with provable generalization?
Allen-Zhu, Z. and Li, Y · 2019
Closest in time.
Compound Probabilistic Context-Free Grammars for Grammar Induction
Kim, Y., Dyer, C., and Rush, A. M · 2019
Closest in time.