Fetching the paper…
Reading the bibliography…
Statistical learning relies upon data sampled from a distribution, and we usually do not care what actually generated it in the first place.
Adaptive mixtures of local experts
Jacobs, R. A., Jordan, M. I., Nowlan, S. J., and Hinton, G. E · 1991
Earlier work this paper cites.
Hierarchical mixtures of experts and the EM algorithm
Jordan, M. I. and Jacobs, R. A · 1994
Earlier work this paper cites.
Causality
Pearl, J · 2000
Earlier work this paper cites.
Inferring deterministic causal relations
Daniušis, P., Janzing, D., Mooij, J., Zscheischler, J., Steudel, B., Zhang, K., and Schölkopf, B · 2010
Earlier work this paper cites.
Causal inference using the algorithmic Markov condition
Janzing, D. and Schölkopf, B · 2010
Earlier work this paper cites.
On causal and anticausal learning
Schölkopf, B., Janzing, D., Peters, J., Sgouritsa, E., Zhang, K., and Mooij, J. M · 2012
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. P. and Welling, M · 2013
Earlier work this paper cites.
Generative adversarial nets
Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. and Ba, J · 2014
Cited alongside, same era.
Fast and accurate deep network learning by exponential linear units (elus)
Clevert, D.-A., Unterthiner, T., and Hochreiter, S · 2015
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C · 2015
Cited alongside, same era.
Human-level concept learning through probabilistic program induction
Lake, B. M., Salakhutdinov, R., and Tenenbaum, J. B · 2015
Cited alongside, same era.
Causal transfer in machine learning
Rojas-Carulla, M., Schölkopf, B., Turner, R., and Peters, J · 2015
Cited alongside, same era.
Stochastic multiple choice learning for training diverse deep ensembles
Lee, S., Purushwalkam Shiva Prakash, S., Cogswell, M., Ranjan, V., Crandall, D., and Batra, D · 2016
Later among the works it cites.
Expert gate: Lifelong learning with a network of experts
Aljundi, R., Chakravarty, P., and Tuytelaars, T · 2017
Closest in time.
Unsupervised pixel-level domain adaptation with generative adversarial networks
Bousmalis, K., Silberman, N., Dohan, D., Erhan, D., and Krishnan, D · 2017
Closest in time.
beta-vae: Learning basic visual concepts with a constrained variational framework
Higgins, I., Matthey, L., Pal, A., Burgess, C., Glorot, X., Botvinick, M., Mohamed, S., and Lerchner, A · 2017
Closest in time.
Elements of Causal Inference
Peters, J., Janzing, D., and Schölkopf, B · 2017
Closest in time.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Distinguishing cause from effect based on exogeneity
Zhang, K., Zhang, J., and Schölkopf, B · 2015
Cited alongside, same era.
Infogan: Interpretable representation learning by information maximizing generative adversarial nets
Chen, X., Duan, Y., Houthooft, R., Schulman, J., Sutskever, I., and Abbeel, P · 2016
Cited alongside, same era.
Unsupervised feature extraction by time-contrastive learning and nonlinear ica
Hyvarinen, A. and Morioka, H · 2016
Cited alongside, same era.
Shazeer, N., Mirhoseini, A., Maziarz, K., Davis, A., Le, Q., Hinton, G., and Dean, J · 2017
Closest in time.
Adversarial discriminative domain adaptation
Tzeng, E., Hoffman, J., Saenko, K., and Darrell, T · 2017
Closest in time.