On causal and anticausal learning
B. Schölkopf, D. Janzing, J. Peters, E. Sgouritsa, K. Zhang, and J. M. Mooij · 2012
Cited alongside, same era.
Machine learning in non-stationary environments: Introduction to covariate shift adaptation
M. Sugiyama and M. Kawanabe · 2012
Cited alongside, same era.
Making a science of model search: Hyperparameter optimization in hundreds of dimensions for vision architectures
J. Bergstra, D. Yamins, and D. Cox · 2013
Cited alongside, same era.
Domain generalization via invariant feature representation
K. Muandet, D. Balduzzi, and B. Schölkopf · 2013
Cited alongside, same era.
Adam: A method for stochastic optimization
Original
D. P. Kingma and J. Ba · 2014
Cited alongside, same era.
Domain-adversarial training of neural networks
Y. Ganin, E. Ustinova, H. Ajakan, P. Germain, H. Larochelle, F. Laviolette, M. Marchand, and V. Lempitsky · 2016
Cited alongside, same era.
Gradient descent converges to minimizers
Original
J. D. Lee, M. Simchowitz, M. I. Jordan, and B. Recht · 2016
Cited alongside, same era.
Causal inference by using invariant prediction: identification and confidence intervals
J. Peters, P. Bühlmann, and N. Meinshausen · 2016
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Original
C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals · 2016
Cited alongside, same era.
Conditional variance penalties and domain shift robustness
Original
C. Heinze-Deml and N. Meinshausen · 2017
Cited alongside, same era.
Three factors influencing minima in sgd
Original
S. Jastrzębski, Z. Kenton, D. Arpit, N. Ballas, A. Fischer, Y. Bengio, and A. Storkey · 2017
Cited alongside, same era.
Stochastic gradient descent as approximate bayesian inference
S. Mandt, M. D. Hoffman, and D. M. Blei · 2017
Cited alongside, same era.