Fetching the paper…
Reading the bibliography…
The cross-entropy loss commonly used in deep learning is closely related to the defining properties of optimal representations, but does not enforce some of the key properties.
G. E. Hinton and D. Van Camp, “Keeping the neural networks simple by minimizing the description length of the weights,” in Proceedings of the 6th annual conference on Computational learning theory . ACM, 1993, pp. 5–13
1993
Earlier work this paper cites.
N. Tishby, F. C. Pereira, and W. Bialek, “The information bottleneck method,” in The 37th annual Allerton Conference on Communication, Control, and Computing , 1999, pp. 368–377
1999
Earlier work this paper cites.
G. Sundaramoorthi, P. Petersen, V. S. Varadarajan, and S. Soatto, “On the set of images modulo viewpoint and contrast changes,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , June 2009
2009
Earlier work this paper cites.
A. Krizhevsky and G. Hinton, “Learning multiple layers of features from tiny images,” Technical report, University of Toronto , 2009
2009
Earlier work this paper cites.
J. Bruna and S. Mallat, “Classification with scattering operators,” in Proceedings of the 2011 IEEE Conference on Computer Vision and Pattern Recognition , ser. CVPR ’11, 2011, pp. 1561–1566
2011
Earlier work this paper cites.
D. P. Kingma and M. Welling, “Auto-encoding variational bayes,” in Proceedings of the 2nd International Conference on Learning Representations (ICLR) , no. 2014, 2013
2013
Earlier work this paper cites.
S. Wang and C. Manning, “Fast dropout training,” in Proceedings of the 30th International Conference on Machine Learning (ICML) , 2013, pp. 118–126
2013
Cited alongside, same era.
N. Srivastava, G. E. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov, “Dropout: a simple way to prevent neural networks from overfitting.” Journal of Machine Learning Research , vol. 15, no. 1, pp. 1929–1958, 2014
2014
Cited alongside, same era.
V. Mnih, N. Heess, A. Graves et al. , “Recurrent models of visual attention,” in Advances in Neural Information Processing Systems , 2014, pp. 2204–2212
2014
Cited alongside, same era.
2014
Cited alongside, same era.
2015
Later among the works it cites.
M. Abadi, A. Agarwal, P. Barham, E. Brevdo, Z. Chen, C. Citro, G. S. Corrado, A. Davis, J. Dean, M. Devin, S. Ghemawat, I. Goodfellow, A. Harp, G. Irving, M. Isard, Y. Jia, R. Jozefowicz, L. Kaiser, M. Kudlur, J. Levenberg, D. Mané, R. Monga, S. Moore, D. Murray, C. Olah, M. Schuster, J. Shlens, B. Steiner, I. Sutskever, K. Talwar, P. Tucker, V. Vanhoucke, V. Vasudevan, F. Viégas, O. Vinyals, P. Warden, M. Wattenberg, M. Wicke, Y. Yu, and X. Zheng, “TensorFlow: Large-scale machine learning on heterogeneous systems,” 2015, software available from tensorflow.org. [Online]. Available: http://tensorflow.org/
2015
Later among the works it cites.
S. Soatto and A. Chiuso, “Visual representations: Defining properties and deep approximations,” Proceedings of the International Conference on Learning Representations (ICLR); ArXiv: 1411.7676 , May 2016
2016
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
D. P. Kingma, T. Salimans, and M. Welling, “Variational dropout and the local reparameterization trick,” in Proceedings of the 28th International Conference on Neural Information Processing Systems , ser. NIPS’15, 2015, pp. 2575–2583
2015
Cited alongside, same era.
N. Tishby and N. Zaslavsky, “Deep learning and the information bottleneck principle,” in Information Theory Workshop (ITW), 2015 IEEE . IEEE, 2015, pp. 1–5
2015
Cited alongside, same era.
2016
Closest in time.
F. Anselmi, L. Rosasco, and T. Poggio, “On invariance and selectivity in representation learning,” Information and Inference , 2016
2016
Closest in time.
I. Higgins, L. Matthey, A. Pal, C. Burgess, X. Glorot, M. Botvinick, S. Mohamed, and A. Lerchner, “beta-vae: Learning basic visual concepts with a constrained variational framework,” in Proceedings of the International Conference on Learning Representations (ICLR) , 2017
2017
Closest in time.