Fetching the paper…
Reading the bibliography…
Mutual information is widely applied to learn latent representations of observations, whilst its implication in classification neural networks remain to be better explained.
Maximum mutual information estimation of hidden markov model parameters for speech recognition
Lalit R Bahl, Peter F Brown, Peter V De Souza, and Robert L Mercer · 1986
Earlier work this paper cites.
Multilayer feedforward networks are universal approximators
Kurt Hornik, Maxwell Stinchcombe, and Halbert White · 1989
Earlier work this paper cites.
Empirical Processes in M-estimation
Sara A Geer and Sara van de Geer · 2000
Earlier work this paper cites.
The im algorithm: a variational approach to information maximization
David Barber and Felix V Agakov · 2003
Earlier work this paper cites.
A tutorial on energy-based learning
Yann Lecun, Sumit Chopra, Raia Hadsell, Marc Aurelio Ranzato, and Fu Jie Huang · 2006
Earlier work this paper cites.
Mnist handwritten digit database
Yann LeCun, Corinna Cortes, and CJ Burges · 2010
Earlier work this paper cites.
Rectified linear units improve restricted boltzmann machines
Vinod Nair and Geoffrey E Hinton · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
The caltech-ucsd birds-200-2011 dataset
Catherine Wah, Steve Branson, Peter Welinder, Pietro Perona, and Serge Belongie · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Rectifier nonlinearities improve neural network acoustic models
Andrew L Maas, Awni Y Hannun, and Andrew Y Ng · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Cited alongside, same era.
Network in network
Min Lin, Qiang Chen, and Shuicheng Yan · 2014
Cited alongside, same era.
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Cited alongside, same era.
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton · 2015
Cited alongside, same era.
Going deeper with convolutions
Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich · 2015
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Mutual information neural estimation
Mohamed Ishmael Belghazi, Aristide Baratin, Sai Rajeshwar, Sherjil Ozair, Yoshua Bengio, Aaron Courville, and Devon Hjelm · 2018
Later among the works it cites.
Improved techniques for weakly-supervised object localization
Junsuk Choe, Joo Hyun Park, and Hyunjung Shim · 2018
Later among the works it cites.
Large scale fine-grained categorization and domain-specific transfer learning
Yin Cui, Yang Song, Chen Sun, Andrew Howard, and Serge Belongie · 2018
Later among the works it cites.
Pairwise confusion for fine-grained visual classification
Abhimanyu Dubey, Otkrist Gupta, Pei Guo, Ramesh Raskar, Ryan Farrell, and Nikhil Naik · 2018
Later among the works it cites.
Formal limitations on the measurement of mutual information
David McAllester and Karl Statos · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Rethinking the inception architecture for computer vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna · 2016
Cited alongside, same era.
Learning deep features for discriminative localization
Bolei Zhou, Aditya Khosla, Agata Lapedriza, Aude Oliva, and Antonio Torralba · 2016
Cited alongside, same era.
Conditional image synthesis with auxiliary classifier gans
Augustus Odena, Christopher Olah, and Jonathon Shlens · 2017
Cited alongside, same era.
Automatic differentiation in pytorch
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer · 2017
Cited alongside, same era.
Infovae: Information maximizing variational autoencoders
Shengjia Zhao, Jiaming Song, and Stefano Ermon · 2017
Cited alongside, same era.
https://tiny-imagenet.herokuapp.com/
Tiny imagenet visual recognition challenge · 2019
Closest in time.
Attention-based dropout layer for weakly supervised object localization
Junsuk Choe and Hyunjung Shim · 2019
Closest in time.
Learning deep representations by mutual information estimation and maximization
R Devon Hjelm, Alex Fedorov, Samuel Lavoie-Marchildon, Karan Grewal, Phil Bachman, Adam Trischler, and Yoshua Bengio · 2019
Closest in time.
On variational bounds of mutual information
Ben Poole, Sherjil Ozair, Aaron van den Oord, Alexander A Alemi, and George Tucker · 2019
Closest in time.
Camdrop: A new explanation of dropout and a guided regularization method for deep neural networks
Hongjun Wang, Guangrun Wang, Guanbin Li, and Liang Lin · 2019
Closest in time.