Fetching the paper…
Reading the bibliography…
Variational mutual information (MI) estimators are widely used in unsupervised representation learning methods such as contrastive predictive coding (CPC).
Asymptotic evaluation of certain markov process expectations for large time, I
Monroe D Donsker and S R Srinivasa Varadhan · 1975
Earlier work this paper cites.
Self-organization in a perceptual network
Ralph Linsker · 1988
Earlier work this paper cites.
An information-maximization approach to blind separation and blind deconvolution
Anthony J Bell and Terrence J Sejnowski · 1995
Earlier work this paper cites.
The IM algorithm: a variational approach to information maximization
David Barber and Felix V Agakov · 2003
Earlier work this paper cites.
Estimating divergence functionals and the likelihood ratio by convex risk minimization
Xuanlong Nguyen, Martin J Wainwright, and Michael I Jordan · 2008
Earlier work this paper cites.
Curriculum learning
Yoshua Bengio, Jérôme Louradour, Ronan Collobert, and Jason Weston · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Noise-Contrastive estimation of unnormalized statistical models, with applications to natural image statistics
Michael U Gutmann and Aapo Hyvärinen · 2012
Earlier work this paper cites.
Density ratio estimation in machine learning
Masashi Sugiyama, Taiji Suzuki, and Takafumi Kanamori · 2012
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean · 2013
Earlier work this paper cites.
Learning word embeddings efficiently with noise-contrastive estimation
Andriy Mnih and Koray Kavukcuoglu · 2013
Earlier work this paper cites.
Explaining and harnessing adversarial examples
Ian J Goodfellow, Jonathon Shlens, and Christian Szegedy · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
FaceNet: A unified embedding for face recognition and clustering
Florian Schroff, Dmitry Kalenichenko, and James Philbin · 2015
Earlier work this paper cites.
Learning deep representation for imbalanced classification
Chen Huang, Yining Li, Chen Change Loy, and Xiaoou Tang · 2016
Earlier work this paper cites.
Large-Margin softmax loss for convolutional neural networks
Weiyang Liu, Yandong Wen, Zhiding Yu, and Meng Yang · 2016
Earlier work this paper cites.
Conditional image generation with PixelCNN decoders
Aaron van den Oord, Nal Kalchbrenner, Oriol Vinyals, Lasse Espeholt, Alex Graves, and Koray Kavukcuoglu · 2016
Earlier work this paper cites.
Like what you like: Knowledge distill via neuron selectivity transfer
Zehao Huang and Naiyan Wang · 2017
Cited alongside, same era.
Rethinking feature discrimination and polymerization for large-scale recognition
Yu Liu, Hongyang Li, and Xiaogang Wang · 2017
Cited alongside, same era.
Learning to model the tail
Yu-Xiong Wang, Deva Ramanan, and Martial Hebert · 2017
Cited alongside, same era.
MINE: Mutual information neural estimation
Mohamed Ishmael Belghazi, Aristide Baratin, Sai Rajeswar, Sherjil Ozair, Yoshua Bengio, Aaron Courville, and R Devon Hjelm · 2018
Cited alongside, same era.
What is the effect of importance weighting in deep learning?
Jonathon Byrd and Zachary C Lipton · 2018
Class-Balanced loss based on effective number of samples
Yin Cui, Menglin Jia, Tsung-Yi Lin, Yang Song, and Serge Belongie · 2019
Later among the works it cites.
Video representation learning by dense predictive coding
Tengda Han, Weidi Xie, and Andrew Zisserman · 2019
Later among the works it cites.
Momentum contrast for unsupervised visual representation learning
Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick · 2019
Later among the works it cites.
Data-Efficient image recognition with contrastive predictive coding
Olivier J Hénaff, Ali Razavi, Carl Doersch, S M Ali Eslami, and Aaron van den Oord · 2019
Later among the works it cites.
Deep imbalanced learning for face recognition and attribute prediction
Chen Huang, Yining Li, Change Loy Chen, and Xiaoou Tang · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Cited alongside, same era.
Learning deep representations by mutual information estimation and maximization
R Devon Hjelm, Alex Fedorov, Samuel Lavoie-Marchildon, Karan Grewal, Phil Bachman, Adam Trischler, and Yoshua Bengio · 2018
Cited alongside, same era.
Paraphrasing complex network: Network compression via factor transfer
Jangho Kim, SeongUk Park, and Nojun Kwak · 2018
Cited alongside, same era.
Stronger data poisoning attacks break data sanitization defenses
Pang Wei Koh, Jacob Steinhardt, and Percy Liang · 2018
Cited alongside, same era.
Zhuang Ma and Michael Collins · 2018
Cited alongside, same era.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever · 2018
Cited alongside, same era.
Learning controllable fair representations
Jiaming Song, Pratyusha Kalluri, Aditya Grover, Shengjia Zhao, and Stefano Ermon · 2018
Cited alongside, same era.
Later among the works it cites.
Striking the right balance with uncertainty
Salman Khan, Munawar Hayat, Waqas Zamir, Jianbing Shen, and Ling Shao · 2019
Later among the works it cites.
Lit: Learned intermediate representation training for model compression
Animesh Koratana, Daniel Kang, Peter Bailis, and Matei Zaharia · 2019
Later among the works it cites.
Overfitting of neural nets under class imbalance: Analysis and improvements for segmentation
Zeju Li, Konstantinos Kamnitsas, and Ben Glocker · 2019
Later among the works it cites.
On variational bounds of mutual information
Ben Poole, Sherjil Ozair, Aaron van den Oord, Alexander A Alemi, and George Tucker · 2019
Later among the works it cites.
Bridging the gap between f f -gans and wasserstein gans
Jiaming Song and Stefano Ermon · 2019
Later among the works it cites.
Understanding the limitations of variational mutual information estimators
Jiaming Song and Stefano Ermon · 2019
Later among the works it cites.
Yonglong Tian, Dilip Krishnan, and Phillip Isola · 2019
Later among the works it cites.
Local aggregation for unsupervised learning of visual embeddings
Chengxu Zhuang, Alex Lin Zhai, and Daniel Yamins · 2019
Later among the works it cites.
Language models are few-shot learners
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Closest in time.
A simple framework for contrastive learning of visual representations
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton · 2020
Closest in time.
Improved baselines with momentum contrastive learning
Xinlei Chen, Haoqi Fan, Ross Girshick, and Kaiming He · 2020
Closest in time.
Formal limitations on the measurement of mutual information
David McAllester and Karl Stratos · 2020
Closest in time.
A theory of usable information under computational constraints
Yilun Xu, Shengjia Zhao, Jiaming Song, Russell Stewart, and Stefano Ermon · 2020
Closest in time.