Fetching the paper…
Reading the bibliography…
Deep neural networks excel in supervised learning tasks but are constrained by the need for extensive labeled data.
Relations between two sets of variates
Harold Hotelling · 1936
Earlier work this paper cites.
On distributions admitting a sufficient statistic
Bernard Osgood Koopman · 1936
Earlier work this paper cites.
Asymptotic evaluation of certain markov process expectations for large time, i
Monroe D Donsker and SR Srinivasa Varadhan · 1975
Earlier work this paper cites.
Maximum likelihood from incomplete data via the em algorithm
Arthur P Dempster, Nan M Laird, and Donald B Rubin · 1977
Earlier work this paper cites.
Sample estimate of the entropy of a random vector
Lyudmyla F Kozachenko and Nikolai N Leonenko · 1987
Earlier work this paper cites.
Self-organization in a perceptual network
Ralph Linsker · 1988
Earlier work this paper cites.
Self-organizing neural network that discovers surfaces in random-dot stereograms
Suzanna Becker and Geoffrey E Hinton · 1992
Earlier work this paper cites.
Signature verification using a” siamese” time delay neural network
Jane Bromley, Isabelle Guyon, Yann LeCun, Eduard Säckinger, and Roopak Shah · 1993
Earlier work this paper cites.
An information-maximization approach to blind separation and blind deconvolution
Anthony J Bell and Terrence J Sejnowski · 1995
Earlier work this paper cites.
Elements of information theory
Thomas M Cover · 1999
Earlier work this paper cites.
On the convergence of markovian stochastic algorithms with rapidly decreasing ergodicity rates
Laurent Younes · 1999
Earlier work this paper cites.
Kernel independent component analysis
Francis R Bach and Michael I Jordan · 2002
Earlier work this paper cites.
Information theoretic learning: Renyi’s entropy and its applications to adaptive system training
Deniz Erdogmus · 2002
Earlier work this paper cites.
Slow feature analysis: Unsupervised learning of invariances
Laurenz Wiskott and Terrence J. Sejnowski · 2002
Earlier work this paper cites.
Kernel independent component analysis
Francis R. Bach and Michael I. Jordan · 2003
Earlier work this paper cites.
Improved baselines with momentum contrastive learning
Xinlei Chen, Haoqi Fan, Ross Girshick, and Kaiming He · 2003
Earlier work this paper cites.
Estimation of entropy and mutual information
Liam Paninski · 2003
Earlier work this paper cites.
Canonical correlation analysis: An overview with application to learning methods
David R. Hardoon, Sandor Szedmak, and John Shawe-Taylor · 2004
Earlier work this paper cites.
Learning a similarity metric discriminatively, with application to face verification
Sumit Chopra, Raia Hadsell, and Yann LeCun · 2005
Earlier work this paper cites.
Entropy regularization., 2006
Yves Grandvalet and Yoshua Bengio · 2006
Earlier work this paper cites.
Dimensionality reduction by learning an invariant mapping
Raia Hadsell, Sumit Chopra, and Yann LeCun · 2006
Earlier work this paper cites.
Some extensions of score matching, 2006
Aapo Hyvärinen · 2006
Earlier work this paper cites.
Efficient sparse coding algorithms
Honglak Lee, Alexis Battle, Rajat Raina, and Andrew Ng · 2006
Earlier work this paper cites.
Scaling learning algorithms towards ai ,in l. bottou, o. chapelle, d. decoste, and j. weston, editors,
Yoshua Bengio and Yann LeCun · 2007
Earlier work this paper cites.
A maximum-likelihood interpretation for slow feature analysis
Richard Turner and Maneesh Sahani · 2007
Earlier work this paper cites.
Learning and generalization with the information bottleneck
Ohad Shamir, Sivan Sabato, and Naftali Tishby · 2008
Earlier work this paper cites.
An information theoretic framework for multi-view learning
Karthik Sridharan and Sham Kakade · 2008
Earlier work this paper cites.
Extracting and composing robust features with denoising autoencoders
Pascal Vincent, Hugo Larochelle, Yoshua Bengio, and Pierre-Antoine Manzagol · 2008
Earlier work this paper cites.
Semi-supervised learning (chapelle, o. et al., eds.; 2006)[book reviews]
Olivier Chapelle, Bernhard Scholkopf, and Alexander Zien · 2009
Earlier work this paper cites.
Speaker recognition by gaussian information bottleneck
Ron M Hecht, Elad Noor, and Naftali Tishby · 2009
Earlier work this paper cites.
A spiking neuron as information bottleneck
Lars Buesing and Wolfgang Maass · 2010
Earlier work this paper cites.
Factorized latent spaces with structured sparsity
Yangqing Jia, Mathieu Salzmann, and Trevor Darrell · 2010
Earlier work this paper cites.
A scalable two-stage approach for a class of dimensionality reduction techniques
Liang Sun, Betul Ceran, and Jieping Ye · 2010
Earlier work this paper cites.
A co-training approach for multi-view spectral clustering
Abhishek Kumar and Hal Daumé · 2011
Earlier work this paper cites.
Sparse autoencoder
Andrew Ng et al · 2011
Earlier work this paper cites.
Multimodal deep learning
Jiquan Ngiam, Aditya Khosla, Mingyu Kim, Juhan Nam, Honglak Lee, and Andrew Y. Ng · 2011
Earlier work this paper cites.
A connection between score matching and denoising autoencoders
Pascal Vincent · 2011
Earlier work this paper cites.
The information bottleneck em algorithm
Gal Elidan and Nir Friedman · 2012
Earlier work this paper cites.
Deep canonical correlation analysis
Galen Andrew, Raman Arora, Jeff Bilmes, and Karen Livescu · 2013
Earlier work this paper cites.
Representation learning: A review and new perspectives
Yoshua Bengio, Aaron Courville, and Pascal Vincent · 2013
Earlier work this paper cites.
Robust multimodal dictionary learning
Tian Cao, Vladimir Jojic, Shannon Modla, Debbie Powell, Kirk Czymmek, and Marc Niethammer · 2013
Earlier work this paper cites.
Multivariate information bottleneck
Nir Friedman, Ori Mosenzon, Noam Slonim, and Naftali Tishby · 2013
Earlier work this paper cites.
Generating sequences with recurrent neural networks
Alex Graves · 2013
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P Kingma and Max Welling · 2013
Earlier work this paper cites.
Pseudo-label: The simple and efficient semi-supervised learning method for deep neural networks
Dong-Hyun Lee et al · 2013
Earlier work this paper cites.
Multiview hessian discriminative sparse coding for image annotation
Weifeng Liu, Dacheng Tao, Jun Cheng, and Yuanyan Tang · 2013
Earlier work this paper cites.
The multi-feature information bottleneck with application to unsupervised image categorization
Zhengzheng Lou, Yangdong Ye, and Xiaoqiang Yan · 2013
Earlier work this paper cites.
A survey of multi-view machine learning
Shiliang Sun · 2013
Earlier work this paper cites.
Semi-supervised learning with deep generative models
Durk P Kingma, Shakir Mohamed, Danilo Jimenez Rezende, and Max Welling · 2014
Earlier work this paper cites.
Deep captioning with multimodal recurrent neural networks (m-rnn)
Junhua Mao, Wei Xu, Yi Yang, Jiang Wang, Zhiheng Huang, and Alan Yuille · 2014
Earlier work this paper cites.
C4. 5: programs for machine learning
J Ross Quinlan · 2014
Earlier work this paper cites.
Mutual information between discrete and continuous data sets
Brian C Ross · 2014
Earlier work this paper cites.
Multimodal learning with deep boltzmann machines
Nitish Srivastava and Ruslan Salakhutdinov · 2014
Earlier work this paper cites.
Large-margin multi-viewinformation bottleneck
Chang Xu, Dacheng Tao, and Chao Xu · 2014
Cited alongside, same era.
Long-term recurrent convolutional networks for visual recognition and description
Jeffrey Donahue, Lisa Anne Hendricks, Sergio Guadarrama, Marcus Rohrbach, Subhashini Venugopalan, Kate Saenko, and Trevor Darrell · 2015
Cited alongside, same era.
Efficient estimation of mutual information for strongly dependent variables
Shuyang Gao, Greg Ver Steeg, and Aram Galstyan · 2015
Cited alongside, same era.
Made: Masked autoencoder for distribution estimation
Mathieu Germain, Karol Gregor, Iain Murray, and Hugo Larochelle · 2015
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Cited alongside, same era.
On mutual information maximization for representation learning
Michael Tschannen, Josip Djolonga, Paul K Rubenstein, Sylvain Gelly, and Mario Lucic · 2019
Later among the works it cites.
Deep Multi-view Information Bottleneck , pages 37–45
Qi Wang, Claire Boudreau, Qixing Luo, Pang-Ning Tan, and Jiayu Zhou · 2019
Later among the works it cites.
Deep low-rank subspace ensemble for multi-view clustering
Zhe Xue, Junping Du, Dawei Du, and Siwei Lyu · 2019
Later among the works it cites.
S4l: Self-supervised semi-supervised learning
Xiaohua Zhai, Avital Oliver, Alexander Kolesnikov, and Lucas Beyer · 2019
Later among the works it cites.
Survey on deep neural networks in speech and vision systems
Mahbubul Alam, Manar D Samad, Lasitha Vidyaratne, Alexander Glandon, and Khan M Iftekharuddin · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Andrej Karpathy and Li Fei-Fei · 2015
Cited alongside, same era.
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton · 2015
Cited alongside, same era.
Alireza Makhzani, Jonathon Shlens, Navdeep Jaitly, Ian Goodfellow, and Brendan Frey · 2015
Cited alongside, same era.
Predictive information in a sensory population
Stephanie E Palmer, Olivier Marre, Michael J Berry, and William Bialek · 2015
Cited alongside, same era.
Variational inference with normalizing flows
Danilo Rezende and Shakir Mohamed · 2015
Cited alongside, same era.
Unsupervised and semi-supervised learning with categorical generative adversarial networks
Jost Tobias Springenberg · 2015
Cited alongside, same era.
On deep multi-view representation learning
Weiran Wang, Raman Arora, Karen Livescu, and Jeff Bilmes · 2015
Cited alongside, same era.
Mathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal, Piotr Bojanowski, and Armand Joulin · 2020
Later among the works it cites.
What information does a resnet compress?
Luke Nicholas Darlow and Amos Storkey · 2020
Later among the works it cites.
Learning robust representations via multi-view information bottleneck
Marco Federici, Anjan Dutta, Patrick Forré, Nate Kushman, and Zeynep Akata · 2020
Later among the works it cites.
The conditional entropy bottleneck
Ian Fischer · 2020
Later among the works it cites.
On information plane analyses of neural network classifiers–a review
Bernhard C Geiger · 2020
Later among the works it cites.
Bootstrap your own latent-a new approach to self-supervised learning
Jean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec, Pierre Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Avila Pires, Zhaohan Guo, Mohammad Gheshlaghi Azar, et al · 2020
Later among the works it cites.
Momentum contrast for unsupervised visual representation learning
Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick · 2020
Later among the works it cites.
Data-efficient image recognition with contrastive predictive coding
Olivier Henaff · 2020
Later among the works it cites.
Self-supervised learning of pretext-invariant representations
Ishan Misra and Laurens van der Maaten · 2020
Later among the works it cites.
The dual information bottleneck
Zoe Piran, Ravid Shwartz-Ziv, and Naftali Tishby · 2020
Later among the works it cites.
Multimodal topic learning for video recommendation
Shi Pu, Yijiang He, Zheng Li, and Mao Zheng · 2020
Later among the works it cites.
Information in infinite ensembles of infinitely-wide neural networks
Ravid Shwartz-Ziv and Alexander A Alemi · 2020
Later among the works it cites.
Fixmatch: Simplifying semi-supervised learning with consistency and confidence
Kihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang, Han Zhang, Colin A Raffel, Ekin Dogus Cubuk, Alexey Kurakin, and Chun-Liang Li · 2020
Later among the works it cites.
Score-based generative modeling through stochastic differential equations
Yang Song, Jascha Sohl-Dickstein, Diederik P Kingma, Abhishek Kumar, Stefano Ermon, and Ben Poole · 2020
Later among the works it cites.
Reasoning about generalization via conditional mutual information
Thomas Steinke and Lydia Zakynthinou · 2020
Later among the works it cites.
Self-supervised learning from a multi-view perspective
Yao-Hung Hubert Tsai, Yue Wu, Ruslan Salakhutdinov, and Louis-Philippe Morency · 2020
Later among the works it cites.
Variational information bottleneck for unsupervised clustering: Deep gaussian mixture embedding
Yiğit Uğur, George Arvanitakis, and Abdellatif Zaidi · 2020
Later among the works it cites.
Variational information bottleneck for semi-supervised classification
Slava Voloshynovskiy, Olga Taran, Mouad Kondah, Taras Holotyak, and Danilo Rezende · 2020
Later among the works it cites.
Understanding contrastive representation learning through alignment and uniformity on the hypersphere
Tongzhou Wang and Phillip Isola · 2020
Later among the works it cites.
How good is the bayes posterior in deep neural networks really?
Florian Wenzel, Kevin Roth, Bastiaan S Veeling, Jakub Świkatkowski, Linh Tran, Stephan Mandt, Jasper Snoek, Tim Salimans, Rodolphe Jenatton, and Sebastian Nowozin · 2020
Later among the works it cites.
Unsupervised data augmentation for consistency training
Qizhe Xie, Zihang Dai, Eduard Hovy, Thang Luong, and Quoc Le · 2020
Later among the works it cites.
A theory of usable information under computational constraints
Yilun Xu, Shengjia Zhao, Jiaming Song, Russell Stewart, and Stefano Ermon · 2020
Later among the works it cites.
Tabnet: Attentive interpretable tabular learning
Sercan Ö Arik and Tomas Pfister · 2021
Later among the works it cites.
Vicreg: Variance-invariance-covariance regularization for self-supervised learning
Adrien Bardes, Jean Ponce, and Yann LeCun · 2021
Later among the works it cites.
A geometric perspective on information plane analysis
Mina Basirat, Bernhard C. Geiger, and Peter M. Roth · 2021
Later among the works it cites.
Emerging properties in self-supervised vision transformers
Mathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou, Julien Mairal, Piotr Bojanowski, and Armand Joulin · 2021
Later among the works it cites.
Exploring simple siamese representation learning
Xinlei Chen and Kaiming He · 2021
Later among the works it cites.
Lossy compression for lossless prediction
Yann Dubois, Benjamin Bloem-Reddy, Karen Ullrich, and Chris J Maddison · 2021
Later among the works it cites.
Sliced mutual information: A scalable measure of statistical dependence
Ziv Goldfeld and Kristjan Greenewald · 2021
Later among the works it cites.
Understanding dimensional collapse in contrastive self-supervised learning
Li Jing, Pascal Vincent, Yann LeCun, and Yuandong Tian · 2021
Later among the works it cites.
How to train your energy-based models
Yang Song and Diederik P Kingma · 2021
Later among the works it cites.
Subtab: Subsetting features of tabular data for self-supervised representation learning
Talip Ucar, Ehsan Hajiramezanali, and Lindsay Edwards · 2021
Later among the works it cites.
Deep multi-view learning methods: A review
Xiaoqiang Yan, Shizhe Hu, Yiqiao Mao, Yangdong Ye, and Hui Yu · 2021
Later among the works it cites.
Barlow twins: Self-supervised learning via redundancy reduction
Jure Zbontar, Li Jing, Ishan Misra, Yann LeCun, and Stéphane Deny · 2021
Later among the works it cites.
Contrastive learning inverts the data generating process
Roland S Zimmermann, Yash Sharma, Steffen Schneider, Matthias Bethge, and Wieland Brendel · 2021
Later among the works it cites.
Detreg: Unsupervised pretraining with region priors for object detection
Amir Bar, Xin Wang, Vadim Kantorov, Colorado J Reed, Roei Herzig, Gal Chechik, Anna Rohrbach, Trevor Darrell, and Amir Globerson · 2022
Later among the works it cites.
An information maximization based blind source separation approach for dependent and independent sources
Alper T Erdogan · 2022
Later among the works it cites.
Jonas Geiping, Micah Goldblum, Gowthami Somepalli, Ravid Shwartz-Ziv, Tom Goldstein, and Andrew Gordon Wilson · 2022
Later among the works it cites.
k-sliced mutual information: A quantitative study of scalability with dimension
Ziv Goldfeld, Kristjan Greenewald, Theshani Nuradha, and Galen Reeves · 2022
Later among the works it cites.
Masked autoencoders are scalable vision learners
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick · 2022
Later among the works it cites.
The physics of energy-based models
Patrick Huembeli, Juan Miguel Arrazola, Nathan Killoran, Masoud Mohseni, and Peter Wittek · 2022
Later among the works it cites.
A contrastive objective for learning disentangled representations
Jonathan Kahana and Yedid Hoshen · 2022
Later among the works it cites.
Self-supervised learning with an information maximization criterion
Serdar Ozsoy, Shadi Hamdan, Sercan Arik, Deniz Yuret, and Alper Erdogan · 2022
Later among the works it cites.
Information flow in deep neural networks
Ravid Shwartz-Ziv · 2022
Later among the works it cites.
Tabular data: Deep learning is not all you need
Ravid Shwartz-Ziv and Amitai Armon · 2022
Later among the works it cites.
Rethinking minimal sufficient representation in contrastive learning
Haoqing Wang, Xun Guo, Zhi-Hong Deng, and Yan Lu · 2022
Later among the works it cites.
Reverse engineering self-supervised learning
Ido Ben-Shaul, Ravid Shwartz-Ziv, Tomer Galanti, Shai Dekel, and Yann LeCun · 2023
Closest in time.
An information-theoretic perspective on variance-invariance-covariance regularization
Ravid Shwartz-Ziv, Randall Balestriero, Kenji Kawaguchi, Tim GJ Rudner, and Yann LeCun · 2023
Closest in time.