Fetching the paper…
Reading the bibliography…
In this paper, we contend that the objective of representation learning is to compress and transform the distribution of the data, say sets of tokens, towards a mixture of low-dimensional Gaussian distributions supported on incoherent subspaces.
“Language models are few-shot learners”
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry and Amanda Askell · 1901
Earlier work this paper cites.
“Generative Modeling by Estimating Gradients of the Data Distribution”, 2019
Yang Song and Stefano Ermon · 1907
Earlier work this paper cites.
“Estimation of the Mean of a Multivariate Normal Distribution”
Charles Stein · 1981
Earlier work this paper cites.
“Sparse coding with an overcomplete basis set: A strategy employed by V1?”
Bruno Olshausen and David Field · 1997
Earlier work this paper cites.
“Image Manifolds which are Isometric to Euclidean Space”
David Donoho and Carrie Grimes · 2005
Earlier work this paper cites.
“Estimation of Non-Normalized Statistical Models by Score Matching”
Aapo Hyvärinen · 2005
Earlier work this paper cites.
“The multiscale structure of non-differentiable image manifolds”
Michael Wakin, David Donoho, Hyeokho Choi and Richard Baraniuk · 2005
Earlier work this paper cites.
“End-to-End Object Detection with Transformers”, 2020
Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov and Sergey Zagoruyko · 2005
Earlier work this paper cites.
“Segmentation of multivariate mixed data via lossy data coding and compression”
Yi Ma, Harm Derksen, Wei Hong and John Wright · 2007
Earlier work this paper cites.
“Solving Linear Inverse Problems Using the Prior Implicit in a Denoiser”, 2020
Zahra Kadkhodaie and Eero Simoncelli · 2007
Earlier work this paper cites.
“Automated flower classification over a large number of classes”
Maria-Elena Nilsback and Andrew Zisserman · 2008
Earlier work this paper cites.
“A fast iterative shrinkage-thresholding algorithm for linear inverse problems”
Amir Beck and Marc Teboulle · 2009
Earlier work this paper cites.
“Imagenet: A large-scale hierarchical image database”
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li and Li Fei-Fei · 2009
Earlier work this paper cites.
“Learning multiple layers of features from tiny images”
Alex Krizhevsky and Geoffrey Hinton · 2009
Earlier work this paper cites.
“Learning fast approximations of sparse coding”
Karol Gregor and Yann LeCun · 2010
Earlier work this paper cites.
“A Distribution-Free Theory of Nonparametric Regression”
László Györfi, Michael Kohler, Adam Krzyzak and Harro Walk · 2010
Earlier work this paper cites.
“Denoising Diffusion Implicit Models”, 2020
Jiaming Song, Chenlin Meng and Stefano Ermon · 2010
Earlier work this paper cites.
“Tweedie’s Formula and Selection Bias”
Bradley Efron · 2011
Earlier work this paper cites.
“Least squares estimation without priors or supervision”
Martin Raphan and Eero Simoncelli · 2011
Earlier work this paper cites.
“A connection between score matching and denoising autoencoders”
Pascal Vincent · 2011
Earlier work this paper cites.
“A Tour of Modern Image Filtering: New Insights and Methods, Both Practical and Theoretical”
Peyman Milanfar · 2011
Earlier work this paper cites.
“Score-Based Generative Modeling through Stochastic Differential Equations”, 2020
Yang Song, Jascha Sohl-Dickstein, Diederik Kingma, Abhishek Kumar, Stefano Ermon and Ben Poole · 2011
Earlier work this paper cites.
“Cats and dogs”
Omkar Parkhi, Andrea Vedaldi, Andrew Zisserman and CV Jawahar · 2012
Earlier work this paper cites.
“Exact Recovery of Sparsely-Used Dictionaries”, 2012
Daniel Spielman, Huan Wang and John Wright · 2012
Earlier work this paper cites.
“Invariant scattering convolution networks”
Joan Bruna and Stéphane Mallat · 2012
Earlier work this paper cites.
“Plug-and-Play priors for model based reconstruction”
Singanallur Venkatakrishnan, Charles Bouman and Brendt Wohlberg · 2013
Earlier work this paper cites.
“Sparse and spurious: dictionary learning with noise and outliers”, 2014
Rémi Gribonval, Rodolphe Jenatton and Francis Bach · 2014
Cited alongside, same era.
“Deep Unsupervised Learning using Nonequilibrium Thermodynamics”, 2015
Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan and Surya Ganguli · 2015
Cited alongside, same era.
“Deep Residual Learning for Image Recognition”
Kaiming He, Xiangyu Zhang, Shaoqing Ren and Jian Sun · 2016
Cited alongside, same era.
“Generalized Principal Component Analysis”
René Vidal, Yi Ma and Shankar Sastry · 2016
Cited alongside, same era.
Kaiming He, Georgia Gkioxari, Piotr Dollár and Ross Girshick · 2017
Cited alongside, same era.
“MLP-Mixer: An all-MLP Architecture for Vision”, 2021
Ilya Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Thomas Unterthiner, Jessica Yung, Andreas Steiner, Daniel Keysers, Jakob Uszkoreit, Mario Lucic and Alexey Dosovitskiy · 2021
Later among the works it cites.
“ReduNet: A White-box Deep Network from the Principle of Maximizing Rate Reduction”
Kwan Chan, Yaodong Yu, Chong You, Haozhi Qi, John Wright and Yi Ma · 2022
Later among the works it cites.
Hongrui Chen, Holden Lee and Jianfeng Lu · 2022
Later among the works it cites.
“Contrastive audio-visual masked autoencoder”
Yuan Gong, Andrew Rouditchenko, Alexander Liu, David Harwath, Leonid Karlinsky, Hilde Kuehne and James Glass · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ilya Loshchilov and Frank Hutter · 2017
Cited alongside, same era.
“The Little Engine That Could: Regularization by Denoising (RED)”
Yaniv Romano, Michael Elad and Peyman Milanfar · 2017
Cited alongside, same era.
“Attention is all you need”
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan Gomez, Łukasz Kaiser and Illia Polosukhin · 2017
Cited alongside, same era.
“The sparse manifold transform”
Yubei Chen, Dylan Paiton and Bruno Olshausen · 2018
Cited alongside, same era.
“Bert: Pre-training of deep bidirectional transformers for language understanding”
Jacob Devlin, Ming-Wei Chang, Kenton Lee and Kristina Toutanova · 2018
Cited alongside, same era.
“A Style-Based Generator Architecture for Generative Adversarial Networks”, 2018
Tero Karras, Samuli Laine and Timo Aila · 2018
Cited alongside, same era.
“Theoretical Foundations of Deep Learning via Sparse Representations: A Multilayer Sparse Model and Its Connection to Convolutional Neural Networks”
Vardan Papyan, Yaniv Romano, Jeremias Sulam and Michael Elad · 2018
Cited alongside, same era.
Hangfeng He and Weijie Su · 2022
Later among the works it cites.
“The Forward-Forward Algorithm: Some Preliminary Investigations”, 2022
Geoffrey Hinton · 2022
Later among the works it cites.
“Elucidating the design space of diffusion-based generative models”
Tero Karras, Miika Aittala, Timo Aila and Samuli Laine · 2022
Later among the works it cites.
“Statistical Efficiency of Score Matching: The View from Isoperimetry”, 2022
Frederic Koehler, Alexander Heckett and Andrej Risteski · 2022
Later among the works it cites.
“On the principles of parsimony and self-consistency for the emergence of intelligence”
Yi Ma, Doris Tsao and Heung-Yeung Shum · 2022
Later among the works it cites.
“Pursuit of a discriminative representation for multiple subspaces via sequential games”
Druv Pai, Michael Psenka, Chih-Yuan Chiu, Manxi Wu, Edgar Dobriban and Yi Ma · 2022
Later among the works it cites.
“Formal algorithms for transformers”
Mary Phuong and Marcus Hutter · 2022
Later among the works it cites.
“High-resolution image synthesis with latent diffusion models”
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser and Björn Ommer · 2022
Later among the works it cites.
“Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding”, 2022
Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily Denton, Seyed Ghasemipour, Burcu Ayan, S Sara, Rapha Lopes, Tim Salimans, Jonathan Ho, David Fleet and Mohammad Norouzi · 2022
Later among the works it cites.
“Understanding the Covariance Structure of Convolutional Filters”, 2022
Asher Trockman, Devin Willmott and J Zico · 2022
Later among the works it cites.
“Attention: Self-Expression Is All You Need” Unpublished; available: https://openreview.net/forum?id=MmujBClawFo , 2022
Rene Vidal · 2022
Later among the works it cites.
“Rethinking minimal sufficient representation in contrastive learning”
Haoqing Wang, Xun Guo, Zhi-Hong Deng and Yan Lu · 2022
Later among the works it cites.
“High-Dimensional Data Analysis with Low-Dimensional Models: Principles, Computation, and Applications”
John Wright and Yi Ma · 2022
Later among the works it cites.
Sitan Chen, Giannis Daras and Alexandros Dimakis · 2023
Closest in time.
“Symbolic discovery of optimization algorithms”
Xiangning Chen, Chen Liang, Da Huang, Esteban Real, Kaiyuan Wang, Yao Liu, Hieu Pham, Xuanyi Dong, Thang Luong and Cho-Jui Hsieh · 2023
Closest in time.
“Scaling vision transformers to 22 billion parameters”
Mostafa Dehghani, Josip Djolonga, Basil Mustafa, Piotr Padlewski, Jonathan Heek, Justin Gilmer, Andreas Steiner, Mathilde Caron, Robert Geirhos and Ibrahim Alabdulmohsin · 2023
Closest in time.
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander Berg, Wan-Yen Lo, Piotr Dollár and Ross Girshick · 2023
Closest in time.
Hongkang Li, Meng Wang, Sijia Liu and Pin-Yu Chen · 2023
Closest in time.
“The Lazy Neuron Phenomenon: On Emergence of Activation Sparsity in Transformers”
Zonglin Li, Chong You, Srinadh Bhojanapalli, Daliang Li, Ankit Rawat, Sashank Reddi, Ke Ye, Felix Chern, Felix Yu, Ruiqi Guo and Sanjiv Kumar · 2023
Closest in time.
“To Compress or Not to Compress–Self-Supervised Learning and Information Theory: A Review”
Ravid Shwartz-Ziv and Yann LeCun · 2023
Closest in time.
Yang Song, Prafulla Dhariwal, Mark Chen and Ilya Sutskever · 2023
Closest in time.