Fetching the paper…
Reading the bibliography…
Deep neural networks are typically trained under a supervised learning framework where a model learns a single task using labeled data.
Cloze-driven pretraining of self-attention networks
A. Baevski, S. Edunov, Y. Liu, L. Zettlemoyer, and M. Auli · 1903
Earlier work this paper cites.
Ernie: Enhanced representation through knowledge integration
Y. Sun, S. Wang, Y. Li, S. Feng, X. Chen, H. Zhang, X. Tian, D. Zhu, H. Tian, and H. Wu · 1904
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, and V. Stoyanov · 1907
Earlier work this paper cites.
vq-wav2vec: Self-supervised learning of discrete speech representations
A. Baevski, S. Schneider, and M. Auli · 1910
Earlier work this paper cites.
Laws of organization in perceptual forms
M. Wertheimer · 1938
Earlier work this paper cites.
A learning algorithm for boltzmann machines
D. H. Ackley, G. E. Hinton, and T. J. Sejnowski · 1985
Earlier work this paper cites.
Learning internal representations from gray-scale images: An example of extensional programming
G. W. Cottrell, P. Munro, and D. Zipser · 1987
Earlier work this paper cites.
Principal component analysis
S. Wold, K. Esbensen, and P. Geladi · 1987
Earlier work this paper cites.
Learning from hints in neural networks
Y. S. Abu-Mostafa · 1990
Earlier work this paper cites.
Face recognition using unsupervised feature extraction
G. W. Cottrell and M. Fleming · 1990
Earlier work this paper cites.
Indexing by latent semantic analysis
S. C. Deerwester, S. T. Dumais, T. K. Landauer, G. W. Furnas, and R. A. Harshman · 1990
Earlier work this paper cites.
Categorization of faces using unsupervised feature extraction
M. K. Fleming and G. W. Cottrell · 1990
Earlier work this paper cites.
Empath: Face, emotion, and gender recognition using holons
G. W. Cottrell and J. Metcalfe · 1991
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
R. J. Williams · 1992
Earlier work this paper cites.
Non-linear dimensionality reduction
D. DeMers and G. W. Cottrell · 1993
Earlier work this paper cites.
An information-maximization approach to blind separation and blind deconvolution
A. J. Bell and T. J. Sejnowski · 1995
Earlier work this paper cites.
Multitask learning
R. Caruana · 1997
Earlier work this paper cites.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Sparse coding with an overcomplete basis set: A strategy employed by v1?
B. A. Olshausen and D. J. Field · 1997
Earlier work this paper cites.
Learning to recognize objects
G. Wallis and H. Bülthoff · 1999
Earlier work this paper cites.
A model of inductive bias learning
J. Baxter · 2000
Earlier work this paper cites.
A perspective view and survey of meta-learning
R. Vilalta and Y. Drissi · 2002
Earlier work this paper cites.
Slow feature analysis: Unsupervised learning of invariances
L. Wiskott and T. J. Sejnowski · 2002
Earlier work this paper cites.
Language models are few-shot learners
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei · 2005
Earlier work this paper cites.
Language models are few-shot learners
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2005
Earlier work this paper cites.
Learning a similarity metric discriminatively, with application to face verification
S. Chopra, R. Hadsell, and Y. LeCun · 2005
Earlier work this paper cites.
Dimensionality reduction by learning an invariant mapping
R. Hadsell, S. Chopra, and Y. LeCun · 2006
Earlier work this paper cites.
Reducing the dimensionality of data with neural networks
G. E. Hinton and R. R. Salakhutdinov · 2006
Earlier work this paper cites.
A fast learning algorithm for deep belief nets
G. E. Hinton, S. Osindero, and Y. W. Teh · 2006
Earlier work this paper cites.
Scaling learning algorithms towards AI
Y. Bengio and Y. Lecun · 2007
Earlier work this paper cites.
Greedy layer-wise training of deep networks
Y. Bengio, P. Lamblin, D. Popovici, and H. Larochelle · 2007
Earlier work this paper cites.
Sparse feature learning for deep belief networks
M. Ranzato, Y.-L. Boureau, and Y. L. Cun · 2008
Earlier work this paper cites.
Extracting and composing robust features with denoising autoencoders
P. Vincent, H. Larochelle, Y. Bengio, and P. Manzagol · 2008
Earlier work this paper cites.
A survey on transfer learning
S. J. Pan and Q. Yang · 2009
Earlier work this paper cites.
Why does unsupervised pre-training help deep learning?
D. Erhan, Y. Bengio, A. C. Courville, P. Manzagol, P. Vincent, and S. Bengio · 2010
Earlier work this paper cites.
Rectified linear units improve restricted boltzmann machines
V. Nair and G. E. Hinton · 2010
Earlier work this paper cites.
What the no free lunch theorems really mean; how to improve search algorithms
D. H. Wolpert · 2012
Earlier work this paper cites.
Representation learning: A review and new perspectives
Y. Bengio, A. Courville, and P. Vincent · 2013
Earlier work this paper cites.
Auto-encoding variational bayes
D. P. Kingma and M. Welling · 2013
Earlier work this paper cites.
A. Makhzani and B. Frey · 2013
Cited alongside, same era.
Efficient estimation of word representations in vector space
T. Mikolov, K. Chen, G. Corrado, and J. Dean · 2013
Cited alongside, same era.
Distributed representations of words and phrases and their compositionality
T. Mikolov, I. Sutskever, K. Chen, G. S. Corrado, and J. Dean · 2013
Cited alongside, same era.
Generative adversarial nets
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio · 2014
Cited alongside, same era.
Semi-supervised sequence learning
A. M. Dai and Q. V. Le · 2015
Cited alongside, same era.
Unsupervised visual representation learning by context prediction
C. Doersch, A. Gupta, and A. A. Efros · 2015
Representation learning with contrastive predictive coding
A. van den Oord, Y. Li, and O. Vinyals · 2018
Later among the works it cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding
A. Wang, A. Singh, J. Michael, F. Hill, O. Levy, and S. R. Bowman · 2018
Later among the works it cites.
Deep visual domain adaptation: A survey
M. Wang and W. Deng · 2018
Later among the works it cites.
Unsupervised feature learning via non-parametric instance discrimination
Z. Wu, Y. Xiong, S. X. Yu, and D. Lin · 2018
Later among the works it cites.
Uniter: Learning universal image-text representations
Y.-C. Chen, L. Li, L. Yu, A. E. Kholy, F. Ahmed, Z. Gan, Y. Cheng, and J. Liu · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Deep residual learning for image recognition. corr abs/1512.03385 (2015), 2015
K. He, X. Zhang, S. Ren, and J. Sun · 2015
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
S. Ioffe and C. Szegedy · 2015
Cited alongside, same era.
Skip-thought vectors
R. Kiros, Y. Zhu, R. R. Salakhutdinov, R. Zemel, R. Urtasun, A. Torralba, and S. Fidler · 2015
Cited alongside, same era.
Deep learning
Y. LeCun, Y. Bengio, and G. Hinton · 2015
Cited alongside, same era.
Semi-supervised learning with ladder networks
A. Rasmus, M. Berglund, M. Honkala, H. Valpola, and T. Raiko · 2015
Cited alongside, same era.
Unsupervised learning of video representations using lstms
N. Srivastava, E. Mansimov, and R. Salakhudinov · 2015
Cited alongside, same era.
Lakhnes: Improving multi-instrumental music generation with cross-domain pre-training
C. Donahue, H. H. Mao, Y. E. Li, G. W. Cottrell, and J. McAuley · 2019
Later among the works it cites.
R. Eloff, A. Nortje, B. van Niekerk, A. Govender, L. Nortje, A. Pretorius, E. Van Biljon, E. van der Westhuizen, L. van Staden, and H. Kamper · 2019
Later among the works it cites.
Video representation learning by dense predictive coding
T. Han, W. Xie, and A. Zisserman · 2019
Later among the works it cites.
Data-efficient image recognition with contrastive predictive coding
O. J. Hénaff, A. Srinivas, J. De Fauw, A. Razavi, C. Doersch, S. Eslami, and A. v. d. Oord · 2019
Later among the works it cites.
Self-supervised visual feature learning with deep neural networks: A survey
L. Jing and Y. Tian · 2019
Later among the works it cites.
Big transfer (bit): General visual representation learning
A. Kolesnikov, L. Beyer, X. Zhai, J. Puigcerver, J. Yung, S. Gelly, and N. Houlsby · 2019
Later among the works it cites.
A mutual information maximization perspective of language representation learning
L. Kong, C. d. M. d’Autume, W. Ling, L. Yu, Z. Dai, and D. Yogatama · 2019
Later among the works it cites.
M. Lewis, Y. Liu, N. Goyal, M. Ghazvininejad, A. Mohamed, O. Levy, V. Stoyanov, and L. Zettlemoyer · 2019
Later among the works it cites.
Improving neural story generation by targeted common sense grounding
H. H. Mao, B. P. Majumder, J. McAuley, and G. W. Cottrell · 2019
Later among the works it cites.
Continual lifelong learning with neural networks: A review
G. I. Parisi, R. Kemker, J. L. Part, C. Kanan, and S. Wermter · 2019
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu · 2019
Later among the works it cites.
Studying the inductive biases of rnns with synthetic variations of natural languages
S. Ravfogel, Y. Goldberg, and T. Linzen · 2019
Later among the works it cites.
Neural transfer learning for natural language processing
S. Ruder · 2019
Later among the works it cites.
wav2vec: Unsupervised pre-training for speech recognition
S. Schneider, A. Baevski, R. Collobert, and M. Auli · 2019
Later among the works it cites.
Mass: Masked sequence to sequence pre-training for language generation
K. Song, X. Tan, T. Qin, J. Lu, and T.-Y. Liu · 2019
Later among the works it cites.
A deep learning approach with deep contextualized word representations for chemical-protein interaction extraction from biomedical literature
C. Sun, Z. Yang, L. Luo, L. Wang, Y. Zhang, H. Lin, and J. Wang · 2019
Later among the works it cites.
The bitter lesson
R. Sutton · 2019
Later among the works it cites.
Selfie: Self-supervised pretraining for image embedding
T. H. Trinh, M. Luong, and Q. V. Le · 2019
Later among the works it cites.
Generalizing from a few examples: A survey on few-shot learning
Y. Wang, Q. Yao, J. Kwok, and L. M. Ni · 2019
Later among the works it cites.
Xlnet: Generalized autoregressive pretraining for language understanding
Z. Yang, Z. Dai, Y. Yang, J. Carbonell, R. R. Salakhutdinov, and Q. V. Le · 2019
Later among the works it cites.
Rezero is all you need: Fast convergence at large depth
T. Bachlechner, B. P. Majumder, H. H. Mao, G. W. Cottrell, and J. McAuley · 2020
Closest in time.
A simple framework for contrastive learning of visual representations
T. Chen, S. Kornblith, M. Norouzi, and G. Hinton · 2020
Closest in time.
Audio albert: A lite bert for self-supervised learning of audio representation
P.-H. Chi, P.-H. Chung, T.-H. Wu, C.-C. Hsieh, S.-W. Li, and H.-y. Lee · 2020
Closest in time.
Electra: Pre-training text encoders as discriminators rather than generators
K. Clark, M.-T. Luong, Q. V. Le, and C. D. Manning · 2020
Closest in time.
Jukebox: A generative model for music
P. Dhariwal, H. Jun, C. Payne, J. W. Kim, A. Radford, and I. Sutskever · 2020
Closest in time.
Momentum contrast for unsupervised visual representation learning
K. He, H. Fan, Y. Wu, S. Xie, and R. Girshick · 2020
Closest in time.
A survey on generative adversarial networks: Variants, applications, and training
A. Jabbar, X. Li, and B. Omar · 2020
Closest in time.
Spanbert: Improving pre-training by representing and predicting spans
M. Joshi, D. Chen, Y. Liu, D. S. Weld, L. Zettlemoyer, and O. Levy · 2020
Closest in time.
Unsupervised translation of programming languages
M. Lachaux, B. Rozière, L. Chanussot, and G. Lample · 2020
Closest in time.
Z. Li, E. Wallace, S. Shen, K. Lin, K. Keutzer, D. Klein, and J. E. Gonzalez · 2020
Closest in time.
Automatic shortcut removal for self-supervised representation learning
M. Minderer, O. Bachem, N. Houlsby, and M. Tschannen · 2020
Closest in time.
Movement pruning: Adaptive sparsity by fine-tuning
V. Sanh, T. Wolf, and A. M. Rush · 2020
Closest in time.
A survey on semi-, self-and unsupervised techniques in image classification
L. Schmarje, M. Santarossa, S.-M. Schröder, and R. Koch · 2020
Closest in time.
On mutual information maximization for representation learning
M. Tschannen, J. Djolonga, P. K. Rubenstein, S. Gelly, and M. Lucic · 2020
Closest in time.