Fetching the paper…
Reading the bibliography…
Self-supervised learning excels in learning representations from large amounts of unlabeled data, demonstrating success across multiple data modalities.
Smiles, a chemical language and information system. 1. introduction to methodology and encoding rules
D. Weininger · 1988
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
The million song dataset challenge
B. McFee, T. Bertin-Mahieux, D. P. Ellis, and G. R. Lanckriet · 2012
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Earlier work this paper cites.
Deep learning
Y. LeCun, Y. Bengio, and G. Hinton · 2015
Earlier work this paper cites.
Xgboost: A scalable tree boosting system
T. Chen and C. Guestrin · 2016
Earlier work this paper cites.
Deep networks with stochastic depth
G. Huang, Y. Sun, Z. Liu, D. Sedra, and K. Q. Weinberger · 2016
Earlier work this paper cites.
Self-normalizing neural networks
G. Klambauer, T. Unterthiner, A. Mayr, and S. Hochreiter · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
I. Loshchilov and F. Hutter · 2017
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
N. Shazeer, A. Mirhoseini, K. Maziarz, A. Davis, Q. Le, G. Hinton, and J. Dean · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
mixup: Beyond empirical risk minimization
H. Zhang, M. Cisse, Y. N. Dauphin, and D. Lopez-Paz · 2017
Earlier work this paper cites.
Autoaugment: Learning augmentation policies from data
E. D. Cubuk, B. Zoph, D. Mane, V. Vasudevan, and Q. V. Le · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Earlier work this paper cites.
R. Geirhos, P. Rubisch, C. Michaelis, M. Bethge, F. A. Wichmann, and W. Brendel · 2018
Earlier work this paper cites.
Representation learning with contrastive predictive coding
A. v. d. Oord, Y. Li, and O. Vinyals · 2018
Earlier work this paper cites.
Catboost: unbiased boosting with categorical features
L. Prokhorenkova, G. Gusev, A. Vorobev, A. V. Dorogush, and A. Gulin · 2018
Earlier work this paper cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding
A. Wang, A. Singh, J. Michael, F. Hill, O. Levy, and S. R. Bowman · 2018
Earlier work this paper cites.
Moleculenet: a benchmark for molecular machine learning
Z. Wu, B. Ramsundar, E. N. Feinberg, J. Gomes, C. Geniesse, A. S. Pappu, K. Leswing, and V. Pande · 2018
Earlier work this paper cites.
Guacamol: benchmarking models for de novo molecular design
N. Brown, M. Fiscato, M. H. Segler, and A. C. Vaucher · 2019
Earlier work this paper cites.
Set transformer: A framework for attention-based permutation-invariant neural networks
J. Lee, Y. Lee, J. Kim, A. Kosiorek, S. Choi, and Y. W. Teh · 2019
Earlier work this paper cites.
M. Lewis, Y. Liu, N. Goyal, M. Ghazvininejad, A. Mohamed, O. Levy, V. Stoyanov, and L. Zettlemoyer · 2019
Earlier work this paper cites.
Fast autoaugment
S. Lim, I. Kim, T. Kim, C. Kim, and S. Kim · 2019
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, and V. Stoyanov · 2019
Cited alongside, same era.
Neural oblivious decision ensembles for deep learning on tabular data
S. Popov, S. Morozov, and A. Babenko · 2019
Cited alongside, same era.
Deepchem: democratizing deep-learning for drug discovery, quantum chemistry
B. Ramsundar, P. Eastman, E. Feinberg, J. Gomes, K. Leswing, A. Pappu, M. Wu, and V. Pande · 2019
Cited alongside, same era.
Evaluating protein transfer learning with tape
R. Rao, N. Bhattacharya, N. Thomas, Y. Duan, P. Chen, J. Canny, P. Abbeel, and Y. Song · 2019
Cited alongside, same era.
wav2vec: Unsupervised pre-training for speech recognition
S. Schneider, A. Baevski, R. Collobert, and M. Auli · 2019
Dabs: A domain-agnostic benchmark for self-supervised learning
A. Tamkin, V. Liu, R. Lu, D. Fein, C. Schultz, and N. Goodman · 2021
Later among the works it cites.
Dcn v2: Improved deep & cross network and practical lessons for web-scale learning to rank systems
R. Wang, R. Shivanna, D. Cheng, S. Jain, D. Lin, L. Hong, and E. Chi · 2021
Later among the works it cites.
Chemberta-2: Towards chemical foundation models
W. Ahmad, E. Simon, S. Chithrananda, G. Grand, and B. Ramsundar · 2022
Later among the works it cites.
Data2vec: A general framework for self-supervised learning in speech, vision and language
A. Baevski, W.-N. Hsu, Q. Xu, A. Babu, J. Gu, and M. Auli · 2022
Later among the works it cites.
Proteinbert: a universal deep-learning model of protein sequence and function
N. Brandes, D. Ofer, Y. Peleg, N. Rappoport, and M. Linial · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Autoint: Automatic feature interaction learning via self-attentive neural networks
W. Song, C. Shi, Z. Xiao, Z. Duan, Y. Xu, M. Zhang, and J. Tang · 2019
Cited alongside, same era.
The bitter lesson
R. Sutton · 2019
Cited alongside, same era.
Cutmix: Regularization strategy to train strong classifiers with localizable features
S. Yun, D. Han, S. J. Oh, S. Chun, J. Choe, and Y. Yoo · 2019
Cited alongside, same era.
Gradient boosting neural networks: Grownet
S. Badirli, X. Liu, Z. Xing, A. Bhowmik, K. Doan, and S. S. Keerthi · 2020
Cited alongside, same era.
wav2vec 2.0: A framework for self-supervised learning of speech representations
A. Baevski, Y. Zhou, A. Mohamed, and M. Auli · 2020
Cited alongside, same era.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Cited alongside, same era.
A simple framework for contrastive learning of visual representations
T. Chen, S. Kornblith, M. Norouzi, and G. Hinton · 2020
Cited alongside, same era.
Later among the works it cites.
Palm: Scaling language modeling with pathways
A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann, et al · 2022
Later among the works it cites.
Masked autoencoders as spatiotemporal learners
C. Feichtenhofer, H. Fan, Y. Li, and K. He · 2022
Later among the works it cites.
On embeddings for numerical features in tabular deep learning
Y. Gorishniy, I. Rubachev, and A. Babenko · 2022
Later among the works it cites.
Masked autoencoders are scalable vision learners
K. He, X. Chen, S. Xie, Y. Li, P. Dollár, and R. Girshick · 2022
Later among the works it cites.
Masked autoencoders for point cloud self-supervised learning
Y. Pang, W. Wang, F. E. Tay, W. Liu, Y. Tian, and L. Yuan · 2022
Later among the works it cites.
Adversarial masking for self-supervised learning
Y. Shi, N. Siddharth, P. Torr, and A. R. Kosiorek · 2022
Later among the works it cites.
Teachaugment: Data augmentation optimization using teacher knowledge
T. Suzuki · 2022
Later among the works it cites.
Deit iii: Revenge of the vit
H. Touvron, M. Cord, and H. Jégou · 2022
Later among the works it cites.
Image as a foreign language: Beit pretraining for all vision and vision-language tasks
W. Wang, H. Bao, L. Dong, J. Bjorck, Z. Peng, Q. Liu, K. Aggarwal, O. K. Mohammed, S. Singhal, S. Som, et al · 2022
Later among the works it cites.
Simmim: A simple framework for masked image modeling
Z. Xie, Z. Zhang, Y. Cao, Y. Lin, J. Bao, Z. Yao, Q. Dai, and H. Hu · 2022
Later among the works it cites.
Masked autoencoders that listen
H. Xu, J. Li, A. Baevski, M. Auli, W. Galuba, F. Metze, C. Feichtenhofer, et al · 2022
Later among the works it cites.
Byt5: Towards a token-free future with pre-trained byte-to-byte models
L. Xue, A. Barua, N. Constant, R. Al-Rfou, S. Narang, M. Kale, A. Roberts, and C. Raffel · 2022
Later among the works it cites.
Point-bert: Pre-training 3d point cloud transformers with masked point modeling
X. Yu, L. Tang, Y. Rao, T. Huang, J. Zhou, and J. Lu · 2022
Later among the works it cites.
Adamae: Adaptive masking for efficient spatiotemporal learning with masked autoencoders
W. G. C. Bandara, N. Patel, A. Gholami, M. Nikkhah, M. Agrawal, and V. M. Patel · 2023
Later among the works it cites.
Dinov2: Learning robust visual features without supervision
M. Oquab, T. Darcet, T. Moutakanni, H. Vo, M. Szafraniec, V. Khalidov, P. Fernandez, D. Haziza, F. Massa, A. El-Nouby, et al · 2023
Later among the works it cites.
Megabyte: Predicting million-byte sequences with multiscale transformers
L. Yu, D. Simig, C. Flaherty, A. Aghajanyan, L. Zettlemoyer, and M. Lewis · 2023
Later among the works it cites.
Uni-mol: A universal 3d molecular representation learning framework
G. Zhou, Z. Gao, Q. Ding, H. Zheng, H. Xu, Z. Wei, L. Zhang, and G. Ke · 2023
Later among the works it cites.