Fetching the paper…
Reading the bibliography…
Pretraining a neural network on a large dataset is becoming a cornerstone in machine learning that is within the reach of only a few communities with large-resources.
The hungarian method for the assignment problem
Kuhn, H. W · 1955
Earlier work this paper cites.
Object detection combining recognition and segmentation
Wang, L., Shi, J., Song, G., and Shen, I.-f · 2007
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A. et al · 2009
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Saxe, A. M., McClelland, J. L., and Ganguli, S · 2013
Earlier work this paper cites.
Net2net: Accelerating learning via knowledge transfer
Chen, T., Goodfellow, I., and Shlens, J · 2015
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
He, K., Zhang, X., Ren, S., and Sun, J · 2015
Earlier work this paper cites.
Gated graph sequence neural networks
Li, Y., Tarlow, D., Brockschmidt, M., and Zemel, R · 2015
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., et al · 2015
Earlier work this paper cites.
Ha, D., Dai, A., and Le, Q. V · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Smash: one-shot model architecture search through hypernetworks
Brock, A., Lim, T., Ritchie, J. M., and Weston, N · 2017
Earlier work this paper cites.
Model-agnostic meta-learning for fast adaptation of deep networks
Finn, C., Abbeel, P., and Levine, S · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2017
Earlier work this paper cites.
Group sparse regularization for deep neural networks
Scardapane, S., Comminiello, D., Hussain, A., and Uncini, A · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Darts: Differentiable architecture search
Liu, H., Simonyan, K., and Yang, Y · 2018
Earlier work this paper cites.
Graph hypernetworks for neural architecture search
Zhang, C., Ren, M., and Urtasun, R · 2018
Earlier work this paper cites.
Metainit: Initializing learning by learning to initialize
Dauphin, Y. and Schoenholz, S. S · 2019
Cited alongside, same era.
Do better imagenet models transfer better?
Kornblith, S., Shlens, J., and Le, Q. V · 2019
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., et al · 2019
Cited alongside, same era.
Fast and flexible multi-task classification using conditional neural adaptive processes
Requeima, J., Gordon, J., Bronskill, J., Nowozin, S., and Turner, R. E · 2019
Cited alongside, same era.
Energy and policy considerations for deep learning in nlp
Strubell, E., Ganesh, A., and McCallum, A · 2019
Cited alongside, same era.
Resnet strikes back: An improved training procedure in timm
Wightman, R., Touvron, H., and Jégou, H · 2021
Later among the works it cites.
Do transformers really perform badly for graph representation?
Ying, C., Cai, T., Luo, S., Zheng, S., Ke, G., He, D., Shen, Y., and Liu, T.-Y · 2021
Later among the works it cites.
Gradinit: Learning to initialize neural networks for stable and efficient training
Zhu, C., Ni, R., Xu, Z., Kong, K., Huang, W. R., and Goldstein, T · 2021
Later among the works it cites.
Nern–learning neural representations for neural networks
Ashkenazi, M., Rimon, Z., Vainshtein, R., Levi, S., Richardson, E., Mintz, P., and Treister, E · 2022
Later among the works it cites.
Structure-aware transformer for graph representation learning
Chen, D., O’Bray, L., and Borgwardt, K · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zhai, X., Puigcerver, J., Kolesnikov, A., Ruyssen, P., Riquelme, C., Lucic, M., Djolonga, J., Pinto, A. S., Neumann, M., Dosovitskiy, A., et al · 2019
Cited alongside, same era.
Fixup initialization: Residual learning without normalization
Zhang, H., Dauphin, Y. N., and Ma, T · 2019
Cited alongside, same era.
Catch: Context-based meta reinforcement learning for transferrable architecture search
Chen, X., Duan, Y., Chen, Z., Xu, H., Chen, Z., Liang, X., Zhang, T., and Li, Z · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al · 2020
Cited alongside, same era.
A generalization of transformer networks to graphs
Dwivedi, V. P. and Bresson, X · 2020
Cited alongside, same era.
Meta-learning of neural architectures for few-shot learning
Elsken, T., Staffler, B., Metzen, J. H., and Hutter, F · 2020
Cited alongside, same era.
Towards fast adaptation of neural architectures with meta learning
Lian, D., Zheng, Y., Xu, Y., Lu, Y., Lin, L., Zhao, P., Huang, J., and Gao, S · 2020
Cited alongside, same era.
Later among the works it cites.
Breaking the architecture barrier: A method for efficient knowledge transfer across networks
Czyzewski, M. A., Nowak, D., and Piechowiak, K · 2022
Later among the works it cites.
Gradmax: Growing neural networks using gradient information
Evci, U., Vladymyrov, M., Unterthiner, T., van Merriënboer, B., and Pedregosa, F · 2022
Later among the works it cites.
Pure transformers are powerful graph learners
Kim, J., Nguyen, T. D., Min, S., Cho, S., Lee, M., Lee, H., and Hong, S · 2022
Later among the works it cites.
Pretraining a neural network before knowing its architecture
Knyazev, B · 2022
Later among the works it cites.
A convnet for the 2020s
Liu, Z., Mao, H., Wu, C.-Y., Feichtenhofer, C., Darrell, T., and Xie, S · 2022
Later among the works it cites.
Learning to learn with generative models of neural network checkpoints
Peebles, W., Radosavovic, I., Brooks, T., Efros, A. A., and Malik, J · 2022
Later among the works it cites.
Hyper-representations as generative models: Sampling unseen neural network weights
Schürholt, K., Knyazev, B., Giró-i Nieto, X., and Borth, D · 2022
Later among the works it cites.
One hyper-initializer for all network architectures in medical image analysis
Shang, F., Yang, Y., Yang, D., Wu, J., Wang, X., and Xu, Y · 2022
Later among the works it cites.
Towards theoretically inspired neural initialization optimization
Yang, Y., Wang, H., Yuan, H., and Lin, Z · 2022
Later among the works it cites.
Scaling vision transformers
Zhai, X., Kolesnikov, A., Houlsby, N., and Beyer, L · 2022
Later among the works it cites.
Hypertransformer: Model generation for supervised and semi-supervised few-shot learning
Zhmoginov, A., Sandler, M., and Vladymyrov, M · 2022
Later among the works it cites.