Fetching the paper…
Reading the bibliography…
Most uses of machine learning today involve training a model from scratch for a particular task, or sometimes starting with a model pretrained on a related task and then fine-tuning on a downstream task.
Catastrophic interference in connectionist networks: The sequential learning problem
M. McCloskey and N. J. Cohen · 1989
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
Catastrophic forgetting in connectionist networks
R. M. French · 1999
Earlier work this paper cites.
To transfer or not to transfer
M. T. Rosenstein · 2005
Earlier work this paper cites.
Gaussian process optimization in the bandit setting: No regret and experimental design
N. Srinivas, A. Krause, S. M. Kakade, and M. W. Seeger · 2010
Earlier work this paper cites.
Algorithms for hyper-parameter optimization
J. Bergstra, R. Bardenet, Y. Bengio, and B. Kégl · 2011
Earlier work this paper cites.
Evolutionary computation meets machine learning: A survey
J. Zhang, Z. hui Zhan, Y. Lin, N. Chen, Y. jiao Gong, J. Zhong, H. S. hung Chung, Y. Li, and Y. hui Shi · 2011
Earlier work this paper cites.
Practical bayesian optimization of machine learning algorithms
J. Snoek, H. Larochelle, and R. P. Adams · 2012
Earlier work this paper cites.
Human-level concept learning through probabilistic program induction
B. M. Lake, R. Salakhutdinov, and J. B. Tenenbaum · 2015
Earlier work this paper cites.
J. Ba, J. R. Kiros, and G. E. Hinton · 2016
Earlier work this paper cites.
Gaussian error linear units (gelus)
D. Hendrycks and K. Gimpel · 2016
Earlier work this paper cites.
A. A. Rusu, N. C. Rabinowitz, G. Desjardins, H. Soyer, J. Kirkpatrick, K. Kavukcuoglu, R. Pascanu, and R. Hadsell · 2016
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. van den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, S. Dieleman, D. Grewe, J. Nham, N. Kalchbrenner, I. Sutskever, T. P. Lillicrap, M. Leach, K. Kavukcuoglu, T. Graepel, and D. Hassabis · 2016
Earlier work this paper cites.
Visual domain decathlon
H. Bilen, S. Rebuffi, and T. Jakab · 2017
Earlier work this paper cites.
Emnist: Extending mnist to handwritten letters
G. Cohen, S. Afshar, J. C. Tapson, and A. van Schaik · 2017
Earlier work this paper cites.
Pathnet: Evolution channels gradient descent in super neural networks
C. Fernando, D. S. Banarse, C. Blundell, Y. Zwols, D. R. Ha, A. A. Rusu, A. Pritzel, and D. Wierstra · 2017
Earlier work this paper cites.
Population based training of neural networks
M. Jaderberg, V. Dalibard, S. Osindero, W. M. Czarnecki, J. Donahue, A. Razavi, O. Vinyals, T. Green, I. Dunning, K. Simonyan, C. Fernando, and K. Kavukcuoglu · 2017
Cited alongside, same era.
In-datacenter performance analysis of a tensor processing unit
N. P. Jouppi, C. Young, N. Patil, D. A. Patterson, G. Agrawal, R. S. Bajwa, S. Bates, S. Bhatia, N. J. Boden, A. Borchers, R. Boyle, P. luc Cantin, C. Chao, C. Clark, J. Coriell, M. Daley, M. Dau, J. Dean, B. Gelb, T. V. Ghaemmaghami, R. Gottipati, W. Gulland, R. B. Hagmann, C. R. Ho, D. Hogberg, J. Hu, R. Hundt, D. Hurt, J. Ibarz, A. Jaffey, A. Jaworski, A. Kaplan, H. Khaitan, D. Killebrew, A. Koch, N. Kumar, S. Lacy, J. Laudon, J. Law, D. Le, C. Leary, Z. Liu, K. A. Lucke, A. Lundin, G. MacKean, A. Maggiore, M. Mahony, K. Miller, R. Nagarajan, R. Narayanaswami, R. Ni, K. Nix, T. Norrie, M. Omernick, N. Penukonda, A. Phelps, J. Ross, M. Ross, A. Salek, E. Samadiani, C. Severn, G. Sizikov, M. Snelham, J. Souter, D. Steinberg, A. Swing, M. Tan, G. Thorson, B. Tian, H. Toma, E. Tuttle, V. Vasudevan, R. Walter, W. Wang, E. Wilcox, and D. H. Yoon · 2017
Cited alongside, same era.
Learning multiple visual domains with residual adapters
S.-A. Rebuffi, H. Bilen, and A. Vedaldi · 2017
Cited alongside, same era.
A. Vaswani, N. M. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Regularized evolution for image classifier architecture search
E. Real, A. Aggarwal, Y. Huang, and Q. V. Le · 2019
Later among the works it cites.
Mnasnet: Platform-aware neural architecture search for mobile
M. Tan, B. Chen, R. Pang, V. Vasudevan, and Q. V. Le · 2019
Later among the works it cites.
Grandmaster level in starcraft ii using multi-agent reinforcement learning
O. Vinyals, I. Babuschkin, W. M. Czarnecki, M. Mathieu, A. Dudzik, J. Chung, D. H. Choi, R. Powell, T. Ewalds, P. Georgiev, J. Oh, D. Horgan, M. Kroiss, I. Danihelka, A. Huang, L. Sifre, T. Cai, J. P. Agapiou, M. Jaderberg, A. S. Vezhnevets, R. Leblond, T. Pohlen, V. Dalibard, D. Budden, Y. Sulsky, J. Molloy, T. L. Paine, C. Gulcehre, Z. Wang, T. Pfaff, Y. Wu, R. Ring, D. Yogatama, D. Wünsch, K. McKinney, O. Smith, T. Schaul, T. P. Lillicrap, K. Kavukcuoglu, D. Hassabis, C. Apps, and D. Silver · 2019
Later among the works it cites.
Characterizing and avoiding negative transfer
Z. Wang, Z. Dai, B. Póczos, and J. G. Carbonell · 2019
Later among the works it cites.
Language models are few-shot learners
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. J. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Neural architecture search with reinforcement learning
B. Zoph and Q. V. Le · 2017
Cited alongside, same era.
Gradnorm: Gradient normalization for adaptive loss balancing in deep multitask networks
Z. Chen, V. Badrinarayanan, C.-Y. Lee, and A. Rabinovich · 2018
Cited alongside, same era.
Deep learning for classical japanese literature
T. Clanuwat, M. Bober-Irizar, A. Kitamoto, A. Lamb, K. Yamamoto, and D. Ha · 2018
Cited alongside, same era.
Multi-task learning using uncertainty to weigh losses for scene geometry and semantics
A. Kendall, Y. Gal, and R. Cipolla · 2018
Cited alongside, same era.
Evolutionary-neural hybrid agents for architecture search
K. Maziarz, A. Khorlin, Q. de Laroussilhe, and A. Gesmundo · 2018
Cited alongside, same era.
Efficient neural architecture search via parameter sharing
H. Pham, M. Y. Guan, B. Zoph, Q. V. Le, and J. Dean · 2018
Cited alongside, same era.
Multi-task learning as multi-objective optimization
O. Sener and V. Koltun · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2019
Cited alongside, same era.
Later among the works it cites.
Chip placement with deep reinforcement learning
A. Mirhoseini, A. Goldie, M. Yazgan, J. W. Jiang, E. M. Songhori, S. Wang, Y.-J. Lee, E. Johnson, O. Pathak, S. Bae, A. Nazi, J. Pak, A. Tong, K. Srinivasa, W. Hang, E. Tuncer, A. Babu, Q. V. Le, J. Laudon, R. Ho, R. Carpenter, and J. Dean · 2020
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
C. Raffel, N. M. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu · 2020
Later among the works it cites.
Incremental learning through deep adaptation
A. Rosenfeld and J. K. Tsotsos · 2020
Later among the works it cites.
Improved protein structure prediction using potentials from deep learning
A. W. Senior, R. Evans, J. M. Jumper, J. Kirkpatrick, L. Sifre, T. Green, C. Qin, A. Zídek, A. W. R. Nelson, A. Bridgland, H. Penedones, S. Petersen, K. Simonyan, S. Crossan, P. Kohli, D. T. Jones, D. Silver, K. Kavukcuoglu, and D. Hassabis · 2020
Later among the works it cites.
Ernie 2.0: A continual pre-training framework for language understanding
Y. Sun, S. Wang, Y. Li, S. Feng, H. Tian, H. Wu, and H. Wang · 2020
Later among the works it cites.
Gradient surgery for multi-task learning
T. Yu, S. Kumar, A. Gupta, S. Levine, K. Hausman, and C. Finn · 2020
Later among the works it cites.
A survey on negative transfer
W. Zhang, L. Deng, L. Zhang, and D. Wu · 2020
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby · 2021
Later among the works it cites.
Factors of influence for transfer learning across diverse appearance domains and task types
T. Mensink, J. R. R. Uijlings, A. Kuznetsova, M. Gygli, and V. Ferrari · 2021
Later among the works it cites.
How to train your vit? data, augmentation, and regularization in vision transformers
A. Steiner, A. Kolesnikov, X. Zhai, R. Wightman, J. Uszkoreit, and L. Beyer · 2021
Later among the works it cites.