Fetching the paper…
Reading the bibliography…
Few-shot algorithms aim at learning new tasks provided only a handful of training examples.
Using fast weights to deblur old memories
Hinton, G. E. and Plaut, D. C · 1987
Earlier work this paper cites.
Evolutionary principles in self-referential learning, or on learning how to learn: the meta-meta-… hook
Schmidhuber, J · 1987
Earlier work this paper cites.
Learning to control fast-weight memories: An alternative to dynamic recurrent networks
Schmidhuber, J · 1992
Earlier work this paper cites.
Learning to learn using gradient descent
Hochreiter, S., Younger, A. S., and Conwell, P. R · 2001
Earlier work this paper cites.
Object classification from a single example utilizing class relevance metrics
Fink, M · 2005
Earlier work this paper cites.
One-shot learning of object categories
Fei-Fei, L., Fergus, R., and Perona, P · 2006
Earlier work this paper cites.
Learning task grouping and overlap in multi-task learning
Kumar, A. and Daume III, H · 2012
Earlier work this paper cites.
Sparse coding for multitask and transfer learning
Maurer, A., Pontil, M., and Romera-Paredes, B · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Learning to learn by gradient descent by gradient descent
Andrychowicz, M., Denil, M., Gomez, S., Hoffman, M. W., Pfau, D., Schaul, T., Shillingford, B., and De Freitas, N · 2016
Earlier work this paper cites.
Using fast weights to attend to the recent past
Ba, J., Hinton, G. E., Mnih, V., Leibo, J. Z., and Ionescu, C · 2016
Earlier work this paper cites.
Matching networks for one shot learning
Vinyals, O., Blundell, C., Lillicrap, T., Wierstra, D., et al · 2016
Earlier work this paper cites.
Model-agnostic meta-learning for fast adaptation of deep networks
Finn, C., Abbeel, P., and Levine, S · 2017
Cited alongside, same era.
Exploring the structure of a real-time, arbitrary neural artistic stylization network
Ghiasi, G., Lee, H., Kudlur, M., Dumoulin, V., and Shlens, J · 2017
Cited alongside, same era.
Prototypical networks for few-shot learning
Snell, J., Swersky, K., and Zemel, R · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2018
Cited alongside, same era.
Ctrl: A conditional transformer language model for controllable generation
Keskar, N. S., McCann, B., Varshney, L. R., Xiong, C., and Socher, R · 2019
Later among the works it cites.
Compositional generalization through meta sequence-to-sequence learning
Lake, B. M · 2019
Later among the works it cites.
Multiple-attribute text rewriting
Lample, G., Subramanian, S., Smith, E., Denoyer, L., Ranzato, M., and Boureau, Y.-L · 2019
Later among the works it cites.
Stabilizing transformers for reinforcement learning
Parisotto, E., Song, F., Rae, J., Pascanu, R., Gulcehre, C., Jayakumar, S., Jaderberg, M., Kaufman, R. L., Clark, A., Noury, S., Botvinick, M., Heess, N., and Hadsell, R · 2019
Later among the works it cites.
Task-driven modular networks for zero-shot compositional learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mishra, N., Rohaninejad, M., Chen, X., and Abbeel, P · 2018
Cited alongside, same era.
Reptile: a scalable metalearning algorithm
Nichol, A. and Schulman, J · 2018
Cited alongside, same era.
Film: Visual reasoning with a general conditioning layer
Perez, E., Strub, F., De Vries, H., Dumoulin, V., and Courville, A · 2018
Cited alongside, same era.
Few-shot text classification with distributional signatures
Bao, Y., Wu, M., Chang, S., and Barzilay, R · 2019
Cited alongside, same era.
Findings of the 2019 conference on machine translation (WMT19)
Barrault, L., Bojar, O., Costa-jussà, M. R., Federmann, C., Fishel, M., Graham, Y., Haddow, B., Huck, M., Koehn, P., Malmasi, S., Monz, C., Müller, M., Pal, S., Post, M., and Zampieri, M · 2019
Cited alongside, same era.
Parameter-efficient transfer learning for nlp
Houlsby, N., Giurgiu, A., Jastrzebski, S., Morrone, B., De Laroussilhe, Q., Gesmundo, A., Attariyan, M., and Gelly, S · 2019
Cited alongside, same era.
Purushwalkam, S., Nickel, M., Gupta, A., and Ranzato, M · 2019
Later among the works it cites.
Meta-learning with latent embedding optimization
Rusu, A. A., Rao, D., Sygnowski, J., Vinyals, O., Pascanu, R., Osindero, S., and Hadsell, R · 2019
Later among the works it cites.
Mixture models for diverse machine translation: Tricks of the trade
Shen, T., Ott, M., Auli, M., and Ranzato, M · 2019
Later among the works it cites.
Bert and pals: Projected attention layers for efficient adaptation in multi-task learning
Stickland, A. C. and Murray, I · 2019
Later among the works it cites.
Defending against neural fake news
Zellers, R., Holtzman, A., Rashkin, H., Bisk, Y., Farhadi, A., Roesner, F., and Choi, Y · 2019
Later among the works it cites.
Fast context adaptation via meta-learning
Zintgraf, L. M., Shiarlis, K., Kurin, V., Hofmann, K., and Whiteson, S · 2019
Later among the works it cites.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Closest in time.