Fetching the paper…
Reading the bibliography…
We study the problem of meta-learning through the lens of online convex optimization, developing a meta-algorithm bridging the gap between popular gradient-based meta-learning and classical regularization-based multi-task transfer methods.
An algorithm for quadratic programming
Frank, M. and Wolfe, P · 1956
Earlier work this paper cites.
Gradient methods for minimizing functionals
Polyak, B. T · 1963
Earlier work this paper cites.
Weighted sums of certain dependent random variables
Azuma, K · 1967
Earlier work this paper cites.
The relaxation method of finding the common point of convex sets and its application to the solution of problems in convex programming
Bregman, L. M · 1967
Earlier work this paper cites.
On tail probabilities for martingales
Freedman, D. A · 1975
Earlier work this paper cites.
WordNet: An Electronic Lexical Database
Fellbaum, C · 1998
Earlier work this paper cites.
Learning to Learn
Thrun, S. and Pratt, L · 1998
Earlier work this paper cites.
A model of inductive bias learning
Baxter, J · 2000
Earlier work this paper cites.
Online convex programming and generalized infinitesimal gradient ascent
Zinkevich, M · 2003
Earlier work this paper cites.
On the generalization ability of on-line learning algorithms
Cesa-Bianchi, N., Conconi, A., and Gentile, C · 2004
Earlier work this paper cites.
Regularized multi-task learning
Evgeniou, T. and Pontil, M · 2004
Earlier work this paper cites.
Clustering with Bregman divergences
Banerjee, A., Merugu, S., Dhillon, I. S., and Ghosh, J · 2005
Earlier work this paper cites.
Algorithmic stability and meta-learning
Maurer, A · 2005
Earlier work this paper cites.
Multitask learning with expert advice
Abernethy, J., Bartlett, P., and Rakhlin, A · 2007
Earlier work this paper cites.
Online learning of multiple tasks with a shared loss
Dekel, O., Long, P. M., and Singer, Y · 2007
Earlier work this paper cites.
Optimal strategies and minimax lower bounds for online convex games
Abernethy, J., Bartlett, P. L., Rakhlin, A., and Tewari, A · 2008
Earlier work this paper cites.
Adaptive online gradient descent
Bartlett, P. L., Hazan, E., and Rakhlin, A · 2008
Cited alongside, same era.
Mind the duality gap: Logarithmic regret algorithms for online optimization
Kakade, S. and Shalev-Shwartz, S · 2008
Cited alongside, same era.
Linear algorithms for online multitask classification
Cavallanti, G., Cesa-Bianchi, N., and Gentile, C · 2010
Cited alongside, same era.
Online learning and online convex optimization
Shalev-Shwartz, S · 2011
Cited alongside, same era.
Baselines and bigrams: Simple, good sentiment and topic classification
Wang, S. and Manning, C. D · 2012
Cited alongside, same era.
Stability and hypothesis transfer learning
Kuzborskij, I. and Orabona, F · 2013
Cited alongside, same era.
Optimization as a model for few-shot learning
Ravi, S. and Larochelle, H · 2017
Later among the works it cites.
Prototypical networks for few-shot learning
Snell, J., Swersky, K., and Zemel, R. S · 2017
Later among the works it cites.
Continuous adaptation via meta-learning in nonstationary and competitive environments
Al-Shedivat, M., Bansal, T., Burda, Y., Sutskever, I., Mordatch, I., and Abbeel, P · 2018
Later among the works it cites.
Meta-learning by adjusting priors based on extended PAC-Bayes theory
Amit, R. and Meir, R · 2018
Later among the works it cites.
A compressed sensing view of unsupervised text embeddings, bag-of-n-grams, and LSTMs
Arora, S., Khodak, M., Saunshi, N., and Vodrahalli, K · 2018
Later among the works it cites.
Federated meta-learning for recommendation
Chen, F., Dong, Z., Li, Z., and He, X · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
ELLA: An efficient lifelong learning algorithm
Ruvolo, P. and Eaton, E · 2013
Cited alongside, same era.
Glove: Global vectors for word representation
Pennington, J., Socher, R., and Manning, C. D · 2014
Cited alongside, same era.
A PAC-Bayesian bound for lifelong learning
Pentina, A. and Lampert, C. H · 2014
Cited alongside, same era.
Efficient representations for lifelong learning and autoencoding
Balcan, M.-F., Blum, A., and Vempala, S · 2015
Cited alongside, same era.
Introduction to online convex optimization
Hazan, E · 2015
Cited alongside, same era.
Linear convergence of gradient and proximal-gradient methods under the Polyak-Łojasiewicz condition
Karimi, H., Nutini, J., and Schmidt, M · 2016
Cited alongside, same era.
Meta-learning and universality: Deep representations and gradient descent can approximate any learning algorithm
Finn, C. and Levine, S · 2018
Later among the works it cites.
Bilevel programming for hyperparameter optimization and meta-learning
Franceschi, L., Frasconi, P., Salzo, S., Grazzi, R., and Pontil, M · 2018
Later among the works it cites.
Recasting gradient-baed meta-learning as hierarchical Bayes
Grant, E., Finn, C., Levine, S., Darrell, T., and Griffiths, T · 2018
Later among the works it cites.
Online gradient-based mixtures for transfer modulation in meta-learning
Jerfel, G., Grant, E., Griffiths, T. L., and Heller, K · 2018
Later among the works it cites.
Auto-Meta: Automated gradient based meta learner search
Kim, J., Lee, S., Kim, S., Cha, M., Lee, J. K., Choi, Y., Choi, Y., Choi, D.-Y., and Kim, J · 2018
Later among the works it cites.
On first-order meta-learning algorithms
Nichol, A., Achiam, J., and Schulman, J · 2018
Later among the works it cites.
A theoretical analysis of contrastive unsupervised representation learning
Arora, S., Khandeparkar, H., Khodak, M., Plevrakis, O., and Saunshi, N · 2019
Closest in time.
Learning-to-learn stochastic gradient descent with biased regularization
Denevi, G., Ciliberto, C., Grazzi, R., and Pontil, M · 2019
Closest in time.
Online meta-learning
Finn, C., Rajeswaran, A., Kakade, S., and Levine, S · 2019
Closest in time.
Fast rates for online gradient descent without strong convexity via Hoffman’s bound
Garber, D · 2019
Closest in time.