Fetching the paper…
Reading the bibliography…
This paper introduces Meta-Q-Learning (MQL), a new off-policy algorithm for meta-Reinforcement Learning (meta-RL).
The need for biases in learning generalizations
Tom M Mitchell · 1980
Earlier work this paper cites.
Shift of bias for inductive concept learning
Paul E Utgoff · 1986
Earlier work this paper cites.
Evolutionary principles in self-referential learning
Jurgen Schmidhuber · 1987
Earlier work this paper cites.
On the optimization of a synaptic learning rule
Samy Bengio, Yoshua Bengio, Jocelyn Cloutier, and Jan Gecsei · 1992
Earlier work this paper cites.
A note on importance sampling using standardized weights
Augustine Kong · 1992
Earlier work this paper cites.
Learning internal representations
Jonathan Baxter · 1995
Earlier work this paper cites.
Is learning the n-th thing any easier than learning the first?
Sebastian Thrun · 1996
Earlier work this paper cites.
On the optimization of a synaptic learning rule, 1997
Samy Bengio, Yoshua Bengio, Jocelyn Cloutier, and Jan Gecsei · 1997
Earlier work this paper cites.
Shifting inductive bias with success-story algorithm, adaptive levin search, and incremental self-improvement
Jürgen Schmidhuber, Jieyu Zhao, and Marco Wiering · 1997
Earlier work this paper cites.
A model of inductive bias learning
Jonathan Baxter · 2000
Earlier work this paper cites.
Learning to learn using gradient descent
Sepp Hochreiter, A Steven Younger, and Peter R Conwell · 2001
Earlier work this paper cites.
Doubly robust estimation in missing data and causal inference models
Heejung Bang and James M Robins · 2005
Earlier work this paper cites.
Demystifying double robustness: A comparison of alternative strategies for estimating a population mean from incomplete data
Joseph DY Kang, Joseph L Schafer, et al · 2007
Earlier work this paper cites.
Dataset shift in machine learning
Joaquin Quionero-Candela, Masashi Sugiyama, Anton Schwaighofer, and Neil D Lawrence · 2009
Earlier work this paper cites.
Linear-time estimators for propensity scores
Deepak Agarwal, Lihong Li, and Alexander Smola · 2011
Earlier work this paper cites.
Doubly robust policy evaluation and learning
Miroslav Dudík, John Langford, and Lihong Li · 2011
Earlier work this paper cites.
Learning to learn
Sebastian Thrun and Lorien Pratt · 2012
Cited alongside, same era.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Cited alongside, same era.
A probability path
Sidney I Resnick · 2013
Cited alongside, same era.
Monte Carlo statistical methods
Christian Robert and George Casella · 2013
Cited alongside, same era.
Sequential Monte Carlo methods in practice
Adrian Smith · 2013
Cited alongside, same era.
Learning phrase representations using rnn encoder-decoder for statistical machine translation
Kyunghyun Cho, Bart Van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio · 2014
Cited alongside, same era.
Optimization as a model for few-shot learning
Sachin Ravi and Hugo Larochelle · 2016
Later among the works it cites.
Learning to reinforcement learn
Jane X. Wang, Zeb Kurth-Nelson, Dhruva Tirumala, Hubert Soyer, Joel Z. Leibo, Rémi Munos, Charles Blundell, Dharshan Kumaran, and Matthew Botvinick · 2016
Later among the works it cites.
Model-agnostic meta-learning for fast adaptation of deep networks
Chelsea Finn, Pieter Abbeel, and Sergey Levine · 2017
Later among the works it cites.
Prototypical networks for few-shot learning
Jake Snell, Kevin Swersky, and Richard Zemel · 2017
Later among the works it cites.
A closer look at few-shot classification
Wei-Yu Chen, Yen-Cheng Liu, Zsolt Kira, Yu-Chiang Frank Wang, and Jia-Bin Huang · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Diederik P Kingma and Jimmy Ba · 2014
Cited alongside, same era.
Deterministic policy gradient algorithms
David Silver, Guy Lever, Nicolas Heess, Thomas Degris, Daan Wierstra, and Martin Riedmiller · 2014
Cited alongside, same era.
Deep recurrent q-learning for partially observable mdps
Matthew Hausknecht and Peter Stone · 2015
Cited alongside, same era.
Memory-based control with recurrent neural networks
Nicolas Heess, Jonathan J Hunt, Timothy P Lillicrap, and David Silver · 2015
Cited alongside, same era.
Doubly robust off-policy value evaluation for reinforcement learning
Nan Jiang and Lihong Li · 2015
Cited alongside, same era.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Cited alongside, same era.
Víctor Elvira, Luca Martino, and Christian P Robert · 2018
Later among the works it cites.
Dynamic few-shot visual learning without forgetting
Spyros Gidaris and Nikos Komodakis · 2018
Later among the works it cites.
Deep reinforcement learning that matters
Peter Henderson, Riashat Islam, Philip Bachman, Joelle Pineau, Doina Precup, and David Meger · 2018
Later among the works it cites.
Evolved policy gradients
Rein Houthooft, Yuhua Chen, Phillip Isola, Bradly Stadie, Filip Wolski, OpenAI Jonathan Ho, and Pieter Abbeel · 2018
Later among the works it cites.
Are deep policy gradient algorithms truly policy gradient algorithms?
Andrew Ilyas, Logan Engstrom, Shibani Santurkar, Dimitris Tsipras, Firdaus Janoos, Larry Rudolph, and Aleksander Madry · 2018
Later among the works it cites.
On first-order meta-learning algorithms
Alex Nichol, Joshua Achiam, and John Schulman · 2018
Later among the works it cites.
Promp: Proximal meta-policy search
Jonas Rothfuss, Dennis Lee, Ignasi Clavera, Tamim Asfour, and Pieter Abbeel · 2018
Later among the works it cites.
A baseline for few-shot image classification
Guneet S Dhillon, Pratik Chaudhari, Avinash Ravichandran, and Stefano Soatto · 2019
Closest in time.
P3o: Policy-on policy-off policy optimization
Rasool Fakoor, Pratik Chaudhari, and Alexander J Smola · 2019
Closest in time.
Meta-learning with differentiable convex optimization
Kwonjoon Lee, Subhransu Maji, Avinash Ravichandran, and Stefano Soatto · 2019
Closest in time.
Efficient off-policy meta-reinforcement learning via probabilistic context variables
Kate Rakelly, Aurick Zhou, Deirdre Quillen, Chelsea Finn, and Sergey Levine · 2019
Closest in time.