Fetching the paper…
Reading the bibliography…
An agent that efficiently accumulates knowledge to develop increasingly sophisticated skills over a long lifetime could advance the frontier of artificial intelligence capabilities.
Funzione caratteristica di un fenomeno aleatorio
Bruno de Finetti · 1929
Earlier work this paper cites.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
William R Thompson · 1933
Earlier work this paper cites.
A mathematical theory of communication
Claude E Shannon · 1948
Earlier work this paper cites.
A new approach to linear filtering and prediction problems
Rudolph Emil Kalman · 1960
Earlier work this paper cites.
Adaptive switching circuits
Bernard Widrow and Marcian E. Hoff, Jr · 1960
Earlier work this paper cites.
de Finetti’s theorem for Markov chains
Persi Diaconis and David Freedman · 1980
Earlier work this paper cites.
Random sampling with a reservoir
Jeffrey S Vitter · 1985
Earlier work this paper cites.
Catastrophic interference in connectionist networks: The sequential learning problem
Michael McCloskey and Neal J. Cohen · 1989
Earlier work this paper cites.
Learning from delayed rewards
Christopher John Cornish Hellaby Watkins · 1989
Earlier work this paper cites.
Connectionist models of recognition memory: constraints imposed by learning and forgetting functions
Roger Ratcliff · 1990
Earlier work this paper cites.
Integrated architectures for learning, planning, and reacting based on approximating dynamic programming
Richard S Sutton · 1990
Earlier work this paper cites.
Adapting bias by gradient descent: An incremental version of delta-bar-delta
Richard S. Sutton · 1992
Earlier work this paper cites.
Continual Learning in Reinforcement Environments
Mark B. Ring · 1994
Earlier work this paper cites.
Instance-based utile distinctions for reinforcement learning with hidden state
R Andrew McCallum · 1995
Earlier work this paper cites.
Neuro-dynamic programming
Dimitri Bertsekas and John N Tsitsiklis · 1996
Earlier work this paper cites.
Learning to learn: Introduction and overview
Sebastian Thrun and Lorien Pratt · 1998
Earlier work this paper cites.
Computational capacity of the universe
Seth Lloyd · 2002
Earlier work this paper cites.
Toward a formal framework for continual learning
Mark B Ring · 2005
Earlier work this paper cites.
Multi-armed bandit, dynamic environments and meta-bandits
Cédric Hartland, Sylvain Gelly, Nicolas Baskiotis, Olivier Teytaud, and Michele Sebag · 2006
Earlier work this paper cites.
Discounted UCB
Levente Kocsis and Csaba Szepesvári · 2006
Earlier work this paper cites.
Universal algorithmic intelligence: A mathematical top → \rightarrow down approach
Marcus Hutter · 2007
Earlier work this paper cites.
On the role of tracking in stationary environments
Richard S Sutton, Anna Koop, and David Silver · 2007
Earlier work this paper cites.
On upper-confidence bound policies for non-stationary bandit problems
Aurélien Garivier and Eric Moulines · 2008
Earlier work this paper cites.
Adapting to a changing environment: the Brownian restless bandits
Aleksandrs Slivkins and Eli Upfal · 2008
Earlier work this paper cites.
Algorithms for reinforcement learning
Csaba Szepesvári · 2010
Earlier work this paper cites.
Thompson sampling for dynamic multi-armed bandits
Neha Gupta, Ole-Christoffer Granmo, and Ashok Agrawala · 2011
Earlier work this paper cites.
Elements of Information Theory
Thomas M Cover and Joy A Thomas · 2012
Earlier work this paper cites.
Q-learning for history-based reinforcement learning
Mayank Daswani, Peter Sunehag, and Marcus Hutter · 2013
Earlier work this paper cites.
An empirical investigation of catastrophic forgetting in gradient-based neural networks
Ian J Goodfellow, Mehdi Mirza, Da Xiao, Aaron Courville, and Yoshua Bengio · 2013
Cited alongside, same era.
Thompson sampling in switching environments with Bayesian online change detection
Joseph Mellor and Jonathan Shapiro · 2013
Cited alongside, same era.
Thompson sampling for Bayesian bandits with resets
Paolo Viappiani · 2013
Cited alongside, same era.
Feature reinforcement learning: state of the art
Mayank Daswani, Peter Sunehag, Marcus Hutter, et al · 2014
Cited alongside, same era.
Exploration vs exploitation with partially observable gaussian autoregressive arms
Julia Kuhn, Michel Mandjes, and Yoni Nazarathy · 2014
Cited alongside, same era.
Wireless channel selection with reward-observing restless multi-armed bandits
Continual backprop: Stochastic gradient descent with persistent randomness
Shibhansh Dohare, Richard S Sutton, and A Rupam Mahmood · 2021
Later among the works it cites.
Sebastian Flennerhag, Yannick Schroecker, Tom Zahavy, Hado van Hasselt, David Silver, and Satinder Singh · 2021
Later among the works it cites.
The clear benchmark: Continual learning on real-world imagery
Zhiqiu Lin, Jia Shi, Deepak Pathak, and Deva Ramanan · 2021
Later among the works it cites.
Reinforcement learning, bit by bit
Xiuyuan Lu, Benjamin Van Roy, Vikranth Dwaracherla, Morteza Ibrahimi, Ian Osband, and Zheng Wen · 2021
Later among the works it cites.
A survey of reinforcement learning algorithms for dynamically varying environments
Sindhu Padakandla · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Julia Kuhn and Yoni Nazarathy · 2015
Cited alongside, same era.
Rl2̂: Fast reinforcement learning via slow reinforcement learning
Yan Duan, John Schulman, Xi Chen, Peter L Bartlett, Ilya Sutskever, and Pieter Abbeel · 2016
Cited alongside, same era.
Overcoming catastrophic forgetting in neural networks
James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, and Demis Hassibis · 2017
Cited alongside, same era.
Taming non-stationary bandits: A Bayesian approach
Vishnu Raj and Sheetal Kalyani · 2017
Cited alongside, same era.
A tutorial on Thompson sampling
Daniel J Russo, Benjamin Van Roy, Abbas Kazerouni, Ian Osband, Zheng Wen, et al · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Cited alongside, same era.
Reinforcement learning: Theory and algorithms
Alekh Agarwal, Nan Jiang, Sham M Kakade, and Wen Sun · 2019
Cited alongside, same era.
Later among the works it cites.
Autonomous reinforcement learning: Formalism and benchmarking
Archit Sharma, Kelvin Xu, Nikhil Sardana, Abhishek Gupta, Karol Hausman, Sergey Levine, and Chelsea Finn · 2021
Later among the works it cites.
Online coreset selection for rehearsal-based continual learning
Jaehong Yoon, Divyam Madaan, Eunho Yang, and Sung Ju Hwang · 2021
Later among the works it cites.
Beyond supervised continual learning: a review, 2022
Benedikt Bagus, Alexander Gepperth, and Timothée Lesort · 2022
Later among the works it cites.
You only live once: Single-life reinforcement learning
Annie Chen, Archit Sharma, Sergey Levine, and Chelsea Finn · 2022
Later among the works it cites.
Simple agent, complex environment: Efficient reinforcement learning with agent states
Shi Dong, Benjamin Van Roy, and Zhengyuan Zhou · 2022
Later among the works it cites.
Training compute-optimal large language models, 2022
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, Tom Hennigan, Eric Noland, Katie Millican, George van den Driessche, Bogdan Damoc, Aurelia Guy, Simon Osindero, Karen Simonyan, Erich Elsen, Jack W. Rae, Oriol Vinyals, and Laurent Sifre · 2022
Later among the works it cites.
Drinking from a firehose: Continual learning with web-scale natural language
Hexiang Hu, Ozan Sener, Fei Sha, and Vladlen Koltun · 2022
Later among the works it cites.
Towards continual reinforcement learning: A review and perspectives
Khimya Khetarpal, Matthew Riemer, Irina Rish, and Doina Precup · 2022
Later among the works it cites.
Control systems and reinforcement learning
Sean Meyn · 2022
Later among the works it cites.
Wide neural networks forget less catastrophically
Seyed Iman Mirzadeh, Arslan Chaudhry, Dong Yin, Huiyi Hu, Razvan Pascanu, Dilan Gorur, and Mehrdad Farajtabar · 2022
Later among the works it cites.
Scaling language models: Methods, analysis & insights from training gopher, 2022
Jack W. Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, Eliza Rutherford, Tom Hennigan, Jacob Menick, Albin Cassirer, Richard Powell, George van den Driessche, Lisa Anne Hendricks, Maribeth Rauh, Po-Sen Huang, Amelia Glaese, Johannes Welbl, Sumanth Dathathri, Saffron Huang, Jonathan Uesato, John Mellor, Irina Higgins, Antonia Creswell, Nat McAleese, Amy Wu, Erich Elsen, Siddhant Jayakumar, Elena Buchatskaya, David Budden, Esme Sutherland, Karen Simonyan, Michela Paganini, Laurent Sifre, Lena Martens, Xiang Lorraine Li, Adhiguna Kuncoro, Aida Nematzadeh, Elena Gribovskaya, Domenic Donato, Angeliki Lazaridou, Arthur Mensch, Jean-Baptiste Lespiau, Maria Tsimpoukelli, Nikolai Grigorev, Doug Fritz, Thibault Sottiaux, Mantas Pajarskas, Toby Pohlen, Zhitao Gong, Daniel Toyama, Cyprien de Masson d’Autume, Yujia Li, Tayfun Terzi, Vladimir Mikulik, Igor Babuschkin, Aidan Clark, Diego de Las Casas, Aurelia Guy, Chris Jones, James Bradbury, Matthew Johnson, Blake Hechtman, Laura Weidinger, Iason Gabriel, William Isaac, Ed Lockhart, Simon Osindero, Laura Rimell, Chris Dyer, Oriol Vinyals, Kareem Ayoub, Jeff Stanway, Lorrayne Bennett, Demis Hassabis, Koray Kavukcuoglu, and Geoffrey Irving · 2022
Later among the works it cites.
Using deepspeed and megatron to train megatron-turing NLG 530b, a large-scale generative language model, 2022
Shaden Smith, Mostofa Patwary, Brandon Norick, Patrick LeGresley, Samyam Rajbhandari, Jared Casper, Zhun Liu, Shrimai Prabhumoye, George Zerveas, Vijay Korthikanti, Elton Zhang, Rewon Child, Reza Yazdani Aminabadi, Julie Bernauer, Xia Song, Mohammad Shoeybi, Yuxiong He, Michael Houston, Saurabh Tiwary, and Bryan Catanzaro · 2022
Later among the works it cites.
Lamda: Language models for dialog applications, 2022
Romal Thoppilan, Daniel De Freitas, Jamie Hall, Noam Shazeer, Apoorv Kulshreshtha, Heng-Tze Cheng, Alicia Jin, Taylor Bos, Leslie Baker, Yu Du, YaGuang Li, Hongrae Lee, Huaixiu Steven Zheng, Amin Ghafouri, Marcelo Menegali, Yanping Huang, Maxim Krikun, Dmitry Lepikhin, James Qin, Dehao Chen, Yuanzhong Xu, Zhifeng Chen, Adam Roberts, Maarten Bosma, Vincent Zhao, Yanqi Zhou, Chung-Ching Chang, Igor Krivokon, Will Rusch, Marc Pickett, Pranesh Srinivasan, Laichee Man, Kathleen Meier-Hellstern, Meredith Ringel Morris, Tulsee Doshi, Renelito Delos Santos, Toju Duke, Johnny Soraker, Ben Zevenbergen, Vinodkumar Prabhakaran, Mark Diaz, Ben Hutchinson, Kristen Olson, Alejandra Molina, Erin Hoffman-John, Josh Lee, Lora Aroyo, Ravi Rajakumar, Alena Butryna, Matthew Lamm, Viktoriya Kuzmina, Joe Fenton, Aaron Cohen, Rachel Bernstein, Ray Kurzweil, Blaise Aguera-Arcas, Claire Cui, Marian Croak, Ed Chi, and Quoc Le · 2022
Later among the works it cites.
Lifelong learning for robust AI systems
Gautam K Vallabha and Jared Markowitz · 2022
Later among the works it cites.
Revealing the real-world applicable setting of online continual learning
Zhenbo Xu, Haimiao Hu, and Liu Liu · 2022
Later among the works it cites.
A definition of continual reinforcement learning
David Abel, André Barreto, Benjamin Van Roy, Doina Precup, Hado van Hasselt, and Satinder Singh · 2023
Closest in time.
Real-time evaluation in online continual learning: A new paradigm
Yasir Ghunaim, Adel Bibi, Kumail Alhamoud, Motasem Alfarra, Hasan Abed Al Kader Hammoud, Ameya Prabhu, Philip HS Torr, and Bernard Ghanem · 2023
Closest in time.
Rapid adaptation in online continual learning: Are we evaluating it right?
Hasan Abed Al Kader Hammoud, Ameya Prabhu, Ser-Nam Lim, Philip HS Torr, Adel Bibi, and Bernard Ghanem · 2023
Closest in time.
An information-theoretic framework for supervised learning, 2023
Hong Jun Jeon, Yifan Zhu, and Benjamin Van Roy · 2023
Closest in time.
Online boundary-free continual learning by scheduled data prior
Hyunseo Koh, Minhyuk Seo, Jihwan Bang, Hwanjun Song, Deokki Hong, Seulki Park, Jung-Woo Ha, and Jonghyun Choi · 2023
Closest in time.
Understanding plasticity in neural networks
Clare Lyle, Zeyu Zheng, Evgenii Nikishin, Bernardo Avila Pires, Razvan Pascanu, and Will Dabney · 2023
Closest in time.
Deep reinforcement learning with plasticity injection
Evgenii Nikishin, Junhyuk Oh, Georg Ostrovski, Clare Lyle, Razvan Pascanu, Will Dabney, and André Barreto · 2023
Closest in time.
A comprehensive survey of continual learning: Theory, method and application
Liyuan Wang, Xingxing Zhang, Hang Su, and Jun Zhu · 2023
Closest in time.