Fetching the paper…
Reading the bibliography…
This paper presents a new family of backpropagation-free neural architectures, Gated Linear Networks (GLNs).
Perceptrons: An Introduction to Computational Geometry
Marvin Minsky and Seymour Papert · 1969
Earlier work this paper cites.
The art of adaptive pattern recognition by a self-organizing neural network
G. A. Carpenter and S. Grossberg · 1988
Earlier work this paper cites.
Neural networks and principal component analysis: Learning from examples without local minima
P. Baldi and K. Hornik · 1989
Earlier work this paper cites.
Catastrophic interference in connectionist networks: The sequential learning problem
Michael McCloskey and Neal J. Cohen · 1989
Earlier work this paper cites.
Approximation capabilities of multilayer feedforward networks
Kurt Hornik · 1991
Earlier work this paper cites.
Catastrophic forgetting, rehearsal and pseudorehearsal
Anthony V. Robins · 1995
Earlier work this paper cites.
Multitask learning
Rich Caruana · 1997
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann Lecun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
Fast text compression with neural networks
Matthew Mahoney · 2000
Earlier work this paper cites.
Similarity estimation techniques from rounding algorithms
M.S. Charikar · 2002
Earlier work this paper cites.
Training products of experts by minimizing contrastive divergence
Geoffrey E. Hinton · 2002
Earlier work this paper cites.
Online convex programming and generalized infinitesimal gradient ascent
Martin Zinkevich · 2003
Earlier work this paper cites.
Adaptive weighing of context models for lossless data compression
Matthew Mahoney · 2005
Earlier work this paper cites.
Logarithmic regret algorithms for online convex optimization
Elad Hazan, Amit Agarwal, and Satyen Kale · 2007
Earlier work this paper cites.
The neural autoregressive distribution estimator
Hugo Larochelle and Iain Murray · 2011
Earlier work this paper cites.
Mixing strategies in data compression
Christopher Mattern · 2012
Cited alongside, same era.
Decaf: A deep convolutional activation feature for generic visual recognition
Jeff Donahue, Yangqing Jia, Oriol Vinyals, Judy Hoffman, Ning Zhang, Eric Tzeng, and Trevor Darrell · 2013
Cited alongside, same era.
An empirical investigation of catastrophic forgetting in gradient-based neural networks, 2013
Ian J. Goodfellow, Mehdi Mirza, Da Xiao, Aaron Courville, and Yoshua Bengio · 2013
Cited alongside, same era.
Data Compression Explained
Matthew Mahoney · 2013
Cited alongside, same era.
Linear and geometric mixtures - analysis
Christopher Mattern · 2013
Cited alongside, same era.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Pixel recurrent neural networks
Aäron Van Den Oord, Nal Kalchbrenner, and Koray Kavukcuoglu · 2016
Later among the works it cites.
https://fsix.github.io/mnist/
Dibya Ghosh and Alvin Wan, 2017 · 2017
Later among the works it cites.
http://www.byronknoll.com/cmix.html
Byron Knoll, 2017 · 2017
Later among the works it cites.
Overcoming catastrophic forgetting in neural networks
James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A. Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, Demis Hassabis, Claudia Clopath, Dharshan Kumaran, and Raia Hadsell · 2017
Later among the works it cites.
icarl: Incremental classifier and representation learning
Sylvestre-Alvise Rebuffi, Alexander Kolesnikov, Georg Sperl, and Christoph H. Lampert · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Andrew M. Saxe, James L. McClelland, and Surya Ganguli · 2013
Cited alongside, same era.
Deep inside convolutional networks: Visualising image classification models and saliency maps
Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman · 2013
Cited alongside, same era.
Stochastic gradient descent for non-smooth optimization: Convergence results and optimal averaging schemes
Ohad Shamir and Tong Zhang · 2013
Cited alongside, same era.
Rich feature hierarchies for accurate object detection and semantic segmentation
Ross Girshick, Jeff Donahue, Trevor Darrell, and Jitendra Malik · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Cited alongside, same era.
Cnn features off-the-shelf: An astounding baseline for recognition
Ali Sharif Razavian, Hossein Azizpour, Josephine Sullivan, and Stefan Carlsson · 2014
Cited alongside, same era.
Convex optimization: Algorithms and complexity
Sébastien Bubeck et al · 2015
Cited alongside, same era.
Joel Veness, Tor Lattimore, Avishkar Bhoopchand, Agnieszka Grabska-Barwinska, Christopher Mattern, and Peter Toth · 2017
Later among the works it cites.
Continual learning through synaptic intelligence
Friedemann Zenke, Ben Poole, and Surya Ganguli · 2017
Later among the works it cites.
JAX: composable transformations of Python+NumPy programs, 2018
James Bradbury, Roy Frostig, Peter Hawkins, Matthew James Johnson, Chris Leary, Dougal Maclaurin, and Skye Wanderman-Milne · 2018
Later among the works it cites.
Progress & compress: A scalable framework for continual learning, 2018
Jonathan Schwarz, Jelena Luketina, Wojciech M. Czarnecki, Agnieszka Grabska-Barwinska, Yee Whye Teh, Razvan Pascanu, and Raia Hadsell · 2018
Later among the works it cites.
Visual interpretability for deep learning: a survey
Quan-shi Zhang and Song-Chun Zhu · 2018
Later among the works it cites.
Anytime Online-to-Batch, optimism and acceleration
Ashok Cutkosky · 2019
Closest in time.
Meta-learning of sequential strategies
Pedro A. Ortega, Jane X. Wang, Mark Rowland, Tim Genewein, Zeb Kurth-Nelson, Razvan Pascanu, Nicolas Heess, Joel Veness, Alexander Pritzel, Pablo Sprechmann, Siddhant M. Jayakumar, Tom McGrath, Kevin Miller, Mohammad Gheshlaghi Azar, Ian Osband, Neil C. Rabinowitz, András György, Silvia Chiappa, Simon Osindero, Yee Whye Teh, Hado van Hasselt, Nando de Freitas, Matthew Botvinick, and Shane Legg · 2019
Closest in time.
RLax: Reinforcement Learning in JAX, 2020
David Budden, Matteo Hessel, John Quan, and Steven Kapturowski · 2020
Closest in time.
Haiku: Sonnet for JAX, 2020
Tom Hennigan, Trevor Cai, Tamara Norman, and Igor Babuschkin · 2020
Closest in time.