Fetching the paper…
Reading the bibliography…
This paper describes a family of probabilistic architectures designed for online learning under the logarithmic loss.
Perceptrons: An Introduction to Computational Geometry
Marvin Minsky and Seymour Papert · 1969
Earlier work this paper cites.
Computation of matrix chain products. part i
T. C. Hu and M. T. Shing · 1982
Earlier work this paper cites.
Combining probability distributions: A critique and an annotated bibliography
Christian Genest and James V. Zidek · 1986
Earlier work this paper cites.
Data compression using dynamic markov modelling
G. V. Cormack and R. N. S. Horspool · 1987
Earlier work this paper cites.
Arithmetic coding for data compression
Ian H. Witten, Radford M. Neal, and John G. Cleary · 1987
Earlier work this paper cites.
Neurocomputing: Foundations of research
David E. Rumelhart, Geoffrey E. Hinton, and Ronald J. Williams · 1988
Earlier work this paper cites.
Neural networks and principal component analysis: Learning from examples without local minima
P. Baldi and K. Hornik · 1989
Earlier work this paper cites.
Text Compression
T. C. Bell, J. G. Cleary, and I. H. Witten · 1990
Earlier work this paper cites.
Approximation capabilities of multilayer feedforward networks
Kurt Hornik · 1991
Earlier work this paper cites.
Adaptive mixtures of local experts
Robert A. Jacobs, Michael I. Jordan, Steven J. Nowlan, and Geoffrey E. Hinton · 1991
Earlier work this paper cites.
Sequential neural text compression
J. Schmidhuber and S. Heil · 1996
Earlier work this paper cites.
A corpus for the evaluation of lossless compression algorithms
Tim Bell and Ross Arnold · 1997
Earlier work this paper cites.
Reflections on ”the context-tree weighting method: Basic properties”
Frans Willems, Yuri Shtarkov, and Tjalling Tjalkens · 1997
Earlier work this paper cites.
Tracking the best expert
Mark Herbster and Manfred K. Warmuth · 1998
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann Lecun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
Managing Gigabytes (2nd Ed.): Compressing and Indexing Documents and Images
Ian H. Witten, Alistair Moffat, and Timothy C. Bell · 1999
Earlier work this paper cites.
Fast text compression with neural networks
Matthew Mahoney · 2000
Cited alongside, same era.
Training products of experts by minimizing contrastive divergence
Geoffrey E. Hinton · 2002
Cited alongside, same era.
Online convex programming and generalized infinitesimal gradient ascent
Martin Zinkevich · 2003
Cited alongside, same era.
Adaptive weighing of context models for lossless data compression
Matthew Mahoney · 2005
Cited alongside, same era.
Prediction, Learning, and Games
Nicolo Cesa-Bianchi and Gabor Lugosi · 2006
Cited alongside, same era.
A closer look at skip-gram modelling
David Guthrie, Ben Allison, W. Liu, Louise Guthrie, and Yorick Wilks · 2006
Cited alongside, same era.
Draw: A recurrent neural network for image generation
Karol Gregor, Ivo Danihelka, Alex Graves, Danilo Rezende, and Daan Wierstra · 2015
Later among the works it cites.
Compress and control
Joel Veness, Marc G. Bellemare, Marcus Hutter, Alvin Chua, and Guillaume Desjardins · 2015
Later among the works it cites.
An architecture for deep, hierarchical generative models
Philip Bachman · 2016
Later among the works it cites.
Deep online convex optimization with gated games
David Balduzzi · 2016
Later among the works it cites.
Xi Chen, Diederik P. Kingma, Tim Salimans, Yan Duan, Prafulla Dhariwal, John Schulman, Ilya Sutskever, and Pieter Abbeel · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Elad Hazan, Amit Agarwal, and Satyen Kale · 2007
Cited alongside, same era.
Catching up faster in bayesian model selection and model averaging
Tim van Erven, Peter Grunwald, and Steven de Rooij · 2007
Cited alongside, same era.
The radon transform on r n
Sigurdur Helgason · 2011
Cited alongside, same era.
A machine learning perspective on predictive coding with paq8
B. Knoll and N. de Freitas · 2012
Cited alongside, same era.
Mixing strategies in data compression
Christopher Mattern · 2012
Cited alongside, same era.
Markov chains and stochastic stability
Sean P Meyn and Richard L Tweedie · 2012
Cited alongside, same era.
Ishaan Gulrajani, Kundan Kumar, Faruk Ahmed, Adrien Ali Taiga, Francesco Visin, David Vázquez, and Aaron C. Courville · 2016
Later among the works it cites.
Introduction to online convex optimization
Elad Hazan · 2016
Later among the works it cites.
Efficient second order online learning by sketching
Haipeng Luo, Alekh Agarwal, Nicolò Cesa-Bianchi, and John Langford · 2016
Later among the works it cites.
On Statistical Data Compression
Christopher Mattern · 2016
Later among the works it cites.
Pixel recurrent neural networks
Aäron Van Den Oord, Nal Kalchbrenner, and Koray Kavukcuoglu · 2016
Later among the works it cites.
Input switched affine networks: An RNN architecture designed for interpretability
Jakob N. Foerster, Justin Gilmer, Jascha Sohl-Dickstein, Jan Chorowski, and David Sussillo · 2017
Closest in time.
URL https://fsix.github.io/mnist/Deskewing.html
Dibya Ghosh and Alvin Wan, 2017 · 2017
Closest in time.
Hutter prize, 2017
Marcus Hutter · 2017
Closest in time.
Improved strongly adaptive online learning using coin betting
Kwang-Sung Jun, Francesco Orabona, Stephen Wright, and Rebecca Willett · 2017
Closest in time.
URL http://www.byronknoll.com/cmix.html
Byron Knoll, 2017 · 2017
Closest in time.
Count-based exploration with neural density models
Georg Ostrovski, Marc G. Bellemare, Aäron van den Oord, and Rémi Munos · 2017
Closest in time.