Fetching the paper…
Reading the bibliography…
In many practical applications of machine learning data arrives sequentially over time in large chunks.
Approach to anytime learning
John J. Grefenstette and Connie Loggia Ramsey · 1992
Earlier work this paper cites.
Case-based anytime learning
Connie Loggia Ramsey and John J. Grefenstette · 1994
Earlier work this paper cites.
Adaptively growing hierarchical mixtures of experts
Jürgen Fritsch, Michael Finke, and Alex Waibel · 1996
Earlier work this paper cites.
Online algorithms and stochastic approximations
Léon Bottou · 1998
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Leon Bottou, and and Patrick Haffner Yoshua Bengio · 1998
Earlier work this paper cites.
Statistical learning theory
Vladimir Vapnik · 1998
Earlier work this paper cites.
Prediction, learning, and games
Nicolo Cesa-Bianchi and Gábor Lugosi · 2006
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Earlier work this paper cites.
A survey on transfer learning
Sinno Jialin Pan and Qiang Yang · 2010
Earlier work this paper cites.
Adadelta: an adaptive learning rate method
Matthew D Zeiler · 2012
Earlier work this paper cites.
Estimating or propagating gradients through stochastic neurons for conditional computation
Yoshua Bengio, Nicholas Léonard, and Aaron C. Courville · 2013
Cited alongside, same era.
Learning factored representations in a deep mixture of experts
David Eigen, Ilya Sutskever, and Marc’Aurelio Ranzato · 2014
Cited alongside, same era.
Deep sequential neural networks
Ludovic Denoyer and Patrick Gallinari · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2015
Cited alongside, same era.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Later among the works it cites.
fairseq: A fast, extensible toolkit for sequence modeling
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Later among the works it cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Later among the works it cites.
Gshard: Scaling giant models with conditional computation and automatic sharding
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Pooria Joulani, Andras Gyorgy, and Csaba Szepesvari · 2016
Cited alongside, same era.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch · 2016
Cited alongside, same era.
Variance-reduced stochastic gradient descent on streaming data
Ellango Jothimurugesan, Ashraf Tahmasbi, Phillip B. Gibbons, and Srikanta Tirthapura · 2018
Cited alongside, same era.
Learning under concept drift: A review
Jie Lu, Anjin Liu, Fan Dong, Feng Gu, Joao Gama, and Guangquan Zhang · 2018
Cited alongside, same era.
Online deep learning: Learning deep neural networks on the fly
Doyen Sahoo, Quang Pham, Jing Lu, and Steven C.H. Hoi · 2018
Cited alongside, same era.
Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer, and Zhifeng Chen · 2020
Later among the works it cites.
Towards automated deep learning: Analysis of the autodl challenge series 2019
Zhengying Liu, Zhen Xu, Shangeth Rajaa, Meysam Madadi, Julio C. S. Jacques Junior, Sergio Escalera, Adrien Pavao, Sebastien Treguer, Wei-Wei Tu, and Isabelle Guyon · 2020
Later among the works it cites.
Firefly neural architecture descent: a general approach for growing neural networks
Lemeng Wu, Bo Liu, Peter Stone, and Qiang Liu · 2020
Later among the works it cites.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
William Fedus, Barret Zoph, and Noam Shazeer · 2021
Closest in time.
Online learning with optimism and delay
Genevieve Flaspohler, Francesco Orabona, Judah Cohen, Soukayna Mouatadid, Miruna Oprescu, Paulo Orenstein, and Lester Mackey · 2021
Closest in time.