Fetching the paper…
Reading the bibliography…
We propose a formulation for nonlinear recurrent models that includes simple parametric models of recurrent neural networks as a special case.
A note on analytic functions in the unit circle
R. E. A. C. Paley and A. Zygmund · 1932
Earlier work this paper cites.
On the uniform convergence of relative frequencies of events to their probabilities
V. N. Vapnik and A. Y. Chervonenkis · 1971
Earlier work this paper cites.
A method for solving the convex programming problem with convergence rate
Y. E. Nesterov · 1983
Earlier work this paper cites.
Decoupling: From dependence to independence
V. H. de la Peña and E. Giné · 1999
Earlier work this paper cites.
Oracle Inequalities in Empirical Risk Minimization and Sparse Recovery Problems
V. Koltchinskii · 2011
Earlier work this paper cites.
A probabilistic theory of pattern recognition , volume 31
L. Devroye, L. Györfi, and G. Lugosi · 2013
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
R. Johnson and T. Zhang · 2013
Cited alongside, same era.
Introductory Lectures on Convex Optimization: A Basic Course
Y. Nesterov · 2013
Cited alongside, same era.
SAGA: A fast incremental gradient method with support for non-strongly convex composite objectives
A. Defazio, F. Bach, and S. Lacoste-Julien · 2014
Cited alongside, same era.
Minimizing finite sums with the stochastic average gradient
M. Schmidt, N. Le Roux, and F. Bach · 2017
Cited alongside, same era.
How many samples are needed to estimate a convolutional neural network?
S. S. Du, Y. Wang, X. Zhai, S. Balakrishnan, R. Salakhutdinov, and A. Singh · 2018
Cited alongside, same era.
Finite time identification in unstable linear systems
Gradient descent learns linear dynamical systems
M. Hardt, T. Ma, and B. Recht · 2018
Later among the works it cites.
Non-asymptotic identification of LTI systems from a single trajectory
S. Oymak and N. Ozay · 2018
Later among the works it cites.
Learning without mixing: Towards a sharp analysis of linear system identification
M. Simchowitz, H. Mania, S. Tu, M. I. Jordan, and B. Recht · 2018
Later among the works it cites.
Stable recurrent models
J. Miller and M. Hardt · 2019
Closest in time.
Stochastic gradient descent learns state equations with nonlinear activations
S. Oymak · 2019
Closest in time.
Near optimal finite time identification of arbitrary linear dynamical systems
T. Sarkar and A. Rakhlin · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. K. S. Faradonbeh, A. Tewari, and G. Michailidis · 2018
Cited alongside, same era.
Closest in time.