Fetching the paper…
Reading the bibliography…
Recursive least squares (RLS) algorithms were once widely used for training small-scale neural networks, due to their fast convergence.
B. T. Polyak, “Some methods of speeding up the convergence of iteration methods,” USSR Computational Mathematics and Mathematical Physics , vol. 4, no. 5, pp. 1–17, 1964
1964
Earlier work this paper cites.
M. R. Azimi-Sadjadi, S. Citrin, and S. Sheedvash, “Supervised learning process of multi-layer perceptron neural networks using fast least squares,” in Proc. of 1990 Int. Conf. Acoust., Speech, and Signal Process. , Albuquerque, New Mexico, USA, 1990, pp. 1381–1384
1990
Earlier work this paper cites.
J. Park and J. Kim, “Online recurrent extreme learning machine and its application to time-series prediction,” in Proc. of 2017 Int. Joint Conf. on Neural Networks , Anchorage, AK, USA, May 2017, pp. 1983–1990
1990
Earlier work this paper cites.
L. Bottou, “Stochastic gradient learning in neural networks,” in Proc. of 1991 Neuro-Nîmes , Nimes, France, 1991
1991
Earlier work this paper cites.
R. Battiti, “First- and second-order methods for learning: Between steepest descent and newton’s method,” Neural Computation , vol. 4, no. 2, pp. 141–166, 1992
1992
Earlier work this paper cites.
M. R. Azimi-Sadjadi and R. Liou, “Fast learning process of multilayer neural networks using recursive least squares method,” IEEE Trans. Signal Process. , vol. 40, no. 2, pp. 446–450, 1992
1992
Earlier work this paper cites.
F. Biegler-König and F. Bärmann, “A learning algorithm for multilayered neural networks based on linear least squares problems,” Neural Networks , vol. 6, no. 1, pp. 127–131, 1993
1993
Earlier work this paper cites.
J. Y. F. Yam and T. W. S. Chow, “Accelerated training algorithm for feedforward neural networks based on least squares method,” Neural Process. Lett. , vol. 2, no. 4, pp. 20–25, 1995
1995
Earlier work this paper cites.
S. Ergezinger and E. Thomsen, “An accelerated learning algorithm for multilayer perceptrons: Optimization layer by layer,” IEEE Trans. Neural Networks , vol. 6, no. 1, pp. 31–42, 1995
1995
Earlier work this paper cites.
J. Y. F. Yam and T. W. S. Chow, “Extended least squares based algorithm for training feedforward networks,” IEEE Trans. Neural Networks , vol. 8, no. 3, pp. 806–810, 1997
1997
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural Comput. , vol. 9, no. 8, pp. 1735–1780, 1997
1997
Earlier work this paper cites.
J. Bilski and L. Rutkowski, “A fast training algorithm for neural networks,” IEEE Trans. on Circuits and Syst.II: Analog and Digital Signal Process. , vol. 45, no. 6, pp. 749–753, 1998
1998
Earlier work this paper cites.
O. Stan and E. W. Kamen, “A local linearized least squares algorithm for training feedforward neural networks,” IEEE Trans. Neural Netw. Learn Syst. , vol. 11, no. 2, pp. 487–495, 2000
2000
Earlier work this paper cites.
S. Cho, T. W. S. Chow, and Y. Fang, “Training recurrent neural networks using optimization layer-by- layer recursive least squares algorithm for vibration signals system identification and fault diagnostic analysis,” J. of Intell. Syst. , vol. 11, no. 2, pp. 125–154, 2001
2001
Earlier work this paper cites.
H. Jaeger, “A tutorial on training recurrent neural networks, covering bppt, rtrl, ekf and the ‘echo state network’ approach,” German National Research Center for Information Technology, Sankt Augustin, Germany, GMD Report 159, 2002
2002
Earlier work this paper cites.
O. Fontenla-Romero, D. Erdogmus, J. C. Príncipe, A. Alonso-Betanzos, and E. F. Castillo, “Linear least-squares based methods for neural networks learning,” in Proc. of 2003 Artif. Neural Networks and Neural Inf. Process. , vol. 2714, Istanbul, Turkey, 2003, pp. 84–91
2003
Earlier work this paper cites.
G. Huang, L. Chen, and C. K. Siew, “Universal approximation using incremental constructive feedforward networks with random hidden nodes,” IEEE Trans. on Neural Networks , vol. 17, no. 4, pp. 879–892, 2006
2006
Cited alongside, same era.
J. Martens, “Deep learning via hessian-free optimization,” in Proc. of 27th Int. Conf. Mach. Learn. , Haifa, Israel, 2010, pp. 735–742
2010
Cited alongside, same era.
M. S. Al-Batah, N. A. M. Isa, K. Z. Zamli, and K. A. Azizli, “Modified recursive least squares algorithm to train the hybrid multilayered perceptron (HMLP) network,” Appl. Soft Comput. , vol. 10, no. 1, pp. 236–244, 2010
2010
Cited alongside, same era.
E. M. Eksioglu, “RLS adaptive filtering with sparsity regularization,” in Proc. of 10th Int. Conf. on Inf. Sciences, Signal Process. and their Appl. , Kuala Lumpur, Malaysia, May 2010, pp. 550–553
2010
Cited alongside, same era.
I. Goodfellow, Y. Bengio, and A. Courville, Deep Learning . Cambridge, MA: MIT Press, 2016
2016
Later among the works it cites.
J. Martens, “Second-order optimization for neural networks,” Ph.D. dissertation, University of Toronto, Toronto, Canada, 2016
2016
Later among the works it cites.
S. Pang and X. Yang, “Deep convolutional extreme learning machine and its application in handwritten digit classification,” Comput. Intell. Neurosci. , vol. 2016, pp. 3 049 632:1–3 049 632:10, 2016
2016
Later among the works it cites.
R. Claser and V. H. N. andYuriy V. Zakharov, “A low-complexity RLS-DCD algorithm for volterra system identification,” in Proc. of 24th European Signal Processing Conference , Budapest, Hungary, Aug. 2016, pp. 6–10
2016
Later among the works it cites.
A. C. Wilson, R. Roelofs, M. Stern, N. Srebro, and B. Recht, “The marginal value of adaptive gradient methods in machine learning,” in Proc. of 31st Neural Inf. Process. Syst. , Long Beach, CA, USA, Dec. 2017, pp. 4148–4158
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Duchi, E. Hazan, and Y. Singer, “Adaptive subgradient methods for online learning and stochastic optimization,” J. Mach. Learn. Res. , vol. 12, pp. 2121–2159, 2011
2011
Cited alongside, same era.
J. Martens and I. Sutskever, “Learning recurrent neural networks with hessian-free optimization,” in Proc. of 28th Int. Conf. Mach. Learn. , Bellevue, Washington, USA, 2011, pp. 1033–1040
2011
Cited alongside, same era.
E. M. Eksioglu and A. K. Tanc, “RLS algorithm with convex regularization,” IEEE Signal Process. Lett. , vol. 18, no. 8, pp. 470–473, 2011
2011
Cited alongside, same era.
T. Tieleman and G. Hinton, “Lecture 6e rmsprop: Divide the gradient by a running average of its recent magnitude,” COURSERA: Neural Networks for Mach. Learn. , vol. 4, pp. 26–30, 2012
2012
Cited alongside, same era.
G. Huang, H. Zhou, X. Ding, and R. Zhang, “Extreme learning machine for regression and multiclass classification,” IEEE Trans. Syst. Man Cybern. Part B , vol. 42, no. 2, pp. 513–529, 2012
2012
Cited alongside, same era.
K. B. Petersen and M. S. Pedersen, The Matrix Cookbook . Technical University of Denmark, 2012
2012
Cited alongside, same era.
F. Albu, “Improved variable forgetting factor recursive least square algorithm,” in Proc. of 12th Int. Conf. on Control Automat. Robot. & Vision , Guangzhou, China, Dec. 2012, pp. 1789–1793
2012
Cited alongside, same era.
I. Sutskever, J. Martens, G. E. Dahl, and G. E. Hinton, “On the importance of initialization and momentum in deep learning,” in Proc. of 30th Int. Conf. Mach. Learn. , Atlanta, GA, USA, Jun. 2013, pp. 1139–1147
2013
Cited alongside, same era.
2017
Later among the works it cites.
A. Botev, H. Ritter, and D. Barber, “Practical gauss-newton optimisation for deep learning,” in Proc. of 34th Int. Conf. Mach. Learn. , Sydney, NSW, Australia, 2017, pp. 557–565
2017
Later among the works it cites.
X. Wang, S. Ma, D. Goldfarb, and W. Liu, “Stochastic quasi-newton methods for nonconvex stochastic optimization,” SIAM J. on Optim. , vol. 27, no. 2, pp. 927–956, 2017
2017
Later among the works it cites.
A. Voulodimos, N. Doulamis, A. Doulamis, and E. Protopapadakis, “Deep learning for computer vision: A brief review,” Comp. Int. and Neurosc. , vol. 2018, pp. 7 068 349:1–7 068 349:13, 2018
2018
Later among the works it cites.
R. Mu, “A survey of recommender systems based on deep learning,” IEEE Access , vol. 6, pp. 69 009–69 022, 2018
2018
Later among the works it cites.
C. P. H and H. Cho-jui, “A comparison of second-order methods for deep convolutional neural networks,” in Proc. of 6th Int. Conf. Learn. Representations , Vancouver, BC, Canada, May 2018
2018
Later among the works it cites.
A. Esteva, A. Robicquet, B. Ramsundar, V. Kuleshov, M. DePristo, K. Chou, C. Cui, G. Corrado, S. Thrun, and J. Dean, “A guide to deep learning in healthcare,” Nature Medicine , vol. 25, no. 1, pp. 24–29, Jan. 2019
2019
Later among the works it cites.
D. Goldfarb, Y. Ren, and A. Bahamou, “Practical quasi-newton methods for training deep neural networks,” in Proc. of 34th Neural Inf. Process. Syst. , Vancouver, BC, Canada, Dec. 2020
2020
Later among the works it cites.
P. Xu, F. Roosta, and M. W. Mahoney, “Newton-type methods for non-convex optimization under inexact hessian information,” Math. Program. , vol. 184, no. 1, pp. 35–70, 2020
2020
Later among the works it cites.
A. N. Sadigh, A. H. Taherinia, and H. S. Yazdi, “Analysis of robust recursive least squares: Convergence and tracking,” Signal Process. , vol. 171, p. 107482, 2020
2020
Later among the works it cites.
M. Malik, M. K. Malik, K. Mehmood, and I. Makhdoom, “Automatic speech recognition: A survey,” Multim. Tools Appl. , vol. 80, no. 6, pp. 9411–9457, 2021
2021
Closest in time.
D. W. Otter, J. R. Medina, and J. K. Kalita, “A survey of the usages of deep learning for natural language processing,” IEEE Trans. Neural Networks Learn. Syst. , vol. 32, no. 2, pp. 604–624, 2021
2021
Closest in time.