Fetching the paper…
Reading the bibliography…
Machine learning develops rapidly, which has made many theoretical breakthroughs and is widely applied in various fields.
V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu, “Asynchronous methods for deep reinforcement learning,” in International Conference on Machine Learning , 2016, pp. 1928–1937
1937
Earlier work this paper cites.
H. Robbins and S. Monro, “A stochastic approximation method,” The Annals of Mathematical Statistics , pp. 400–407, 1951
1951
Earlier work this paper cites.
H. Robbins and S. Monro, “A stochastic approximation method,” The Annals of Mathematical Statistics , pp. 400–407, 1951
1951
Earlier work this paper cites.
M. R. Hestenes and E. Stiefel, Methods of conjugate gradients for solving linear systems . NBS Washington, DC, 1952
1952
Earlier work this paper cites.
M. Frank and P. Wolfe, “An algorithm for quadratic programming,” Naval Research Logistics Quarterly , vol. 3, pp. 95–110, 1956
1956
Earlier work this paper cites.
R. Fletcher and M. J. Powell, “A rapidly convergent descent method for minimization,” The Computer Journal , vol. 6, pp. 163–168, 1963
1963
Earlier work this paper cites.
B. T. Polyak, “Some methods of speeding up the convergence of iteration methods,” USSR Computational Mathematics and Mathematical Physics , vol. 4, pp. 1–17, 1964
1964
Earlier work this paper cites.
G. H. Ball and D. J. Hall, “A clustering technique for summarizing multivariate data,” Behavioral Science , vol. 12, pp. 153–155, 1967
1967
Earlier work this paper cites.
M. J. Powell, “A method for nonlinear constraints in minimization problems,” Optimization , pp. 283–298, 1969
1969
Earlier work this paper cites.
D. F. Shanno, “Conditioning of quasi-Newton methods for function minimization,” Mathematics of Computation , vol. 24, pp. 647–656, 1970
1970
Earlier work this paper cites.
C. G. Broyden, “The convergence of a class of double-rank minimization algorithms: The new algorithm,” IMA Journal of Applied Mathematics , vol. 6, pp. 222–231, 1970
1970
Earlier work this paper cites.
R. Fletcher, “A new approach to variable metric algorithms,” The Computer Journal , vol. 13, pp. 317–322, 1970
1970
Earlier work this paper cites.
D. Goldfarb, “A family of variable-metric methods derived by variational means,” Mathematics of Computation , vol. 24, pp. 23–26, 1970
1970
Earlier work this paper cites.
J. E. Dennis, Jr, and J. J. Moré, “Quasi-Newton methods, motivation and theory,” SIAM Review , vol. 19, pp. 46–89, 1977
1977
Earlier work this paper cites.
J. A. Hartigan and M. A. Wong, “Algorithm AS 136: A k-means clustering algorithm,” Journal of the Royal Statistical Society. Series C (Applied Statistics) , vol. 28, pp. 100–108, 1979
1979
Earlier work this paper cites.
J. Nocedal, “Updating quasi-Newton matrices with limited storage,” Mathematics of Computation , vol. 35, pp. 773–782, 1980
1980
Earlier work this paper cites.
F. Murtagh, “A survey of recent advances in hierarchical clustering algorithms,” The Computer Journal , vol. 26, pp. 354–359, 1983
1983
Earlier work this paper cites.
A. S. Nemirovsky and D. B. Yudin, Problem Complexity and Method Efficiency in Optimization . John Wiley & Sons, 1983
1983
Earlier work this paper cites.
Y. Nesterov, “A method for unconstrained convex minimization problem with the rate of convergence O ( 1 k 2 ) O(\frac{1}{k^{2}}) ,” Doklady Akademii Nauk SSSR , vol. 269, pp. 543–547, 1983
1983
Earlier work this paper cites.
S. Kirkpatrick, C. D. Gelatt, and M. P. Vecchi, “Optimization by simulated annealing,” Science , vol. 220, pp. 671–680, 1983
1983
Earlier work this paper cites.
M. Fukushima, “A modified Frank-Wolfe algorithm for solving the traffic assignment problem,” Transportation Research Part B: Methodological , vol. 18, pp. 169–177, 1984
1984
Earlier work this paper cites.
J. Schmidhuber, “Evolutionary principles in self-referential learning, or on learning how to learn: the meta-meta-… hook,” Ph.D. dissertation, Technische Universität München, München, Germany, 1987
1987
Earlier work this paper cites.
S. Wold, K. Esbensen, and P. Geladi, “Principal component analysis,” Chemometrics and Intelligent Laboratory Systems , vol. 2, pp. 37–52, 1987
1987
Earlier work this paper cites.
S. Duane, A. D. Kennedy, B. J. Pendleton, and D. Roweth, “Hybrid monte carlo,” Physics Letters B , vol. 195, pp. 216–222, 1987
1987
Earlier work this paper cites.
D. C. Liu and J. Nocedal, “On the limited memory BFGS method for large scale optimization,” Mathematical programming , vol. 45, pp. 503–528, 1989
1989
Earlier work this paper cites.
P. T. Harker and J. Pang, “A damped-Newton method for the linear complementarity problem,” Lectures in Applied Mathematics , vol. 26, pp. 265–284, 1990
1990
Earlier work this paper cites.
C. Darken and J. E. Moody, “Note on learning rate schedules for stochastic optimization,” in Advances in Neural Information Processing Systems , 1991, pp. 832–838
1991
Earlier work this paper cites.
W. C. Davidon, “Variable metric method for minimization,” SIAM Journal on Optimization , vol. 1, pp. 1–17, 1991
1991
Earlier work this paper cites.
C. Darken, J. Chang, and J. Moody, “Learning rate schedules for faster stochastic gradient search,” in Neural Networks for Signal Processing , 1992, pp. 3–12
1992
Earlier work this paper cites.
C. J. Watkins and P. Dayan, “Q-learning,” Machine Learning , vol. 8, pp. 279–292, 1992
1992
Earlier work this paper cites.
J. Alspector, R. Meir, B. Yuhas, A. Jayakumar, and D. Lippe, “A parallel gradient descent method for learning in analog VLSI neural networks,” in Advances in Neural Information Processing Systems , 1993, pp. 836–844
1993
Earlier work this paper cites.
J. Bromley, I. Guyon, Y. LeCun, E. Säckinger, and R. Shah, “Signature verification using a ”siamese” time delay neural network,” in Advances in Neural Information Processing Systems , 1994, pp. 737–744
1994
Earlier work this paper cites.
J. R. Shewchuk, “An introduction to the conjugate gradient method without the agonizing pain,” Carnegie Mellon University, Tech. Rep., 1994
1994
Earlier work this paper cites.
G. A. Rummery and M. Niranjan, “On-line Q-learning using connectionist systems,” Cambridge University Engineering Department, Tech. Rep., 1994
1994
Earlier work this paper cites.
L. P. Kaelbling, M. L. Littman, and A. W. Moore, “Reinforcement learning: A survey,” Journal of Artificial Intelligence Research , vol. 4, pp. 237–285, 1996
1996
Earlier work this paper cites.
A. Nagurney and P. Ramanujam, “Transportation network policy modeling with goal targets and generalized penalty functions,” Transportation Science , vol. 30, pp. 3–13, 1996
1996
Earlier work this paper cites.
L. Breiman, “Bagging predictors,” Machine Learning , vol. 24, pp. 123–140, 1996
1996
Earlier work this paper cites.
P. Y. Ayala and H. B. Schlegel, “A combined method for determining reaction paths, minima, and transition state geometries,” The Journal of Chemical Physics , vol. 107, pp. 375–384, 1997
1997
Earlier work this paper cites.
M. Raydan, “The barzilai and borwein gradient method for the large scale unconstrained minimization problem,” SIAM Journal on Optimization , vol. 7, pp. 26–33, 1997
1997
Earlier work this paper cites.
Y. LeCun and L. Bottou, “Gradient-based learning applied to document recognition,” Proceedings of the IEEE , vol. 86, pp. 2278–2324, 1998
1998
Earlier work this paper cites.
R. S. Sutton and A. G. Barto, Reinforcement Learning: An Introduction . MIT press, 1998
1998
Earlier work this paper cites.
S. I. Amari, “Natural gradient works efficiently in learning,” Neural Computation , vol. 10, pp. 251–276, 1998
1998
Earlier work this paper cites.
M. Mitchell, An Introduction to Genetic Algorithms . MIT press, 1998
1998
Earlier work this paper cites.
C. S. Adjiman and S. Dallwig, “A global optimization method, α \alpha bb, for general twice-differentiable constrained NLPs–I. theoretical advances,” Computers & Chemical Engineering , vol. 22, pp. 1137–1158, 1998
1998
Earlier work this paper cites.
C. Adjiman, C. Schweiger, and C. Floudas, “Mixed-integer nonlinear optimization in process synthesis,” in Handbook of combinatorial optimization , 1998, pp. 1–76
1998
Earlier work this paper cites.
A. Demiriz and K. P. Bennett, “Semi-supervised clustering using genetic algorithms,” Artificial Neural Networks in Engineering , vol. 1, pp. 809–814, 1999
1999
Earlier work this paper cites.
K. P. Bennett and A. Demiriz, “Semi-supervised support vector machines,” in Advances in Neural Information processing systems , 1999, pp. 368–374
1999
Earlier work this paper cites.
M. E. Tipping and C. M. Bishop, “Probabilistic principal component analysis,” Journal of the Royal Statistical Society: Series B (Statistical Methodology) , vol. 61, pp. 611–622, 1999
1999
Earlier work this paper cites.
L. C. Baird III and A. W. Moore, “Gradient descent for general reinforcement learning,” in Advances in Neural Information Processing Systems , 1999, pp. 968–974
1999
Earlier work this paper cites.
D. P. Bertsekas, Nonlinear Programming . Athena Scientific Belmont, 1999
1999
Earlier work this paper cites.
T. Huckle, “Approximate sparsity patterns for the inverse of a matrix and preconditioning,” Applied Numerical Mathematics , vol. 30, pp. 291–303, 1999
1999
Earlier work this paper cites.
M. I. Jordan, Z. Ghahramani, T. S. Jaakkola, and L. K. Saul, “An introduction to variational methods for graphical models,” Machine Learning , vol. 37, pp. 183–233, 1999
1999
Earlier work this paper cites.
S. Guha, R. Rastogi, and K. Shim, “ROCK: A robust clustering algorithm for categorical attributes,” Information Systems , vol. 25, pp. 345–366, 2000
2000
Earlier work this paper cites.
V. Castro and J. Yang, “A fast and robust general purpose clustering algorithm,” in Knowledge Discovery in Databases and Data Mining , 2000, pp. 208–218
2000
Earlier work this paper cites.
B. He, H. Yang, and S. Wang, “Alternating direction method with self-adaptive penalty parameters for monotone variational inequalities,” Journal of Optimization Theory and Applications , vol. 106, pp. 337–356, 2000
2000
Earlier work this paper cites.
R. H. Byrd, J. C. Gilbert, and J. Nocedal, “A trust region method based on interior point techniques for nonlinear programming,” Mathematical Programming , vol. 89, pp. 149–185, 2000
2000
Earlier work this paper cites.
A. Likas and A. Stafylopatis, “Training the random neural network using quasi-Newton methods,” European Journal of Operational Research , vol. 126, pp. 331–339, 2000
2000
Earlier work this paper cites.
C. Ding, X. He, H. Zha, and H. D. Simon, “Adaptive dimension reduction for clustering high dimensional data,” in IEEE International Conference on Data Mining , 2002, pp. 147–154
2002
Earlier work this paper cites.
M. Benzi, “Preconditioning techniques for large linear systems: a survey,” Journal of Computational Physics , vol. 182, pp. 418–477, 2002
2002
Earlier work this paper cites.
N. N. Schraudolph, “Fast curvature matrix-vector products for second-order gradient descent,” Neural Computation , vol. 14, pp. 1723–1738, 2002
2002
Earlier work this paper cites.
N. N. Schraudolph and T. Graepel, “Conjugate directions for stochastic gradient descent,” in International Conference on Artificial Neural Networks , 2002, pp. 1351–1356
2002
Earlier work this paper cites.
M. Avriel, Nonlinear Programming: Analysis and Methods . Dover Publications, 2003
2003
Earlier work this paper cites.
S. Boyd and L. Vandenberghe, Convex Optimization . Cambridge University Press, 2004
2004
Earlier work this paper cites.
L. Bottou and Y. L. Cun, “Large scale online learning,” in Advances in Neural Information Processing Systems , 2004, pp. 217–224
2004
Earlier work this paper cites.
O. Chapelle and A. Zien, “Semi-supervised classification by low density separation.” in International Conference on Artificial Intelligence and Statistics , 2005, pp. 57–64
2005
Earlier work this paper cites.
Z.-H. Zhou and M. Li, “Semi-supervised regression with co-training.” in International Joint Conferences on Artificial Intelligence , 2005, pp. 908–913
2005
Earlier work this paper cites.
M. I. Lourakis, “A brief description of the levenberg-marquardt algorithm implemented by levmar,” Foundation of Research and Technology , vol. 4, pp. 1–6, 2005
2005
Earlier work this paper cites.
J. C. Spall, Introduction to Stochastic Search and Optimization: Estimation, Simulation, and Control . Wiley-Interscience, 2005
2005
Earlier work this paper cites.
C. M. Bishop, Pattern Recognition and Machine Learning . Springer, 2006
2006
Earlier work this paper cites.
J. Nocedal and S. J. Wright, Numerical Optimization . Springer, 2006
2006
Earlier work this paper cites.
W. Sun and Y. X. Yuan, Optimization theory and methods: nonlinear programming . Springer Science & Business Media, 2006
2006
Earlier work this paper cites.
D. Zhang and Z.-H. Zhou, “Semi-supervised dimensionality reduction,” in SIAM International Conference on Data Mining , 2007, pp. 629–634
2007
Earlier work this paper cites.
——, “Branch and bound for semi-supervised support vector machines,” in Advances in Neural Information Processing Systems , 2007, pp. 217–224
2007
Earlier work this paper cites.
N. N. Schraudolph, J. Yu, and S. Günter, “A stochastic quasi-Newton method for online convex optimization,” in Artificial Intelligence and Statistics , 2007, pp. 436–443
2007
Earlier work this paper cites.
L. Hei, “Practical techniques for nonlinear optimization,” Ph.D. dissertation, Northwestern University, America, 2007
2007
Earlier work this paper cites.
O. Chapelle, V. Sindhwani, and S. S. Keerthi, “Optimization techniques for semi-supervised support vector machines,” Journal of Machine Learning Research , vol. 9, pp. 203–233, 2008
2008
Earlier work this paper cites.
L. Bottou and O. Bousquet, “The tradeoffs of large scale learning,” in Advances in Neural Information Processing Systems , 2008, pp. 161–168
2008
Cited alongside, same era.
M. Dorigo, M. Birattari, C. Blum, M. Clerc, T. Stützle, and A. Winfield, Ant Colony Optimization and Swarm Intelligence . Springer, 2008
2008
Cited alongside, same era.
M. J. Wainwright and M. I. Jordan, “Graphical models, exponential families, and variational inference,” Foundations and Trends in Machine Learning , vol. 1, pp. 1–305, 2008
2008
Cited alongside, same era.
C. Andrieu and J. Thoms, “A tutorial on adaptive MCMC,” Statistics and Computing , vol. 18, pp. 343–373, 2008
2008
Cited alongside, same era.
Y. Bengio, “Learning deep architectures for AI,” Foundations and Trends in Machine Learning , vol. 2, pp. 1–127, 2009
2009
Cited alongside, same era.
S. Lai, L. Xu, and K. Liu, “Recurrent convolutional neural networks for text classification,” in Association for the Advancement of Artificial Intelligence , 2015, pp. 2267–2273
2015
Later among the works it cites.
2015
Later among the works it cites.
Y. Xia and J. Wang, “A bi-projection neural network for solving constrained quadratic optimization problems,” IEEE Transactions on Neural Networks and Learning Systems , vol. 27, no. 2, pp. 214–224, 2015
2015
Later among the works it cites.
S. Zhang, Y. Xia, and J. Wang, “A complex-valued projection neural network for constrained optimization of real functions in complex variables,” IEEE Transactions on Neural Networks and Learning Systems , vol. 26, no. 12, pp. 3227–3238, 2015
2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Yang, K. Yu, Y. Gong, and T. S. Huang, “Linear spatial pyramid matching using sparse coding for image classification,” in IEEE Conference on Computer Vision and Pattern Recognition , 2009, pp. 1794–1801
2009
Cited alongside, same era.
B. Kulis and S. Basu, “Semi-supervised graph clustering: a kernel approach,” Machine Learning , vol. 74, pp. 1–22, 2009
2009
Cited alongside, same era.
A. Nemirovski, A. Juditsky, G. Lan, and A. Shapiro, “Robust stochastic approximation approach to stochastic programming,” SIAM Journal on Optimization , vol. 19, pp. 1574–1609, 2009
2009
Cited alongside, same era.
A. Agarwal, M. J. Wainwright, P. L. Bartlett, and P. K. Ravikumar, “Information-theoretic lower bounds on the oracle complexity of convex optimization,” in Advances in Neural Information Processing Systems , 2009, pp. 1–9
2009
Cited alongside, same era.
A. R. Conn, K. Scheinberg, and L. N. Vicente, Introduction to Derivative-Free Optimization . Society for Industrial and Applied Mathematics, 2009
2009
Cited alongside, same era.
S. Bhatnagar, R. S. Sutton, M. Ghavamzadeh, and M. Lee, “Natural actor-critic algorithms,” Automatica , vol. 45, pp. 2471–2482, 2009
2009
Cited alongside, same era.
Y. Nesterov, “Primal-dual subgradient methods for convex problems,” Mathematical Programming , vol. 120, pp. 221–259, 2009
2009
Cited alongside, same era.
2015
Later among the works it cites.
V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, and G. Ostrovski, “Human-level control through deep reinforcement learning,” Nature , vol. 518, pp. 529–533, 2015
2015
Later among the works it cites.
G. Koch, R. Zemel, and R. Salakhutdinov, “Siamese neural networks for one-shot image recognition,” in International Conference on Machine Learning WorkShop , 2015, pp. 1–30
2015
Later among the works it cites.
J. Weston, S. Chopra, and A. Bordes, “Memory networks,” in International Conference on Learning Representations , 2015, pp. 1–15
2015
Later among the works it cites.
W. Yin and H. Schütze, “Multichannel variable-size convolution for sentence classification,” in Conference on Computational Language Learning , 2015, pp. 204–214
2015
Later among the works it cites.
R. Ge, F. Huang, C. Jin, and Y. Yuan, “Escaping from saddle points—online stochastic gradient for tensor decomposition,” in Conference on Learning Theory , 2015, pp. 797–842
2015
Later among the works it cites.
M. Patriksson, The Traffic Assignment Problem: Models and Methods . Dover Publications, 2015
2015
Later among the works it cites.
——, “Global convergence of online limited memory BFGS,” Journal of Machine Learning Research , vol. 16, pp. 3151–3181, 2015
2015
Later among the works it cites.
R. Grosse and R. Salakhudinov, “Scaling up natural gradient by sparsely factorizing the inverse fisher matrix,” in International Conference on Machine Learning , 2015, pp. 2304–2313
2015
Later among the works it cites.
J. Martens and R. Grosse, “Optimizing neural networks with Kronecker-factored approximate curvature,” in International Conference on Machine Learning , 2015, pp. 2408–2417
2015
Later among the works it cites.
2015
Later among the works it cites.
J. Hensman, A. G. d. G. Matthews, and Z. Ghahramani, “Scalable variational gaussian process classification,” in International Conference on Artificial Intelligence and Statistics , 2015, pp. 351–360
2015
Later among the works it cites.
M. Betancourt, “The fundamental incompatibility of scalable Hamiltonian monte carlo and naive data subsampling,” in International Conference on Machine Learning , 2015, pp. 533–540
2015
Later among the works it cites.
2015
Later among the works it cites.
2016
Later among the works it cites.
P. Xu, J. Yang, F. Roosta-Khorasani, C. Ré, and M. W. Mahoney, “Sub-sampled Newton methods with non-uniform sampling,” in Advances in Neural Information Processing Systems , 2016, pp. 3000–3008
2016
Later among the works it cites.
P. Liu and X. Qiu, “Recurrent neural network for text classification with multi-task learning,” in International Joint Conferences on Artificial Intelligence , 2016, pp. 2873–2879
2016
Later among the works it cites.
2016
Later among the works it cites.
S. S. Mousavi, M. Schukat, and E. Howley, “Deep reinforcement learning: an overview,” in SAI Intelligent Systems Conference , 2016, pp. 426–440
2016
Later among the works it cites.
O. Vinyals, C. Blundell, T. Lillicrap, D. Wierstra et al. , “Matching networks for one shot learning,” in Advances in Neural Information Processing Systems , 2016, pp. 3630–3638
2016
Later among the works it cites.
A. Santoro, S. Bartunov, M. Botvinick, D. Wierstra, and T. Lillicrap, “Meta-learning with memory-augmented neural networks,” in International Conference on Machine Learning , 2016, pp. 1842–1850
2016
Later among the works it cites.
M. Andrychowicz, M. Denil, S. Gomez, M. W. Hoffman, D. Pfau, T. Schaul, B. Shillingford, and N. De Freitas, “Learning to learn by gradient descent by gradient descent,” in Advances in Neural Information Processing Systems , 2016, pp. 3981–3989
2016
Later among the works it cites.
S. Ravi and H. Larochelle, “Optimization as a model for few-shot learning,” in International Conference on Learning Representations , 2016, pp. 1–11
2016
Later among the works it cites.
2016
Later among the works it cites.
I. Goodfellow, Y. Bengio, and A. Courville, Deep Learning . MIT Press, 2016
2016
Later among the works it cites.
Z. Allen-Zhu and E. Hazan, “Variance reduction for faster non-convex optimization,” in International Conference on Machine Learning , 2016, pp. 699–707
2016
Later among the works it cites.
S. J. Reddi, A. Hefny, S. Sra, B. Poczos, and A. Smola, “Stochastic variance reduction for nonconvex optimization,” in International Conference on Machine Learning , 2016, pp. 314–323
2016
Later among the works it cites.
R. H. Byrd, S. L. Hansen, J. Nocedal, and Y. Singer, “A stochastic quasi- method for large-scale optimization,” SIAM Journal on Optimization , vol. 26, pp. 1008–1031, 2016
2016
Later among the works it cites.
P. Moritz, R. Nishihara, and M. Jordan, “A linearly-convergent stochastic L-BFGS algorithm,” in Artificial Intelligence and Statistics , 2016, pp. 249–258
2016
Later among the works it cites.
A. S. Berahas, J. Nocedal, and M. Takác, “A multi-batch L-BFGS method for machine learning,” in Advances in Neural Information Processing Systems , 2016, pp. 1055–1063
2016
Later among the works it cites.
R. Gower, D. Goldfarb, and P. Richtárik, “Stochastic block BFGS: Squeezing more curvature out of data,” in International Conference on Machine Learning , 2016, pp. 1869–1878
2016
Later among the works it cites.
C. Audet and M. Kokkolaras, Blackbox and Derivative-Free Optimization: Theory, Algorithms and Applications . Springer, 2016
2016
Later among the works it cites.
S. Diamond and S. Boyd, “Cvxpy: A python-embedded modeling language for convex optimization,” Journal of Machine Learning Research , vol. 17, pp. 2909–2913, 2016
2016
Later among the works it cites.
M. Abadi, P. Barham, J. Chen, Z. Chen, A. Davis, J. Dean, M. Devin, S. Ghemawat, G. Irving, and M. Isard, “Tensorflow: a system for large-scale machine learning,” in USENIX Symposium on Operating Systems Design and Implementations , 2016, pp. 265–283
2016
Later among the works it cites.
T. Dozat, “Incorporating nesterov momentum into adam,” in International Conference on Learning Representations , 2016, pp. 1–14
2016
Later among the works it cites.
D. Silver, A. Huang, C. J. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, and M. Lanctot, “Mastering the game of go with deep neural networks and tree search,” Nature , vol. 529, pp. 484–489, 2016
2016
Later among the works it cites.
2016
Later among the works it cites.
2016
Later among the works it cites.
J. Martens, Second-Order Optimization For Neural Networks . University of Toronto (Canada), 2016
2016
Later among the works it cites.
A. Ullah and J. Ahmad, “Action recognition in video sequences using deep bi-directional LSTM with CNN features,” IEEE Access , vol. 6, pp. 1155–1166, 2017
2017
Later among the works it cites.
2017
Later among the works it cites.
C. Finn, P. Abbeel, and S. Levine, “Model-agnostic meta-learning for fast adaptation of deep networks,” in International Conference on Machine Learning , 2017, pp. 1126–1135
2017
Later among the works it cites.
O. Vinyals, “Model vs optimization meta learning,” http://metalearning-symposium.ml/files/vinyals.pdf , 2017
2017
Later among the works it cites.
J. Snell, K. Swersky, and R. Zemel, “Prototypical networks for few-shot learning,” in Advances in Neural Information Processing Systems , 2017, pp. 4077–4087
2017
Later among the works it cites.
C. Ledig, L. Theis, F. Huszár, J. Caballero, A. Cunningham, A. Acosta, A. Aitken, A. Tejani, J. Totz, Z. Wang et al. , “Photo-realistic single image super-resolution using a generative adversarial network,” in Computer Vision and Pattern Recognition , 2017, pp. 4681–4690
2017
Later among the works it cites.
Y. Wu, E. Mansimov, R. B. Grosse, S. Liao, and J. Ba, “Scalable trust-region method for deep reinforcement learning using Kronecker-factored approximation,” in Advances in Neural Information Processing Systems , 2017, pp. 5279–5288
2017
Later among the works it cites.
P. Chen and L. Jiao, “Semi-supervised double sparse graphs based discriminant analysis for dimensionality reduction,” Pattern Recognition , vol. 61, pp. 361–378, 2017
2017
Later among the works it cites.
M. Schmidt, N. Le Roux, and F. Bach, “Minimizing finite sums with the stochastic average gradient,” Mathematical Programming , vol. 162, pp. 83–112, 2017
2017
Later among the works it cites.
D. Hallac, C. Wong, S. Diamond, A. Sharang, S. Boyd, and J. Leskovec, “Snapvx: A network-based convex optimization solver,” Journal of Machine Learning Research , vol. 18, pp. 1–5, 2017
2017
Later among the works it cites.
2017
Later among the works it cites.
2017
Later among the works it cites.
2017
Later among the works it cites.
Y. Li, “Deep reinforcement learning: An overview,” arXiv preprint arXiv:1701.07274 , 2017
2017
Later among the works it cites.
D. M. Blei, A. Kucukelbir, and J. D. McAuliffe, “Variational inference: A review for statisticians,” Journal of the American Statistical Association , vol. 112, pp. 859–877, 2017
2017
Later among the works it cites.
B. Carpenter, A. Gelman, M. D. Hoffman, D. Lee, B. Goodrich, M. Betancourt, M. Brubaker, J. Guo, P. Li, and A. Riddell, “Stan: A probabilistic programming language,” Journal of Statistical Software , vol. 76, pp. 1–37, 2017
2017
Later among the works it cites.
P. Jain and P. Kar, “Non-convex optimization for machine learning,” Foundations and Trends in Machine Learning , vol. 10, pp. 142–336, 2017
2017
Later among the works it cites.
S. Balakrishnan, M. J. Wainwright, and B. Yu, “Statistical guarantees for the em algorithm: From population to sample-based analysis,” The Annals of Statistics , vol. 45, pp. 77–120, 2017
2017
Later among the works it cites.
P. Jain, S. Kakade, R. Kidambi, P. Netrapalli, and A. Sidford, “Parallelizing stochastic gradient descent for least squares regression: mini-batching, averaging, and model misspecification,” Journal of Machine Learning Research , vol. 18, 2018
2018
Later among the works it cites.
R. Bollapragada, R. H. Byrd, and J. Nocedal, “Exact and inexact subsampled newton methods for optimization,” IMA Journal of Numerical Analysis , vol. 1, pp. 1–34, 2018
2018
Later among the works it cites.
Y. Xia and J. Wang, “Robust regression estimation based on low-dimensional recurrent neural networks,” IEEE Transactions on Neural Networks and Learning Systems , vol. 29, no. 12, pp. 5935–5946, 2018
2018
Later among the works it cites.
S. J. Reddi, S. Kale, and S. Kumar, “On the convergence of Adam and beyond,” in International Conference on Learning Representations , 2018, pp. 1–23
2018
Later among the works it cites.
E. Cheung, Optimization Methods for Semi-Supervised Learning . University of Waterloo, 2018
2018
Later among the works it cites.
R. S. Sutton and A. G. Barto, Reinforcement Learning: An Introduction . MIT Press, 2018
2018
Later among the works it cites.
Z. Allen-Zhu, “Natasha 2: Faster non-convex optimization than SGD,” in Advances in Neural Information Processing Systems , 2018, pp. 2675–2686
2018
Later among the works it cites.
L. Bottou, F. E. Curtis, and J. Nocedal, “Optimization methods for large-scale machine learning,” Society for Industrial and Applied Mathematics Review , vol. 60, pp. 223–311, 2018
2018
Later among the works it cites.
X. Liu and S. Liu, “Limited-memory bfgs optimization of recurrent neural network language models for speech recognition,” in International Conference on Acoustics, Speech and Signal Processing , 2018, pp. 6114–6118
2018
Later among the works it cites.
2018
Later among the works it cites.
J. Hu, B. Jiang, L. Lin, Z. Wen, and Y.-x. Yuan, “Structured quasi-newton methods for optimization with orthogonality constraints,” SIAM Journal on Scientific Computing , vol. 41, pp. 2239–2269, 2019
2019
Closest in time.
J. Pajarinen, H. L. Thai, R. Akrour, J. Peters, and G. Neumann, “Compatible natural gradient policy search,” Machine Learning , pp. 1–24, 2019
2019
Closest in time.
A. S. Berahas, R. H. Byrd, and J. Nocedal, “Derivative-free optimization of noisy functions via quasi-newton methods,” SIAM Journal on Optimization , vol. 29, pp. 965–993, 2019
2019
Closest in time.
Y. Xia, J. Wang, and W. Guo, “Two projection neural networks with reduced model complexity for nonlinear programming,” IEEE Transactions on Neural Networks and Learning Systems , pp. 1–10, 2019
2019
Closest in time.
M. Abdullah Jamal and G.-J. Qi, “Task agnostic meta-learning for few-shot learning,” in The IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2019, pp. 1–11
2019
Closest in time.