Fetching the paper…
Reading the bibliography…
Majorization-minimization (MM) is a family of optimization methods that iteratively reduce a loss by minimizing a locally-tight upper bound, called a majorizer.
Minimization of functions having Lipschitz continuous first partial derivatives
Larry Armijo · 1966
Earlier work this paper cites.
Interval Analysis , volume 4
Ramon E Moore · 1966
Earlier work this paper cites.
Linear regression with non-normal error terms
Richard Zeckhauser and Mark Thompson · 1970
Earlier work this paper cites.
Global optimization using interval analysis: the one-dimensional case
Eldon R Hansen · 1979
Earlier work this paper cites.
Monotonicity of quadratic-approximation algorithms
Dankmar Böhning and Bruce G Lindsay · 1988
Earlier work this paper cites.
Convergence of the majorization method for multidimensional scaling
Jan De Leeuw · 1988
Earlier work this paper cites.
Calculus, Volume 1
Tom M Apostol · 1991
Earlier work this paper cites.
The majorization approach to multidimensional scaling for Minkowski distances
Patrick JF Groenen, Rudolf Mathar, and Willem J Heiser · 1995
Earlier work this paper cites.
Numerical Optimization
Jorge Nocedal and Stephen J Wright · 1999
Earlier work this paper cites.
Quantile regression via an MM algorithm
David R Hunter and Kenneth Lange · 2000
Earlier work this paper cites.
Applied Interval Analysis
Luc Jaulin, Michel Kieffer, Olivier Didrit, and Éric Walter · 2001
Earlier work this paper cites.
Global Optimization Using Interval Analysis: Revised and Expanded , volume 264
Eldon Hansen and G William Walster · 2003
Earlier work this paper cites.
Introductory Lectures on Convex Optimization: A Basic Course , volume 87
Yurii Nesterov · 2003
Cited alongside, same era.
Online convex programming and generalized infinitesimal gradient ascent
Martin Zinkevich · 2003
Cited alongside, same era.
MM algorithms for generalized Bradley-Terry models
David R Hunter · 2004
Cited alongside, same era.
A tutorial on MM algorithms
David R Hunter and Kenneth Lange · 2004
Cited alongside, same era.
A software tool for the exponential power distribution: The normalp package
Angelo Mineo and Mariantonietta Ruggieri · 2005
Cited alongside, same era.
Numerical recipes 3rd edition: The art of scientific computing
William H Press, Saul A Teukolsky, William T Vetterling, and Brian P Flannery · 2007
Cited alongside, same era.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
Saeed Ghadimi and Guanghui Lan · 2013
Later among the works it cites.
Stochastic majorization-minimization algorithms for large-scale optimization
Julien Mairal · 2013
Later among the works it cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2015
Later among the works it cites.
Fast DNN training based on auxiliary function technique
Dung T Tran, Nobutaka Ono, and Emmanuel Vincent · 2015
Later among the works it cites.
Block Relaxation Methods in Statistics
Jan de Leeuw · 2016
Later among the works it cites.
MM Optimization Algorithms
Kenneth Lange · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The complexity of optimizing over a simplex, hypercube or sphere: a short survey
Etienne De Klerk · 2008
Cited alongside, same era.
SVM-Maj: a majorization approach to linear support vector machines with different hinge errors
Patrick JF Groenen, Georgi Nalbantov, and Jan C Bioch · 2008
Cited alongside, same era.
Multidimensional scaling using majorization: SMACOF in R
Jan De Leeuw and Patrick Mair · 2009
Cited alongside, same era.
Adaptive bound optimization for online convex optimization
H Brendan McMahan and Matthew Streeter · 2010
Cited alongside, same era.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Cited alongside, same era.
Wide neural networks of any depth evolve as linear models under gradient descent
Jaehoon Lee, Lechao Xiao, Samuel Schoenholz, Yasaman Bahri, Roman Novak, Jascha Sohl-Dickstein, and Jeffrey Pennington · 2019
Later among the works it cites.
A refined laser method and faster matrix multiplication
Josh Alman and Virginia Vassilevska Williams · 2021
Later among the works it cites.
Evaluation of neural architectures trained with square loss vs cross-entropy in classification tasks
Like Hui and Mikhail Belkin · 2021
Later among the works it cites.
Automatic gradient descent: Deep learning without hyperparameters
Jeremy Bernstein, Chris Mingard, Kevin Huang, Navid Azizan, and Yisong Yue · 2023
Closest in time.
Automatically bounding the Taylor remainder series: Tighter bounds and new applications
Matthew Streeter and Joshua V Dillon · 2023
Closest in time.