Fetching the paper…
Reading the bibliography…
We consider stochastic optimization when one only has access to biased stochastic oracles of the objective and the gradient, and obtaining stochastic gradients with low biases comes at high costs.
The theory of queues with a single server
David V Lindley · 1952
Earlier work this paper cites.
Use of different Monte Carlo sampling techniques
Herman Kahn · 1955
Earlier work this paper cites.
Convex measures of risk and trading constraints
Hans Föllmer and Alexander Schied · 2002
Earlier work this paper cites.
Probability models , volume 24
John Haigh and J Haigh · 2002
Earlier work this paper cites.
Shortfall as a risk measure: properties, optimization and applications
Dimitris Bertsimas, Geoffrey J Lauprete, and Alexander Samarov · 2004
Earlier work this paper cites.
Learning a similarity metric discriminatively, with application to face verification
Sumit Chopra, Raia Hadsell, and Yann LeCun · 2005
Earlier work this paper cites.
Measuring the risk of large losses
Kay Giesecke, Thorsten Schmidt, and Stefan Weber · 2008
Earlier work this paper cites.
Multilevel monte carlo path simulation
Michael B Giles · 2008
Earlier work this paper cites.
Information-theoretic lower bounds on the oracle complexity of convex optimization
Alekh Agarwal, Martin J Wainwright, Peter Bartlett, and Pradeep Ravikumar · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A Krizhevsky · 2009
Earlier work this paper cites.
Robust stochastic approximation approach to stochastic programming
Arkadi Nemirovski, Anatoli Juditsky, Guanghui Lan, and Alexander Shapiro · 2009
Earlier work this paper cites.
Making gradient descent optimal for strongly convex stochastic optimization
Alexander Rakhlin, Ohad Shamir, and Karthik Sridharan · 2012
Earlier work this paper cites.
Subgaussian random variables: An expository note
Omar Rivasplata · 2012
Earlier work this paper cites.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
Saeed Ghadimi and Guanghui Lan · 2013
Earlier work this paper cites.
Multi-digit number recognition from street view imagery using deep convolutional neural networks
Ian J Goodfellow, Yaroslav Bulatov, Julian Ibarz, Sacha Arnoud, and Vinay Shet · 2013
Earlier work this paper cites.
Optimal pricing and capacity sizing for the gi/gi/1 queue
Chihoon Lee and Amy R Ward · 2014
Earlier work this paper cites.
Unbiased monte carlo for optimization and functions of expectations via multi-level randomization
Jose H Blanchet and Peter W Glynn · 2015
Earlier work this paper cites.
Multi-level stochastic approximation algorithms
Noufel Frikha · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
(bandit) convex optimization with biased noisy gradient oracles
Xiaowei Hu, LA Prashanth, András György, and Csaba Szepesvári · 2016
Earlier work this paper cites.
Convex risk measures: efficient computations via monte carlo
Zhaolin Hu and Zhang Dali · 2016
Earlier work this paper cites.
Linear convergence of gradient and proximal-gradient methods under the polyak-łojasiewicz condition
Hamed Karimi, Julie Nutini, and Mark Schmidt · 2016
Earlier work this paper cites.
Unbiased simulation for optimizing stochastic function compositions
Jose Blanchet, Donald Goldfarb, Garud Iyengar, Fengpei Li, and Chaoxu Zhou · 2017
Earlier work this paper cites.
Variational inference: A review for statisticians
David M Blei, Alp Kucukelbir, and Jon D McAuliffe · 2017
Earlier work this paper cites.
Learning from conditional distributions via dual embeddings
Bo Dai, Niao He, Yunpeng Pan, Byron Boots, and Le Song · 2017
Earlier work this paper cites.
UCI machine learning repository, 2017
Dua Dheeru and Efi Karra Taniskidou · 2017
Earlier work this paper cites.
Sarah: A novel method for machine learning problems using stochastic recursive gradient
Lam M Nguyen, Jie Liu, Katya Scheinberg, and Martin Takáč · 2017
Cited alongside, same era.
Optimization methods for large-scale machine learning
Léon Bottou, Frank E Curtis, and Jorge Nocedal · 2018
Cited alongside, same era.
Stochastic gradient descent with biased but consistent gradient estimators
Jie Chen and Ronny Luss · 2018
Cited alongside, same era.
SBEED: Convergent reinforcement learning with nonlinear function approximation
Bo Dai, Albert Shaw, Lihong Li, Lin Xiao, Niao He, Zhen Liu, Jianshu Chen, and Le Song · 2018
Cited alongside, same era.
Spider: Near-optimal non-convex optimization via stochastic path-integrated differential estimator
Cong Fang, Chris Junchi Li, Zhouchen Lin, and Tong Zhang · 2018
Cited alongside, same era.
Understanding contrastive representation learning through alignment and uniformity on the hypersphere
Tongzhou Wang and Phillip Isola · 2020
Later among the works it cites.
Stochastic bias-reduced gradient methods
Hilal Asi, Yair Carmon, Arun Jambulapati, Yujia Jin, and Aaron Sidford · 2021
Later among the works it cites.
Data-driven optimization: A reproducing kernel hilbert space approach
Dimitris Bertsimas and Nihal Koduri · 2021
Later among the works it cites.
Closing the gap: Tighter analysis of alternating stochastic gradient methods for bilevel problems
Tianyi Chen, Yuejiao Sun, and Wotao Yin · 2021
Later among the works it cites.
Multilevel monte carlo variational inference
Masahiro Fujisawa and Issei Sato · 2021
Later among the works it cites.
Online estimation and optimization of utility-based shortfall risk
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bilevel programming for hyperparameter optimization and meta-learning
Luca Franceschi, Paolo Frasconi, Saverio Salzo, Riccardo Grazzi, and Massimiliano Pontil · 2018
Cited alongside, same era.
Approximation methods for bilevel programming
Saeed Ghadimi and Mengdi Wang · 2018
Cited alongside, same era.
Representation learning with contrastive predictive coding
Aaron van den Oord, Yazhe Li, and Oriol Vinyals · 2018
Cited alongside, same era.
Efficient optimization of loops and limits with randomized telescoping sums
Alex Beatson and Ryan P Adams · 2019
Cited alongside, same era.
Jose H Blanchet, Peter W Glynn, and Yanan Pei · 2019
Cited alongside, same era.
Momentum-based variance reduction in non-convex sgd
Ashok Cutkosky and Francesco Orabona · 2019
Cited alongside, same era.
General multilevel adaptations for stochastic approximation algorithms of robbins–monro and polyak–ruppert type
Steffen Dereich and Thomas Müller-Gronbach · 2019
Cited alongside, same era.
Vishwajit Hegde, Arvind S Menon, LA Prashanth, and Krishna Jagannathan · 2021
Later among the works it cites.
On the bias-variance-cost tradeoff of stochastic optimization
Yifan Hu, Xin Chen, and Niao He · 2021
Later among the works it cites.
Page: A simple and optimal probabilistic gradient estimator for nonconvex optimization
Zhize Li, Hongyan Bao, Xiangliang Zhang, and Peter Richtárik · 2021
Later among the works it cites.
Variance reduction for non-convex stochastic optimization: General analysis and new applications
Liang Zhang · 2021
Later among the works it cites.
Recapp: Crafting a more efficient catalyst for convex optimization
Yair Carmon, Arun Jambulapati, Yujia Jin, and Aaron Sidford · 2022
Later among the works it cites.
Shortfall risk models when information on loss function is incomplete
Erick Delage, Shaoyan Guo, and Huifu Xu · 2022
Later among the works it cites.
Adapting to mixing time in stochastic optimization with markovian data
Ron Dorfman and Kfir Yehuda Levy · 2022
Later among the works it cites.
Theoretical convergence of multi-step model-agnostic meta-learning
Kaiyi Ji, Junjie Yang, and Yingbin Liang · 2022
Later among the works it cites.
Stochastic constrained dro with a complexity independent of sample size
Qi Qi, Jiameng Lyu, Er Wei Bai, Tianbao Yang, et al · 2022
Later among the works it cites.
Faster single-loop algorithms for minimax optimization without strong concavity
Junchi Yang, Antonio Orvieto, Aurelien Lucchi, and Niao He · 2022
Later among the works it cites.
Lower bounds for non-convex stochastic optimization
Yossi Arjevani, Yair Carmon, John C Duchi, Dylan J Foster, Nathan Srebro, and Blake Woodworth · 2023
Later among the works it cites.
Regularization for wasserstein distributionally robust optimization
Waïss Azizian, Franck Iutzeler, and Jérôme Malick · 2023
Later among the works it cites.
An online learning approach to dynamic pricing and capacity sizing in service systems
Xinyun Chen, Yunan Liu, and Guiyu Hong · 2023
Later among the works it cites.
Constructing unbiased gradient estimators with finite variance for conditional stochastic optimization
Takashi Goda and Wataru Kitade · 2023
Later among the works it cites.
Contextual stochastic bilevel optimization
Yifan Hu, Jie Wang, Yao Xie, Andreas Krause, and Daniel Kuhn · 2023
Later among the works it cites.
Libsvm, 2023
Chih-Jen Lin · 2023
Later among the works it cites.
Adaptive stochastic optimization algorithms for problems with biased oracles
Yin Liu and Sam Davanloo Tajbakhsh · 2023
Later among the works it cites.
Not all semantics are created equal: contrastive self-supervised learning with automatic temperature individualization
Zi-Hao Qiu, Quanqi Hu, Zhuoning Yuan, Denny Zhou, Lijun Zhang, and Tianbao Yang · 2023
Later among the works it cites.
Sinkhorn distributionally robust optimization
Jie Wang, Rui Gao, and Yao Xie · 2023
Later among the works it cites.
Revisiting inexact fixed-point iterations for min-max problems: Stochasticity and structured nonconvexity
Ahmet Alacaoglu, Donghwan Kim, and Stephen Wright · 2024
Closest in time.
A guide through the zoo of biased SGD
Yury Demidovich, Grigory Malinovsky, Igor Sokolov, and Peter Richtárik · 2024
Closest in time.