Fetching the paper…
Reading the bibliography…
Distributionally robust optimization (DRO) is a widely-used approach to learn models that are robust against distribution shift.
Generalized gradients of lipschitz functionals
Frank H Clarke · 1981
Earlier work this paper cites.
Optimization and nonsmooth analysis
Frank H Clarke · 1990
Earlier work this paper cites.
Ordinal hyperplanes ranker with cost sensitivities for age estimation
Kuang-Yu Chang, Chu-Song Chen, and Yi-Ping Hung · 2011
Earlier work this paper cites.
Cumulative attribute space for age and crowd density estimation
Ke Chen, Shaogang Gong, Tao Xiang, and Chen Change Loy · 2013
Earlier work this paper cites.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
Saeed Ghadimi and Guanghui Lan · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Mini-batch stochastic approximation methods for nonconvex stochastic composite optimization
Saeed Ghadimi, Guanghui Lan, and Hongchao Zhang · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Stochastic gradient methods for distributionally robust optimization with f-divergences
Hongseok Namkoong and John C Duchi · 2016
Earlier work this paper cites.
Ordinal regression with multiple output cnn for age estimation
Zhenxing Niu, Mo Zhou, Le Wang, Xinbo Gao, and Gang Hua · 2016
Earlier work this paper cites.
Stochastic variance reduction for nonconvex optimization
Sashank J Reddi, Ahmed Hefny, Suvrit Sra, Barnabas Poczos, and Alex Smola · 2016
Cited alongside, same era.
Distributionally robust stochastic programming
Alexander Shapiro · 2017
Cited alongside, same era.
Learning models with uniform performance via distributionally robust optimization
John Duchi and Hongseok Namkoong · 2018
Cited alongside, same era.
Fairness without demographics in repeated loss minimization
Tatsunori Hashimoto, Megha Srivastava, Hongseok Namkoong, and Percy Liang · 2018
Cited alongside, same era.
Does distributionally robust supervised learning give robust classifiers?
Weihua Hu, Gang Niu, Issei Sato, and Masashi Sugiyama · 2018
Cited alongside, same era.
Certifiable distributional robustness with principled adversarial training
Remix: Rebalanced mixup
Hsin-Ping Chou, Shih-Chieh Chang, Jia-Yu Pan, Wei Wei, and Da-Cheng Juan · 2020
Later among the works it cites.
Momentum improves normalized sgd
Ashok Cutkosky and Harsh Mehta · 2020
Later among the works it cites.
A stochastic subgradient method for distributionally robust non-convex learning
Mert Gürbüzbalaban, Andrzej Ruszczyński, and Landi Zhu · 2020
Later among the works it cites.
Dionysios S Kalogerias · 2020
Later among the works it cites.
Large-scale methods for distributionally robust optimization
Daniel Levy, Yair Carmon, John C Duchi, and Aaron Sidford · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Aman Sinha, Hongseok Namkoong, and John Duchi · 2018
Cited alongside, same era.
Lower bounds for non-convex stochastic optimization
Yossi Arjevani, Yair Carmon, John C Duchi, Dylan J Foster, Nathan Srebro, and Blake Woodworth · 2019
Cited alongside, same era.
Stochastic model-based minimization of weakly convex functions
Damek Davis and Dmitriy Drusvyatskiy · 2019
Cited alongside, same era.
Distributionally robust optimization: A review
Hamed Rahimian and Sanjay Mehrotra · 2019
Cited alongside, same era.
Why adam beats sgd for attention models
Jingzhao Zhang, Sai Praneeth Karimireddy, Andreas Veit, Seungyeon Kim, Sashank J Reddi, Sanjiv Kumar, and Suvrit Sra · 2019
Cited alongside, same era.
Improved analysis of clipping algorithms for non-convex optimization
Bohang Zhang, Jikai Jin, Cong Fang, and Liwei Wang
Cited in the paper.
Why gradient clipping accelerates training: A theoretical justification for adaptivity
Jingzhao Zhang, Tianxing He, Suvrit Sra, and Ali Jadbabaie
Cited in the paper.
Convergence of a stochastic gradient method with momentum for non-smooth non-convex optimization
Vien Mai and Mikael Johansson · 2020
Later among the works it cites.
A stochastic subgradient method for nonsmooth nonconvex multi-level composition optimization
Andrzej Ruszczynski · 2020
Later among the works it cites.
Distributionally robust neural networks
Shiori Sagawa, Pang Wei Koh, Tatsunori B. Hashimoto, and Percy Liang · 2020
Later among the works it cites.
Statistical learning with conditional value at risk
Tasuku Soma and Yuichi Yoshida · 2020
Later among the works it cites.
Coping with label shift via distributionally robust optimisation
Jingzhao Zhang, Aditya Krishna Menon, Andreas Veit, Srinadh Bhojanapalli, Sanjiv Kumar, and Suvrit Sra · 2021
Closest in time.