Fetching the paper…
Reading the bibliography…
As machine learning models are deployed ever more broadly, it becomes increasingly important that they are not only able to perform well on their training distribution, but also yield accurate predictions when confronted with distribution shift.
Shiori Sagawa, Pang Wei Koh, Tatsunori B Hashimoto, and Percy Liang · 1911
Earlier work this paper cites.
Principles of risk minimization for learning theory
Vladimir Vapnik · 1992
Earlier work this paper cites.
Nash convergence of gradient dynamics in general-sum games
Satinder P. Singh, Michael J. Kearns, and Yishay Mansour · 2000
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Learning bounds for importance weighting
Corinna Cortes, Yishay Mansour, and Mehryar Mohri · 2010
Earlier work this paper cites.
Distributionally robust optimization under moment uncertainty with application to data-driven problems
Erick Delage and Yinyu Ye · 2010
Earlier work this paper cites.
Robust solutions of optimization problems affected by uncertain probabilities
Aharon Ben-Tal, Dick Den Hertog, Anja De Waegenaere, Bertrand Melenberg, and Gijs Rennen · 2013
Earlier work this paper cites.
Kullback-leibler divergence constrained distributionally robust optimization
Zhaolin Hu and L Jeff Hong · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
On the accuracy of self-normalized log-linear models
Jacob Andreas, Maxim Rabinovich, Michael I Jordan, and Dan Klein · 2015
Earlier work this paper cites.
Deep learning face attributes in the wild
Ziwei Liu, Ping Luo, Xiaogang Wang, and Xiaoou Tang · 2015
Earlier work this paper cites.
Deep speech 2 : End-to-end speech recognition in english and mandarin
Dario Amodei, Sundaram Ananthanarayanan, Rishita Anubhai, Jingliang Bai, Eric Battenberg, Carl Case, Jared Casper, Bryan Catanzaro, Qiang Cheng, Guoliang Chen, Jie Chen, Jingdong Chen, Zhijie Chen, Mike Chrzanowski, Adam Coates, Greg Diamos, Ke Ding, Niandong Du, Erich Elsen, Jesse Engel, Weiwei Fang, Linxi Fan, Christopher Fougner, Liang Gao, Caixia Gong, Awni Hannun, Tony Han, Lappi Johannes, Bing Jiang, Cai Ju, Billy Jun, Patrick LeGresley, Libby Lin, Junjie Liu, Yang Liu, Weigao Li, Xiangang Li, Dongpeng Ma, Sharan Narang, Andrew Ng, Sherjil Ozair, Yiping Peng, Ryan Prenger, Sheng Qian, Zongfeng Quan, Jonathan Raiman, Vinay Rao, Sanjeev Satheesh, David Seetapun, Shubho Sengupta, Kavya Srinet, Anuroop Sriram, Haiyuan Tang, Liliang Tang, Chong Wang, Jidong Wang, Kaifu Wang, Yi Wang, Zhijian Wang, Zhiqian Wang, Shuang Wu, Likai Wei, Bo Xiao, Wen Xie, Yan Xie, Dani Yogatama, Bin Yuan, Jun Zhan, and Zhenyao Zhu · 2016
Earlier work this paper cites.
Demographic dialectal variation in social media: A case study of African-American English
Su Lin Blodgett, Lisa Green, and Brendan O’Connor · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Pixel recurrent neural networks
Aaron Van Oord, Nal Kalchbrenner, and Koray Kavukcuoglu · 2016
Earlier work this paper cites.
Learning with average top-k loss
Yanbo Fan, Siwei Lyu, Yiming Ying, and Bao-Gang Hu · 2017
Earlier work this paper cites.
Accurate, large minibatch sgd: Training imagenet in 1 hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Earlier work this paper cites.
The mechanics of n-player differentiable games
David Balduzzi, Sebastien Racaniere, James Martens, Jakob Foerster, Karl Tuyls, and Thore Graepel · 2018
Cited alongside, same era.
Recognition in terra incognita
Sara Beery, Grant Van Horn, and Pietro Perona · 2018
Cited alongside, same era.
Gender shades: Intersectional accuracy disparities in commercial gender classification
Joy Buolamwini and Timnit Gebru · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Cited alongside, same era.
Learning models with uniform performance via distributionally robust optimization
John Duchi and Hongseok Namkoong · 2018
Cited alongside, same era.
Large scale crowdsourcing and characterization of twitter abusive behavior
Distributionally Robust Optimization: A Review
Hamed Rahimian and Sanjay Mehrotra · 2019
Later among the works it cites.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf · 2019
Later among the works it cites.
The risk of racial bias in hate speech detection
Maarten Sap, Dallas Card, Saadia Gabriel, Yejin Choi, and Noah A. Smith · 2019
Later among the works it cites.
Adaptive sampling for stochastic risk-averse learning
Sebastian Curi, Kfir Y. Levy, Stefanie Jegelka, and Andreas Krause · 2020
Later among the works it cites.
Distributionally robust losses for latent covariate mixtures
J. Duchi, T. Hashimoto, and H. Namkoong · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Antigoni-Maria Founta, Constantinos Djouvas, Despoina Chatzakou, Ilias Leontiadis, Jeremy Blackburn, Gianluca Stringhini, Athena Vakali, Michael Sirivianos, and Nicolas Kourtellis · 2018
Cited alongside, same era.
Fairness without demographics in repeated loss minimization
Tatsunori Hashimoto, Megha Srivastava, Hongseok Namkoong, and Percy Liang · 2018
Cited alongside, same era.
Does distributionally robust supervised learning give robust classifiers?
Weihua Hu, Gang Niu, Issei Sato, and Masashi Sugiyama · 2018
Cited alongside, same era.
MTNT: A testbed for machine translation of noisy text
Paul Michel and Graham Neubig · 2018
Cited alongside, same era.
Training tips for the transformer model
Martin Popel and Ondřej Bojar · 2018
Cited alongside, same era.
Improving language understanding with unsupervised learning
Alec Radford, Karthik Narasimhan, Time Salimans, and Ilya Sutskever · 2018
Cited alongside, same era.
Certifying some distributional robustness with principled adversarial training
Aman Sinha, Hongseok Namkoong, and John Duchi · 2018
Cited alongside, same era.
Distributionally robust counterfactual risk minimization
Louis Faury, Ugo Tanielian, Elvis Dohmatob, Elena Smirnova, and Flavian Vasile · 2020
Later among the works it cites.
In search of lost domain generalization
Ishaan Gulrajani and David Lopez-Paz · 2020
Later among the works it cites.
Don’t stop pretraining: Adapt language models to domains and tasks
Suchin Gururangan, Ana Marasović, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A. Smith · 2020
Later among the works it cites.
Distributional robustness with ipms and links to regularization and gans
Hisham Husain · 2020
Later among the works it cites.
Wilds: A benchmark of in-the-wild distribution shifts
Pang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie, Marvin Zhang, Akshay Balsubramani, Weihua Hu, Michihiro Yasunaga, Richard Lanas Phillips, Irena Gao, et al · 2020
Later among the works it cites.
Large-scale methods for distributionally robust optimization
Daniel Levy, Yair Carmon, John C Duchi, and Aaron Sidford · 2020
Later among the works it cites.
Robust bayesian classification using an optimistic score ratio
Viet Anh Nguyen, Nian Si, and Jose Blanchet · 2020
Later among the works it cites.
Simple data balancing achieves competitive worst-group-accuracy
Badr Youbi Idrissi, Martin Arjovsky, Mohammad Pezeshki, and David Lopez-Paz · 2021
Later among the works it cites.
Modeling the second player in distributionally robust optimization
Paul Michel, Tatsunori Hashimoto, and Graham Neubig · 2021
Later among the works it cites.
Examining and combating spurious features under distribution shift
Chunting Zhou, Xuezhe Ma, Paul Michel, and Graham Neubig · 2021
Later among the works it cites.
Tagging performance correlates with author age
Dirk Hovy and Anders Søgaard · 2079
Closest in time.