Fetching the paper…
Reading the bibliography…
Real-life data are often non-IID due to complex distributions and interactions, and the sensitivity to the distribution of samples can differ among learning models.
Eine informationstheoretische ungleichung und ihre anwendung auf beweis der ergodizitaet von markoffschen ketten
Imre Csiszár · 1964
Earlier work this paper cites.
I-divergence geometry of probability distributions and minimization problems
Imre Csiszár · 1975
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
Rademacher and gaussian complexities: Risk bounds and structural results
Peter L. Bartlett and Shahar Mendelson · 2002
Earlier work this paper cites.
Permutation methods: a basis for exact inference
Michael D Ernst · 2004
Earlier work this paper cites.
Detecting change in data streams
Daniel Kifer, Shai Ben-David, and Johannes Gehrke · 2004
Earlier work this paper cites.
A Modern Introduction to Probability and Statistics: Understanding why and how
Frederik Michel Dekking, Cornelis Kraaikamp, Hendrik Paul Lopuhaä, and Ludolf Erwin Meester · 2005
Earlier work this paper cites.
Eulerian calculus for the contraction in the wasserstein distance
Felix Otto and Michael Westdickenberg · 2005
Earlier work this paper cites.
Pattern recognition and machine learning
Christopher M Bishop and Nasser M Nasrabadi · 2006
Earlier work this paper cites.
Direct importance estimation with model selection and its application to covariate shift adaptation
Masashi Sugiyama, Shinichi Nakajima, Hisashi Kashima, Paul von Bünau, and Motoaki Kawanabe · 2007
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Earlier work this paper cites.
A theory of learning from different domains
Shai Ben-David, John Blitzer, Koby Crammer, Alex Kulesza, Fernando Pereira, and Jennifer Wortman Vaughan · 2010
Earlier work this paper cites.
Multiple kernel learning algorithms
Mehmet Gönen and Ethem Alpaydin · 2011
Earlier work this paper cites.
A kernel two-sample test
Arthur Gretton, Karsten M. Borgwardt, Malte J. Rasch, Bernhard Schölkopf, and Alexander J. Smola · 2012
Earlier work this paper cites.
Machine learning - a probabilistic perspective
Kevin P. Murphy · 2012
Earlier work this paper cites.
Learning with noisy labels
Nagarajan Natarajan, Inderjit S. Dhillon, Pradeep Ravikumar, and Ambuj Tewari · 2013
Earlier work this paper cites.
Maximum mean discrepancy for class ratio estimation: Convergence bounds and kernel selection
Arun Shankar Iyer, J. Saketha Nath, and Sunita Sarawagi · 2014
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P. Kingma and Max Welling · 2014
Earlier work this paper cites.
One weird trick for parallelizing convolutional neural networks
Alex Krizhevsky · 2014
Earlier work this paper cites.
Understanding Machine Learning From Theory to Algorithms
Shai Shalev-Shwartz and Shai Ben-David · 2014
Earlier work this paper cites.
Fast two-sample testing with analytic representations of probability measures
Kacper Chwialkowski, Aaditya Ramdas, Dino Sejdinovic, and Arthur Gretton · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Cited alongside, same era.
Training deep neuralnetworks on noisy labels with bootstrapping
Scott E. Reed, Honglak Lee, Dragomir Anguelov, Christian Szegedy, Dumitru Erhan, and Andrew Rabinovich · 2015
Cited alongside, same era.
Learning with symmetric label noise: The importance of being unhinged
Brendan van Rooyen, Aditya Krishna Menon, and Robert C. Williamson · 2015
Cited alongside, same era.
How machine learning won the higgs boson challenge
Claire Adam-Bourdarios, Glen Cowan, Cécile Germain, Isabelle Guyon, Balázs Kégl, and David Rousseau · 2016
Cited alongside, same era.
Towards open set deep networks
Abhijit Bendale and Terrance E. Boult · 2016
mixup: Beyond empirical risk minimization
Hongyi Zhang, Moustapha Cissé, Yann N. Dauphin, and David Lopez-Paz · 2018
Later among the works it cites.
Invariant risk minimization
Martín Arjovsky, Léon Bottou, Ishaan Gulrajani, and David Lopez-Paz · 2019
Later among the works it cites.
Robust learning from untrusted sources
Nikola Konstantinov and Christoph Lampert · 2019
Later among the works it cites.
On the jensen-shannon symmetrization of distances relying on abstract means
Frank Nielsen · 2019
Later among the works it cites.
Learning robust global representationsby penalizing local predictive power
Haohan Wang, Songwei Ge, Zachary C. Lipton, and Eric P. Xing · 2019
Later among the works it cites.
Are anchor points really indispensable in label-noise learning?
Xiaobo Xia, Tongliang Liu, Nannan Wang, Bo Han, Chen Gong, Gang Niu, and Masashi Sugiyama · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Interpretable distribution features with maximum testing power
Wittawat Jitkrittum, Zoltán Szabó, Kacper P. Chwialkowski, and Arthur Gretton · 2016
Cited alongside, same era.
A baseline for detecting misclassified and out-of-distribution examples in neural networks
Dan Hendrycks and Kevin Gimpel · 2017
Cited alongside, same era.
Uniform convergence rates for kernel density estimation
Heinrich Jiang · 2017
Cited alongside, same era.
Deeper, broader and artier domain generalization
Da Li, Yongxin Yang, Yi-Zhe Song, and Timothy M. Hospedales · 2017
Cited alongside, same era.
Revisiting classifier two-sample tests
David Lopez-Paz and Maxime Oquab · 2017
Cited alongside, same era.
Two-sample testing using deep learning
Matthias Kirchler, Shahryar Khorasani, Marius Kloft, and Christoph Lippert · 2020
Later among the works it cites.
Learning deep kernels for non-parametric two-sample tests
Feng Liu, Wenkai Xu, Jie Lu, Guangquan Zhang, Arthur Gretton, and Danica J. Sutherland · 2020
Later among the works it cites.
Test-time training with self-supervision for generalization under distribution shifts
Yu Sun, Xiaolong Wang, Zhuang Liu, John Miller, Alexei A. Efros, and Moritz Hardt · 2020
Later among the works it cites.
Optimal bounds between f-divergences and integral probability metrics
Rohit Agrawal and Thibaut Horel · 2021
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby · 2021
Later among the works it cites.
Regressive domain adaptation for unsupervised keypoint detection
Junguang Jiang, Yifei Ji, Ximei Wang, Yufeng Liu, Jianmin Wang, and Mingsheng Long · 2021
Later among the works it cites.
Beyond i.i.d.: Non-iid thinking, informatics, and learning
Longbing Cao · 2022
Later among the works it cites.
Non-iid learning
Longbing Cao · 2022
Later among the works it cites.
Classification logit two-sample testing by neural networks for differentiating near manifold densities
Xiuyuan Cheng and Alexander Cloninger · 2022
Later among the works it cites.
A convnet for the 2020s
Zhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer, Trevor Darrell, and Saining Xie · 2022
Later among the works it cites.
How do vision transformers work?
Namuk Park and Songkuk Kim · 2022
Later among the works it cites.
To smooth or not? when label smoothing meets noisy labels
Jiaheng Wei, Hangyu Liu, Tongliang Liu, Gang Niu, Masashi Sugiyama, and Yang Liu · 2022
Later among the works it cites.
Comparing distributions by measuring differences that affect decision making
Shengjia Zhao, Abhishek Sinha, Yutong He, Aidan Perreault, Jiaming Song, and Stefano Ermon · 2022
Later among the works it cites.
Sparse fusion mixture-of-experts are domain generalizable learners
Bo Li, Jingkang Yang, Jiawei Ren, Yezhen Wang, and Ziwei Liu · 2023
Closest in time.
Revealing the distributional vulnerability of discriminators by implicit generators
Zhilin Zhao, Longbing Cao, and Kun-Yu Lin · 2023
Closest in time.