Fetching the paper…
Reading the bibliography…
ML models often exhibit unexpectedly poor behavior when they are deployed in real-world domains.
Right for the wrong reasons: Diagnosing syntactic heuristics in natural language inference
R Thomas McCoy, Ellie Pavlick, and Tal Linzen · 1902
Earlier work this paper cites.
Bret Nestor, Matthew B. A. McDermott, Willie Boag, Gabriela Berner, Tristan Naumann, Michael C Hughes, Anna Goldenberg, and Marzyeh Ghassemi · 1908
Earlier work this paper cites.
R Thomas McCoy, Junghyun Min, and Tal Linzen · 1911
Earlier work this paper cites.
Sun and skin
TB Fitzpatrick · 1975
Earlier work this paper cites.
The use of misclassification costs to learn rule-based decision support models for cost-effective hospital admission strategies
R Ambrosino, B G Buchanan, G F Cooper, and M J Fine · 1995
Earlier work this paper cites.
Priors for infinite networks
Radford M Neal · 1996
Earlier work this paper cites.
Long Short-Term Memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
SemEval-2017 task 1: Semantic textual similarity multilingual and crosslingual focused evaluation
Daniel Cer, Mona Diab, Eneko Agirre, Iñigo Lopez-Gazpio, and Lucia Specia · 2001
Earlier work this paper cites.
Principal components analysis corrects for stratification in genome-wide association studies
Alkes L Price, Nick J Patterson, Robert M Plenge, Michael E Weinblatt, Nancy A Shadick, and David Reich · 2006
Earlier work this paper cites.
PLINK: a tool set for whole-genome association and population-based linkage analyses
Shaun Purcell, Benjamin Neale, Kathe Todd-Brown, Lori Thomas, Manuel A R Ferreira, David Bender, Julian Maller, Pamela Sklar, Paul I W de Bakker, Mark J Daly, and Pak C Sham · 2007
Earlier work this paper cites.
Prediction of individual genetic risk to disease from genome-wide association studies
Naomi R Wray, Michael E Goddard, and Peter M Visscher · 2007
Earlier work this paper cites.
Random features for large-scale kernel machines
Ali Rahimi and Benjamin Recht · 2008
Earlier work this paper cites.
Linkage disequilibrium — understanding the evolutionary past and mapping the medical future
Montgomery Slatkin · 2008
Earlier work this paper cites.
ImageNet: A Large-Scale Hierarchical Image Database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Common polygenic variation contributes to risk of schizophrenia and bipolar disorder
International Schizophrenia Consortium, Shaun M Purcell, Naomi R Wray, Jennifer L Stone, Peter M Visscher, Michael C O’Donovan, Patrick F Sullivan, and Pamela Sklar · 2009
Earlier work this paper cites.
Next generation disparities in human genomics: concerns and remedies
Anna C Need and David B Goldstein · 2009
Earlier work this paper cites.
Eigenvectors of some large sample covariance matrix ensembles
Olivier Ledoit and Sandrine Péché · 2011
Earlier work this paper cites.
Improving disease prediction using ICD-9 ontological features
Mihail Popescu and Mohammad Khalilia · 2011
Earlier work this paper cites.
Unachievable region in precision-recall space and its effect on empirical evaluation
Kendrick Boyd, Vítor Santos Costa, Jesse Davis, and C. David Page · 2012
Earlier work this paper cites.
KDIGO clinical practice guidelines for acute kidney injury
Arif Khwaja · 2012
Earlier work this paper cites.
Efficient estimation of word representations in vector space
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean · 2013
Earlier work this paper cites.
Stability
Bin Yu et al · 2013
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning · 2015
Earlier work this paper cites.
Intelligible Models for HealthCare
Rich Caruana, Yin Lou, Johannes Gehrke, Paul Koch, Marc Sturm, and Noemie Elhadad · 2015
Earlier work this paper cites.
Transparent reporting of a multivariable prediction model for individual prognosis or diagnosis (tripod): the tripod statement
Gary S Collins, Johannes B Reitsma, Douglas G Altman, and Karel GM Moons · 2015
Earlier work this paper cites.
OntoNotes: The 90% solution
Eduard Hovy, Mitchell Marcus, Martha Palmer, Lance Ramshaw, and Ralph Weischedel · 2015
Earlier work this paper cites.
Prediction policy problems
Jon Kleinberg, Jens Ludwig, Sendhil Mullainathan, and Ziad Obermeyer · 2015
Earlier work this paper cites.
UK biobank: an open access resource for identifying the causes of a wide range of complex diseases of middle and old age
Cathie Sudlow, John Gallacher, Naomi Allen, Valerie Beral, Paul Burton, John Danesh, Paul Downey, Paul Elliott, Jane Green, Martin Landray, Bette Liu, Paul Matthews, Giok Ong, Jill Pell, Alan Silman, Alan Young, Tim Sprosen, Tim Peakman, and Rory Collins · 2015
Earlier work this paper cites.
Modeling linkage disequilibrium increases accuracy of polygenic risk scores
Bjarni J Vilhjálmsson, Jian Yang, Hilary K Finucane, Alexander Gusev, Sara Lindström, Stephan Ripke, Giulio Genovese, Po-Ru Loh, Gaurav Bhatia, Ron Do, Tristan Hayeck, Hong-Hee Won, Schizophrenia Working Group of the Psychiatric Genomics Consortium, Discovery, Biology, and Risk of Inherited Variants in Breast Cancer (DRIVE) study, Sekar Kathiresan, Michele Pato, Carlos Pato, Rulla Tamimi, Eli Stahl, Noah Zaitlen, Bogdan Pasaniuc, Gillian Belbin, Eimear E Kenny, Mikkel H Schierup, Philip De Jager, Nikolaos A Patsopoulos, Steve McCarroll, Mark Daly, Shaun Purcell, Daniel Chasman, Benjamin Neale, Michael Goddard, Peter M Visscher, Peter Kraft, Nick Patterson, and Alkes L Price · 2015
Earlier work this paper cites.
Man is to computer programmer as woman is to homemaker? debiasing word embeddings
Tolga Bolukbasi, Kai-Wei Chang, James Y Zou, Venkatesh Saligrama, and Adam T Kalai · 2016
Earlier work this paper cites.
Deep learning , volume 1
Ian Goodfellow, Yoshua Bengio, Aaron Courville, and Yoshua Bengio · 2016
Earlier work this paper cites.
Development and validation of a deep learning algorithm for detection of diabetic retinopathy in retinal fundus photographs
Varun Gulshan, Lily Peng, Marc Coram, Martin C Stumpe, Derek Wu, Arunachalam Narayanaswamy, Subhashini Venugopalan, Kasumi Widner, Tom Madams, Jorge Cuadros, et al · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Causal inference by using invariant prediction: identification and confidence intervals
Jonas Peters, Peter Bühlmann, and Nicolai Meinshausen · 2016
Earlier work this paper cites.
Genomics is failing on diversity
Alice B Popejoy and Stephanie M Fullerton · 2016
Earlier work this paper cites.
Beyond prediction: Using big data for policy problems
Susan Athey · 2017
Earlier work this paper cites.
Yoshua Bengio · 2017
Earlier work this paper cites.
Semantics derived automatically from language corpora contain human-like biases
Aylin Caliskan, Joanna J Bryson, and Arvind Narayanan · 2017
Earlier work this paper cites.
GRAM: Graph-based Attention Model for Healthcare Representation Learning
Edward Choi, Mohammad Taha Bahadori, Le Song, Walter F Stewart, and Jimeng Sun · 2017
Earlier work this paper cites.
Capacity and trainability in recurrent neural networks
Jasmine Collins, Jascha Sohl-Dickstein, and David Sussillo · 2017
Earlier work this paper cites.
Dermatologist-level classification of skin cancer with deep neural networks
Andre Esteva, Brett Kuprel, Roberto A Novoa, Justin Ko, Susan M Swetter, Helen M Blau, and Sebastian Thrun · 2017
Earlier work this paper cites.
Counterfactual fairness
Matt J Kusner, Joshua Loftus, Chris Russell, and Ricardo Silva · 2017
Earlier work this paper cites.
Simple and scalable predictive uncertainty estimation using deep ensembles
Balaji Lakshminarayanan, Alexander Pritzel, and Charles Blundell · 2017
Earlier work this paper cites.
Human demographic history impacts genetic risk prediction across diverse populations
Alicia R Martin, Christopher R Gignoux, Raymond K Walters, Genevieve L Wojcik, Benjamin M Neale, Simon Gravel, Mark J Daly, Carlos D Bustamante, and Eimear E Kenny · 2017
Earlier work this paper cites.
Machine learning: an applied econometric approach
Sendhil Mullainathan and Jann Spiess · 2017
Earlier work this paper cites.
Right for the right reasons: training differentiable models by constraining their explanations
Andrew Slavin Ross, Michael C Hughes, and Finale Doshi-Velez · 2017
Cited alongside, same era.
Revisiting unreasonable effectiveness of data in deep learning era
Chen Sun, Abhinav Shrivastava, Saurabh Singh, and Abhinav Gupta · 2017
Cited alongside, same era.
Inception-v4, inception-resnet and the impact of residual connections on learning
Christian Szegedy, Sergey Ioffe, Vincent Vanhoucke, and Alexander A Alemi · 2017
Cited alongside, same era.
Development and validation of a deep learning system for diabetic retinopathy and related eye diseases using retinal images from multiethnic populations with diabetes
Daniel Shu Wei Ting, Carol Yim-Lui Cheung, Gilbert Lim, Gavin Siew Wei Tan, Nguyen D Quang, Alfred Gan, Haslina Hamzah, Renata Garcia-Franco, Ian Yew San Yeo, Shu Yen Lee, et al · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Albert: A lite bert for self-supervised learning of language representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut · 2019
Later among the works it cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 2019
Later among the works it cites.
Clinical use of current polygenic risk scores may exacerbate health disparities
Alicia R Martin, Masahiro Kanai, Yoichiro Kamatani, Yukinori Okada, Benjamin M Neale, and Mark J Daly · 2019
Later among the works it cites.
Predictive multiplicity in classification
Charles T Marx, Flavio du Pin Calmon, and Berk Ustun · 2019
Later among the works it cites.
The generalization error of random features regression: Precise asymptotics and double descent curve
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Machine learning and health care disparities in dermatology
Adewole S Adamson and Avery Smith · 2018
Cited alongside, same era.
Reconciling modern machine learning and the bias-variance trade-off
Mikhail Belkin, Daniel Hsu, Siyuan Ma, and Soumik Mandal · 2018
Cited alongside, same era.
Gender shades: Intersectional accuracy disparities in commercial gender classification
Joy Buolamwini and Timnit Gebru · 2018
Cited alongside, same era.
Loss surfaces, mode connectivity, and fast ensembling of dnns
Timur Garipov, Pavel Izmailov, Dmitrii Podoprikhin, Dmitry P Vetrov, and Andrew G Wilson · 2018
Cited alongside, same era.
Invariant causal prediction for nonlinear models
Christina Heinze-Deml, Jonas Peters, and Nicolai Meinshausen · 2018
Cited alongside, same era.
Universal language model fine-tuning for text classification
Jeremy Howard and Sebastian Ruder · 2018
Cited alongside, same era.
Averaging weights leads to wider optima and better generalization
Pavel Izmailov, Dmitrii Podoprikhin, Timur Garipov, Dmitry Vetrov, and Andrew Gordon Wilson · 2018
Cited alongside, same era.
Song Mei and Andrea Montanari · 2019
Later among the works it cites.
Acute kidney injury: prevention, detection and management
National Institute for Health and Care Excellence (NICE) · 2019
Later among the works it cites.
Dissecting racial bias in an algorithm used to manage the health of populations
Ziad Obermeyer, Brian Powers, Christine Vogeli, and Sendhil Mullainathan · 2019
Later among the works it cites.
Causality for machine learning
Bernhard Schölkopf · 2019
Later among the works it cites.
Lesia Semenova, Cynthia Rudin, and Ronald Parr · 2019
Later among the works it cites.
Can you trust your model’s uncertainty? evaluating predictive uncertainty under dataset shift
Jasper Snoek, Yaniv Ovadia, Emily Fertig, Balaji Lakshminarayanan, Sebastian Nowozin, D Sculley, Joshua Dillon, Jie Ren, and Zachary Nado · 2019
Later among the works it cites.
Learning robust global representations by penalizing local predictive power
Haohan Wang, Songwei Ge, Zachary Lipton, and Eric P Xing · 2019
Later among the works it cites.
Association between surgical skin markings in dermoscopic images and diagnostic performance of a deep learning convolutional neural network for melanoma recognition
Julia K Winkler, Christine Fink, Ferdinand Toberer, Alexander Enk, Teresa Deinlein, Rainer Hofmann-Wellenhof, Luc Thomas, Aimilios Lallas, Andreas Blum, Wilhelm Stolz, et al · 2019
Later among the works it cites.
Xlnet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R Salakhutdinov, and Quoc V Le · 2019
Later among the works it cites.
A fourier perspective on model robustness in computer vision
Dong Yin, Raphael Gontijo Lopes, Jon Shlens, Ekin Dogus Cubuk, and Justin Gilmer · 2019
Later among the works it cites.
Hellaswag: Can a machine really finish your sentence?
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi · 2019
Later among the works it cites.
Skin color in dermatology textbooks: An updated evaluation and analysis
Ademide Adelekun, Ginikanwa Onyekaba, and Jules B Lipoff · 2020
Closest in time.
Quantifying gender bias in different corpora
Marzieh Babaeianjelodar, Stephen Lorenz, Josh Gordon, Jeanna Matthews, and Evan Freitag · 2020
Closest in time.
A human-centered evaluation of a deep learning system deployed in clinics for the detection of diabetic retinopathy
Emma Beede, Elizabeth Baylor, Fred Hersch, Anna Iurchenko, Lauren Wilcox, Paisan Ruamviboonsuk, and Laura M Vardoulakis · 2020
Closest in time.
Climbing towards NLU: On meaning, form, and understanding in the age of data
Emily M. Bender and Alexander Koller · 2020
Closest in time.
On robustness and transferability of convolutional neural networks
Josip Djolonga, Jessica Yung, Michael Tschannen, Rob Romijnders, Lucas Beyer, Alexander Kolesnikov, Joan Puigcerver, Matthias Minderer, Alexander D’Amour, Dan Moldovan, et al · 2020
Closest in time.
Fine-tuning pretrained language models: Weight initializations, data orders, and early stopping
Jesse Dodge, Gabriel Ilharco, Roy Schwartz, Ali Farhadi, Hannaneh Hajishirzi, and Noah Smith · 2020
Closest in time.
Analyzing the role of model uncertainty for electronic health records
Michael W Dusenberry, Dustin Tran, Edward Choi, Jonas Kemp, Jeremy Nixon, Ghassen Jerfel, Katherine Heller, and Andrew M Dai · 2020
Closest in time.
Estimating the effects of non-pharmaceutical interventions on covid-19 in europe
Seth Flaxman, Swapnil Mishra, Axel Gandy, H Juliette T Unwin, Thomas A Mellan, Helen Coupland, Charles Whittaker, Harrison Zhu, Tresnia Berah, Jeffrey W Eaton, et al · 2020
Closest in time.
Linear mode connectivity and the lottery ticket hypothesis
Jonathan Frankle, Gintare Karolina Dziugaite, Daniel M Roy, and Michael Carbin · 2020
Closest in time.
The myth of generalisability in clinical research and machine learning in health care
Joseph Futoma, Morgan Simons, Trishan Panch, Finale Doshi-Velez, and Leo Anthony Celi · 2020
Closest in time.
Shortcut learning in deep neural networks
Robert Geirhos, Jörn-Henrik Jacobsen, Claudio Michaelis, Richard Zemel, Wieland Brendel, Matthias Bethge, and Felix A Wichmann · 2020
Closest in time.
The many faces of robustness: A critical analysis of out-of-distribution generalization
Dan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath, Frank Wang, Evan Dorundo, Rahul Desai, Tyler Zhu, Samyak Parajuli, Mike Guo, et al · 2020
Closest in time.
Sara Hooker · 2020
Closest in time.
Formalizing trust in artificial intelligence: Prerequisites, causes and goals of human trust in ai
Alon Jacovi, Ana Marasović, Tim Miller, and Yoav Goldberg · 2020
Closest in time.
Learning the difference that makes a difference with counterfactually-augmented data
Divyansh Kaushik, Eduard Hovy, and Zachary Lipton · 2020
Closest in time.
How can we accelerate progress towards human-like linguistic generalization?
Tal Linzen · 2020
Closest in time.
Stereoset: Measuring stereotypical bias in pretrained language models
Moin Nadeem, Anna Bethke, and Siva Reddy · 2020
Closest in time.
Deep double descent: Where bigger models and more data hurt
Preetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang, Boaz Barak, and Ilya Sutskever · 2020
Closest in time.
Hidden stratification causes clinically meaningful failures in machine learning for medical imaging
Luke Oakden-Rayner, Jared Dunnmon, Gustavo Carneiro, and Christopher Ré · 2020
Closest in time.
Doctor XAI An ontology-based approach to black-box sequential data classification explanations
Cecilia Panigutti, Alan Perotti, and Dino Pedreschi · 2020
Closest in time.
Understanding and mitigating the tradeoff between robustness and accuracy
Aditi Raghunathan, Sang Michael Xie, Fanny Yang, John Duchi, and Percy Liang · 2020
Closest in time.
Beyond accuracy: Behavioral testing of NLP models with CheckList
Marco Tulio Ribeiro, Tongshuang Wu, Carlos Guestrin, and Sameer Singh · 2020
Closest in time.
Guidelines for clinical trial protocols for interventions involving artificial intelligence: the spirit-ai extension
Samantha Cruz Rivera, Xiaoxuan Liu, An-Wen Chan, Alastair K Denniston, and Melanie J Calvert · 2020
Closest in time.
Measuring robustness to natural distribution shifts in image classification
Rohan Taori, Achal Dave, Vaishaal Shankar, Nicholas Carlini, Benjamin Recht, and Ludwig Schmidt · 2020
Closest in time.
High-frequency component helps explain the generalization of convolutional neural networks
Haohan Wang, Xindi Wu, Zeyi Huang, and Eric P Xing · 2020
Closest in time.
Measuring and reducing gendered correlations in pre-trained models
Kellie Webster, Xuezhi Wang, Ian Tenney, Alex Beutel, Emily Pitler, Ellie Pavlick, Jilin Chen, and Slav Petrov · 2020
Closest in time.
Hyperparameter ensembles for robustness and uncertainty quantification
Florian Wenzel, Jasper Snoek, Dustin Tran, and Rodolphe Jenatton · 2020
Closest in time.
Bayesian deep learning and a probabilistic perspective of generalization
Andrew Gordon Wilson and Pavel Izmailov · 2020
Closest in time.
The curse of performance instability in analysis datasets: Consequences, source, and suggestions
Xiang Zhou, Yixin Nie, Hao Tan, and Mohit Bansal · 2020
Closest in time.