Fetching the paper…
Reading the bibliography…
A recent trend in artificial intelligence is the use of pretrained models for language and vision tasks, which have achieved extraordinary performance but also puzzling failures.
Statistical theory of reliability and life testing: probability models
Richard E Barlow and Frank Proschan · 1975
Earlier work this paper cites.
The well-calibrated Bayesian
A Philip Dawid · 1982
Earlier work this paper cites.
Statistical methods for forecasting , volume 179
Bovas Abraham and Johannes Ledolter · 1983
Earlier work this paper cites.
Neural network ensembles
L.K. Hansen and P. Salamon · 1990
Earlier work this paper cites.
Active learning with statistical models
David A Cohn, Zoubin Ghahramani, and Michael I Jordan · 1996
Earlier work this paper cites.
Lifelong learning algorithms
Sebastian Thrun · 1998
Earlier work this paper cites.
Active learning for natural language parsing and information extraction
Cynthia A Thompson, Mary Elaine Califf, and Raymond J Mooney · 1999
Earlier work this paper cites.
Active hidden markov models for information extraction
Tobias Scheffer, Christian Decomain, and Stefan Wrobel · 2001
Earlier work this paper cites.
Active learning for automatic speech recognition
Dilek Hakkani-Tür, Giuseppe Riccardi, and Allen Gorin · 2002
Earlier work this paper cites.
Learning generative visual models from few training examples: An incremental Bayesian approach tested on 101 object categories
Li Fei-Fei, Rob Fergus, and Pietro Perona · 2004
Earlier work this paper cites.
Active learning: Theory and applications to automatic speech recognition
Giuseppe Riccardi and Dilek Hakkani-Tur · 2005
Earlier work this paper cites.
One-shot learning of object categories
Fei-Fei Li, Rob Fergus, and Pietro Perona · 2006
Earlier work this paper cites.
Margin-based active learning for structured output spaces
Dan Roth and Kevin Small · 2006
Earlier work this paper cites.
Strictly Proper Scoring Rules, Prediction, and Estimation
Tilmann Gneiting and Adrian E Raftery · 2007
Earlier work this paper cites.
Active policy learning for robot planning and exploration under uncertainty
Ruben Martinez-Cantin, Nando de Freitas, Arnaud Doucet, and José A Castellanos · 2007
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Decision theory: principles and approaches
G. Parmigiani and Lurdes Inoue · 2009
Earlier work this paper cites.
Active learning literature survey
Burr Settles · 2009
Earlier work this paper cites.
On the foundations of noise-free selective classification
Ran El-Yaniv and Yair Wiener · 2010
Earlier work this paper cites.
Caltech-ucsd birds 200
Peter Welinder, Steve Branson, Takeshi Mita, Catherine Wah, Florian Schroff, Serge Belongie, and Pietro Perona · 2010
Earlier work this paper cites.
Combining ensembles and data augmentation can harm your calibration
Yeming Wen, Ghassen Jerfel, Rafael Muller, Michael W Dusenberry, Jasper Snoek, Balaji Lakshminarayanan, and Dustin Tran · 2010
Earlier work this paper cites.
Bag-of-visual-words and spatial extensions for land-use classification
Yi Yang and Shawn Newsam · 2010
Earlier work this paper cites.
Practical reliability engineering
Patrick O’Connor and Andre Kleyner · 2012
Earlier work this paper cites.
Cats and dogs
O. M. Parkhi, A. Vedaldi, A. Zisserman, and C. V. Jawahar · 2012
Earlier work this paper cites.
Machine learning in non-stationary environments: Introduction to covariate shift adaptation
Masashi Sugiyama and Motoaki Kawanabe · 2012
Earlier work this paper cites.
Estimating or propagating gradients through stochastic neurons for conditional computation
Yoshua Bengio, Nicholas Léonard, and Aaron Courville · 2013
Earlier work this paper cites.
3d object representations for fine-grained categorization
Jonathan Krause, Michael Stark, Jia Deng, and Li Fei-Fei · 2013
Earlier work this paper cites.
Describing textures in the wild
M. Cimpoi, S. Maji, I. Kokkinos, S. Mohamed, and A. Vedaldi · 2014
Earlier work this paper cites.
Conditional computation in neural networks for faster models
Emmanuel Bengio, Pierre-Luc Bacon, Joelle Pineau, and Doina Precup · 2015
Earlier work this paper cites.
Weight Uncertainty in Neural Networks
Charles Blundell, Julien Cornebise, Koray Kavukcuoglu, and Daan Wierstra · 2015
Earlier work this paper cites.
Probabilistic machine learning and artificial intelligence
Zoubin Ghahramani · 2015
Earlier work this paper cites.
Obtaining well calibrated probabilities using Bayesian binning
Mahdi Pakdaman Naeini, Gregory Cooper, and Milos Hauskrecht · 2015
Earlier work this paper cites.
Research priorities for robust and beneficial artificial intelligence
Stuart Russell, Daniel Dewey, and Max Tegmark · 2015
Earlier work this paper cites.
Concrete problems in AI safety
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané · 2016
Earlier work this paper cites.
Dropout As a Bayesian Approximation: Representing Model Uncertainty in Deep Learning
Yarin Gal and Zoubin Ghahramani · 2016
Earlier work this paper cites.
Multi-class texture analysis in colorectal cancer histology
Jakob Nikolas Kather, Cleo-Aron Weis, Francesco Bianconi, Susanne M Melchers, Lothar R Schad, Timo Gaiser, Alexander Marx, and Frank Gerrit Z”ollner · 2016
Earlier work this paper cites.
Deep Bayesian active learning with image data
Yarin Gal, Riashat Islam, and Zoubin Ghahramani · 2017
Earlier work this paper cites.
A Baseline for Detecting Misclassified and Out-of-Distribution Examples in Neural Networks
Dan Hendrycks and Kevin Gimpel · 2017
Cited alongside, same era.
What uncertainties do we need in Bayesian deep learning for computer vision?
Alex Kendall and Yarin Gal · 2017
Cited alongside, same era.
Simple and scalable predictive uncertainty estimation using deep ensembles
Balaji Lakshminarayanan, Alexander Pritzel, and Charles Blundell · 2017
Cited alongside, same era.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean · 2017
Cited alongside, same era.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel R Bowman · 2017
Cited alongside, same era.
Measuring robustness to natural distribution shifts in image classification
Rohan Taori, Achal Dave, Vaishaal Shankar, Nicholas Carlini, Benjamin Recht, and Ludwig Schmidt · 2020
Later among the works it cites.
Practical uncertainty estimation and out-of-distribution robustness in deep learning
Dustin Tran, Jasper Snoek, and Balaji Lakshminarayanan · 2020
Later among the works it cites.
An empirical study on robustness to spurious correlations using pre-trained language models
Lifu Tu, Garima Lalwani, Spandana Gella, and He He · 2020
Later among the works it cites.
Uncertainty estimation using a single deep deterministic neural network
Joost Van Amersfoort, Lewis Smith, Yee Whye Teh, and Yarin Gal · 2020
Later among the works it cites.
Sparse MoEs meet efficient ensembles
James Urquhart Allingham, Florian Wenzel, Zelda E Mariet, Basil Mustafa, Joan Puigcerver, Neil Houlsby, Ghassen Jerfel, Vincent Fortuin, Balaji Lakshminarayanan, Jasper Snoek, Dustin Tran, Carlos Riquelme Ruiz, and Rodolphe Jenatton · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ex machina: Personal attacks seen at scale
Ellery Wulczyn, Nithum Thain, and Lucas Dixon · 2017
Cited alongside, same era.
The power of ensembles for active learning in image classification
William H Beluch, Tim Genewein, Andreas Nürnberger, and Jan M Köhler · 2018
Cited alongside, same era.
Tom Everitt, Gary Lea, and Marcus Hutter · 2018
Cited alongside, same era.
A simple unified framework for detecting out-of-distribution samples and adversarial attacks
Kimin Lee, Kibok Lee, Honglak Lee, and Jinwoo Shin · 2018
Cited alongside, same era.
Deep Bayesian active learning for natural language processing: Results of a large-scale empirical study
Aditya Siddhant and Zachary C Lipton · 2018
Cited alongside, same era.
Simple, distributed, and accelerated probabilistic programming
Dustin Tran, Matthew W Hoffman, Dave Moore, Christopher Suter, Srinivas Vasudevan, and Alexey Radul · 2018
Cited alongside, same era.
Active model learning and diverse action sampling for task and motion planning
Zi Wang, Caelan Reed Garrett, Leslie Pack Kaelbling, and Tomás Lozano-Pérez · 2018
Cited alongside, same era.
Later among the works it cites.
Filling gaps in trustworthy development of ai
Shahar Avin, Haydn Belfield, Miles Brundage, Gretchen Krueger, Jasmine Wang, Adrian Weller, Markus Anderljung, Igor Krawczuk, David Krueger, Jonathan Lebensold, et al · 2021
Later among the works it cites.
Benchmarking Bayesian deep learning on diabetic retinopathy detection tasks
Neil Band, Tim G. J. Rudner, Qixuan Feng, Angelos Filos, Zachary Nado, Michael W Dusenberry, Ghassen Jerfel, Dustin Tran, and Yarin Gal · 2021
Later among the works it cites.
On the opportunities and risks of foundation models
Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al · 2021
Later among the works it cites.
Batch active learning at scale, 2021
Gui Citovsky, Giulia DeSalvo, Claudio Gentile, Lazaros Karydas, Anand Rajagopalan, Afshin Rostamizadeh, and Sanjiv Kumar · 2021
Later among the works it cites.
Correlated input-dependent label noise in large-scale image classification
Mark Collier, Basil Mustafa, Efi Kokiopoulou, Rodolphe Jenatton, and Jesse Berent · 2021
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2021
Later among the works it cites.
Active learning at the imagenet scale
Zeyad Ali Sami Emam, Hong-Min Chu, Ping-Yeh Chiang, Wojciech Czaja, Richard Leapman, Micah Goldblum, and Tom Goldstein · 2021
Later among the works it cites.
Exploring the limits of out-of-distribution detection
Stanislav Fort, Jie Ren, and Balaji Lakshminarayanan · 2021
Later among the works it cites.
Measuring and improving model-moderator collaboration using uncertainty estimation
Ian D Kivlichan, Zi Lin, Jeremiah Liu, and Lucy Vasserman · 2021
Later among the works it cites.
GShard: Scaling giant models with conditional computation and automatic sharding
Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer, and Zhifeng Chen · 2021
Later among the works it cites.
Benchmarking natural language understanding services for building conversational agents
Xingkun Liu, Arash Eshghi, Pawel Swietojanski, and Verena Rieser · 2021
Later among the works it cites.
Cross-token modeling with conditional computation, 2021
Yuxuan Lou, Fuzhao Xue, Zangwei Zheng, and Yang You · 2021
Later among the works it cites.
Uncertainty estimation in autoregressive structured prediction
Andrey Malinin and Mark Gales · 2021
Later among the works it cites.
Revisiting the calibration of modern neural networks
Matthias Minderer, Josip Djolonga, Rob Romijnders, Frances Hubis, Xiaohua Zhai, Neil Houlsby, Dustin Tran, and Mario Lucic · 2021
Later among the works it cites.
Uncertainty baselines: Benchmarks for uncertainty & robustness in deep learning
Zachary Nado, Neil Band, Mark Collier, Josip Djolonga, Michael W Dusenberry, Sebastian Farquhar, Qixuan Feng, Angelos Filos, Marton Havasi, Rodolphe Jenatton, et al · 2021
Later among the works it cites.
Combined scaling for zero-shot transfer learning
Hieu Pham, Zihang Dai, Golnaz Ghiasi, Hanxiao Liu, Adams Wei Yu, Minh-Thang Luong, Mingxing Tan, and Quoc V Le · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever · 2021
Later among the works it cites.
A simple fix to mahalanobis distance for improving near-ood detection
Jie Ren, Stanislav Fort, Jeremiah Liu, Abhijit Guha Roy, Shreyas Padhy, and Balaji Lakshminarayanan · 2021
Later among the works it cites.
Scaling vision with sparse mixture of experts
Carlos Riquelme, Joan Puigcerver, Basil Mustafa, Maxim Neumann, Rodolphe Jenatton, André Susano Pinto, Daniel Keysers, and Neil Houlsby · 2021
Later among the works it cites.
Tractable Function-Space Variational Inference in Bayesian Neural Networks
Tim G. J. Rudner, Zonghao Chen, Yee Whye Teh, and Yarin Gal · 2021
Later among the works it cites.
Do image classifiers generalize across time?
Vaishaal Shankar, Achal Dave, Rebecca Roelofs, Deva Ramanan, Benjamin Recht, and Ludwig Schmidt · 2021
Later among the works it cites.
X Zhai, A Kolesnikov, N Houlsby, and L Beyer · 2021
Later among the works it cites.
Jian-Guo Zhang, Kazuma Hashimoto, Yao Wan, Ye Liu, Caiming Xiong, and Philip S Yu · 2021
Later among the works it cites.
Palm: Scaling language modeling with pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al · 2022
Closest in time.
Head2toe: Utilizing intermediate representations for better transfer learning
Utku Evci, Vincent Dumoulin, Hugo Larochelle, and Michael C Mozer · 2022
Closest in time.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
William Fedus, Barret Zoph, and Noam Shazeer · 2022
Closest in time.
A simple approach to improve single-model deep uncertainty via distance-awareness
Jeremiah Zhe Liu, Shreyas Padhy, Jie Ren, Zi Lin, Yeming Wen, Ghassen Jerfel, Zack Nado, Jasper Snoek, Dustin Tran, and Balaji Lakshminarayanan · 2022
Closest in time.
Continual Learning via Sequential Function-Space Variational Inference
Tim G. J. Rudner, Freddie Bickford Smith, Qixuan Feng, Yee Whye Teh, and Yarin Gal · 2022
Closest in time.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, et al · 2022
Closest in time.
Active learning helps pretrained models learn the intended task
Alex Tamkin, Dat Nguyen, Salil Deshpande, Jesse Mu, and Noah Goodman · 2022
Closest in time.
Lamda: Language models for dialog applications
Romal Thoppilan, Daniel De Freitas, Jamie Hall, Noam Shazeer, Apoorv Kulshreshtha, Heng-Tze Cheng, Alicia Jin, Taylor Bos, Leslie Baker, Yu Du, et al · 2022
Closest in time.
Using self-supervised pretext tasks for active learning, 2022
John Seon Keun Yi, Minseok Seo, Jongchan Park, and Dong-Geol Choi · 2022
Closest in time.
What do we mean by generalization in federated learning?
Honglin Yuan, Warren Morningstar, Lin Ning, and Karan Singhal · 2022
Closest in time.