Fetching the paper…
Reading the bibliography…
Despite achieving state-of-the-art results in nearly all Natural Language Processing applications, fine-tuning Transformer-based language models still requires a significant amount of labeled data to work.
“A mathematical theory of communication”
Claude. Shannon · 1948
Earlier work this paper cites.
“Query by Committee”, COLT ’92
H.. Seung, M. Opper and H. Sompolinsky · 1992
Earlier work this paper cites.
“A Sequential Algorithm for Training Text Classifiers”
David. Lewis and William. Gale · 1994
Earlier work this paper cites.
“Employing EM and pool-based active learning for text classification”
Andrew McCallumzy and Kamal Nigamy · 1998
Earlier work this paper cites.
“Active Hidden Markov Models for Information Extraction”
Tobias Scheffer, Christian Decomain and Stefan Wrobel · 2001
Earlier work this paper cites.
“Active Learning by Querying Informative and Representative Examples”
Sheng-jun Huang, Rong Jin and Zhi-Hua Zhou · 2010
Earlier work this paper cites.
“Weight uncertainty in neural network”
Charles Blundell, Julien Cornebise, Koray Kavukcuoglu and Daan Wierstra · 2015
Earlier work this paper cites.
“Deep learning”
Yann LeCun, Yoshua Bengio and Geoffrey Hinton · 2015
Earlier work this paper cites.
“Dropout as a bayesian approximation: Representing model uncertainty in deep learning”
Yarin Gal and Zoubin Ghahramani · 2016
Earlier work this paper cites.
“Rethinking the inception architecture for computer vision”
Christian Szegedy et al · 2016
Earlier work this paper cites.
“Simple and scalable predictive uncertainty estimation using deep ensembles”
Balaji Lakshminarayanan, Alexander Pritzel and Charles Blundell · 2017
Earlier work this paper cites.
“Snorkel: Rapid training data creation with weak supervision”
Alexander Ratner et al · 2017
Earlier work this paper cites.
“Attention is all you need”
Ashish Vaswani et al · 2017
Cited alongside, same era.
“Active discriminative text representation learning”
Ye Zhang, Matthew Lease and Byron Wallace · 2017
Cited alongside, same era.
“Deep learning using rectified linear units (relu)”
Abien Agarap · 2018
Cited alongside, same era.
“To Trust Or Not To Trust A Classifier”
Heinrich Jiang, Been Kim, Melody Guan and Maya Gupta · 2018
Cited alongside, same era.
“Inhibited softmax for uncertainty estimation in neural networks”
Marcin Możejko, Mateusz Susik and Rafał Karczewski · 2018
Cited alongside, same era.
“Mix-n-match: Ensemble and compositional methods for uncertainty calibration in deep learning”
Jize Zhang, Bhavya Kailkhura and T-Jin Han · 2020
Later among the works it cites.
“A survey of uncertainty in deep neural networks”
Jakob Gawlikowski et al · 2021
Later among the works it cites.
“Mind Your Outliers! Investigating the Negative Impact of Outliers on Active Learning for Visual Question Answering”
Siddharth Karamcheti, Ranjay Krishna, Li Fei-Fei and Christopher Manning · 2021
Later among the works it cites.
“Datasets: A community library for natural language processing”
Quentin Lhoest et al · 2021
Later among the works it cites.
“Understanding softmax confidence and uncertainty”
Tim Pearce, Alexandra Brintrup and Jun Zhu · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Murat Sensoy, Lance Kaplan and Melih Kandemir · 2018
Cited alongside, same era.
“BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding”
Jacob Devlin, Ming-Wei Chang, Kenton Lee and Kristina Toutanova · 2019
Cited alongside, same era.
“Why relu networks yield high-confidence predictions far away from the training data and how to mitigate the problem”
Matthias Hein, Maksym Andriushchenko and Julian Bitterwolf · 2019
Cited alongside, same era.
“BatchBALD: Efficient and Diverse Batch Acquisition for Deep Bayesian Active Learning”
A. Kirsch, J. v. Amersfoort and Y. Gal · 2019
Cited alongside, same era.
“Roberta: A robustly optimized bert pretraining approach”
Yinhan Liu et al · 2019
Cited alongside, same era.
“Huggingface’s transformers: State-of-the-art natural language processing”
Thomas Wolf et al · 2019
Cited alongside, same era.
Later among the works it cites.
“Small-text: Active Learning for Text Classification in Python”
Christopher Schröder, Lydia Müller, Andreas Niekler and Martin Potthast · 2021
Later among the works it cites.
“Limitations of Active Learning With Deep Transformer Language Models”, 2022
Mike D’Arcy and Doug Downey · 2022
Closest in time.
“Uncertainty Estimation for Language Reward Models”
Adam Gleave and Geoffrey Irving · 2022
Closest in time.
“Revisiting Uncertainty-based Query Strategies for Active Learning with Transformers”
Christopher Schröder, Andreas Niekler and Martin Potthast · 2022
Closest in time.
“BayesFormer: Transformer with Uncertainty Estimation”
Karthik Sankararaman, Sinong Wang and Han Fang · 2022
Closest in time.
Michael Weiss and Paolo Tonella · 2022
Closest in time.