Fetching the paper…
Reading the bibliography…
State-of-the-art pre-trained language models (PLMs) outperform other models when applied to the majority of language processing tasks.
Test-Time Training with Self-Supervision for Generalization under Distribution Shifts
Sun, Y.; Wang, X.; Liu, Z.; Miller, J.; Efros, A. A.; and Hardt, M. 2020 · 1909
Earlier work this paper cites.
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
Sanh, V.; Debut, L.; Chaumond, J.; and Wolf, T. 2019 · 1910
Earlier work this paper cites.
FixMatch: Simplifying Semi-Supervised Learning with Consistency and Confidence
Sohn, K.; Berthelot, D.; Li, C.-L.; Zhang, Z.; Carlini, N.; Cubuk, E. D.; Kurakin, A.; Zhang, H.; and Raffel, C. 2020 · 2001
Earlier work this paper cites.
Greedy Policy Search: A Simple Baseline for Learnable Test-Time Augmentation
Molchanov, D.; Lyzhov, A.; Molchanova, Y.; Ashukha, A.; and Vetrov, D. 2020 · 2002
Earlier work this paper cites.
Data Augmentation using Pre-trained Transformer Models
Kumar, V.; Choudhary, A.; and Cho, E. 2021 · 2003
Earlier work this paper cites.
Dataset shift in machine learning
Quinonero-Candela, J.; Sugiyama, M.; Schwaighofer, A.; and Lawrence, N. D. 2008 · 2008
Earlier work this paper cites.
Better Aggregation in Test-Time Augmentation
Shanmugam, D.; Blalock, D.; Balakrishnan, G.; and Guttag, J. 2021 · 2011
Earlier work this paper cites.
WILDS: A Benchmark of in-the-Wild Distribution Shifts
Koh, P. W.; Sagawa, S.; Marklund, H.; Xie, S. M.; Zhang, M.; Balsubramani, A.; Hu, W.; Yasunaga, M.; Phillips, R. L.; Beery, S.; Leskovec, J.; Kundaje, A.; Pierson, E.; Levine, S.; Finn, C.; and Liang, P. 2020 · 2012
Earlier work this paper cites.
Machine learning in non-stationary environments: Introduction to covariate shift adaptation
Sugiyama, M.; and Kawanabe, M. 2012 · 2012
Earlier work this paper cites.
PPDB: The paraphrase database
Ganitkevitch, J.; Van Durme, B.; and Callison-Burch, C. 2013 · 2013
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Mikolov, T.; Sutskever, I.; Chen, K.; Corrado, G. S.; and Dean, J. 2013 · 2013
Earlier work this paper cites.
Contextual Augmentation: Data Augmentation by Words with Paradigmatic Relations
Kobayashi, S. 2018 · 2018
Earlier work this paper cites.
Nuanced metrics for measuring unintended bias with real data for text classification
Borkan, D.; Dixon, L.; Sorensen, J.; Thain, N.; and Vasserman, L. 2019 · 2019
Cited alongside, same era.
AutoAugment: Learning Augmentation Policies from Data
Cubuk, E. D.; Zoph, B.; Mane, D.; Vasudevan, V.; and Le, Q. V. 2019 · 2019
Cited alongside, same era.
Domain Adaptation with BERT-based Domain Classification and Data Selection
Ma, X.; Xu, P.; Wang, Z.; Nallapati, R.; and Xiang, B. 2019 · 2019
Cited alongside, same era.
On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?
Bender, E. M.; Gebru, T.; McMillan-Major, A.; and Shmitchell, S. 2021 · 2021
Cited alongside, same era.
Quantifying and Alleviating Distribution Shifts in Foundation Models on Review Classification
Chawla, S.; Singh, N.; and Drori, I. 2021 · 2021
Cited alongside, same era.
A Robustly Optimized BERT Pre-training Approach with Post-training
Zhuang, L.; Wayne, L.; Ya, S.; and Jun, Z. 2021 · 2021
Later among the works it cites.
Debiased Self-Training for Semi-Supervised Learning
Chen, B.; Jiang, J.; Wang, X.; Wan, P.; Wang, J.; and Long, M. 2022 · 2022
Closest in time.
Continual Pre-Training Mitigates Forgetting in Language and Vision
Cossu, A.; Tuytelaars, T.; Carta, A.; Passaro, L. C.; Lomonaco, V.; and Bacciu, D. 2022 · 2022
Closest in time.
Lifelong Pretraining: Continually Adapting Language Models to Emerging Corpora
Jin, X.; Zhang, D.; Zhu, H.; Xiao, W.; Li, S.-W.; Wei, X.; Arnold, A.; and Ren, X. 2022 · 2022
Closest in time.
A Broad Study of Pre-training for Domain Generalization and Adaptation
Kim, D.; Wang, K.; Sclaroff, S.; and Saenko, K. 2022 · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A Survey of Data Augmentation Approaches for NLP
Feng, S. Y.; Gangal, V.; Wei, J.; Chandar, S.; Vosoughi, S.; Mitamura, T.; and Hovy, E. 2021 · 2021
Cited alongside, same era.
Mind the Gap: Assessing Temporal Generalization in Neural Language Models
Lazaridou, A.; Kuncoro, A.; Gribovskaya, E.; Agrawal, D.; Liska, A.; Terzi, T.; Gimenez, M.; d’Autume, C. d. M.; Kocisky, T.; Ruder, S.; Yogatama, D.; Cao, K.; Young, S.; and Blunsom, P. 2021 · 2021
Cited alongside, same era.
Unsupervised Paraphrasing with Pretrained Language Models
Niu, T.; Yavuz, S.; Zhou, Y.; Keskar, N. S.; Wang, H.; and Xiong, C. 2021 · 2021
Cited alongside, same era.
Con$^{2}$DA: Simplifying Semi-supervised Domain Adaptation by Learning Consistent and Contrastive Feature Representations
Pérez-Carrasco, M. I.; Protopapas, P.; and Cabrera-Vives, G. 2021a · 2021
Cited alongside, same era.
Con$^{2}$DA: Simplifying Semi-supervised Domain Adaptation by Learning Consistent and Contrastive Feature Representations
Pérez-Carrasco, M. I.; Protopapas, P.; and Cabrera-Vives, G. 2021b · 2021
Cited alongside, same era.
Domain-agnostic Test-time Adaptation by Prototypical Training with Auxiliary Data
Wu, Q.; Yue, X.; and Sangiovanni-Vincentelli, A. 2021 · 2021
Cited alongside, same era.
Closest in time.
Improved Text Classification via Test-Time Augmentation
Lu, H.; Shanmugam, D.; Suresh, H.; and Guttag, J. 2022 · 2022
Closest in time.
NLP Augmentation
Ma, E. 2019 · 2022
Closest in time.
Continual Active Adaptation to Evolving Distributional Shifts
Machireddy, A.; Krishnan, R.; Ahuja, N.; and Tickoo, O. 2022 · 2022
Closest in time.
Continual Learning with Deep Learning Methods in an Application-Oriented Context
Pfülb, B. 2022 · 2022
Closest in time.
A Fine-Grained Analysis on Distribution Shift
Wiles, O.; Gowal, S.; Stimberg, F.; Rebuffi, S.-A.; Ktena, I.; Dvijotham, K. D.; and Cemgil, A. T. 2022 · 2022
Closest in time.
MEMO: Test Time Robustness via Adaptation and Augmentation
Zhang, M.; Levine, S.; and Finn, C. 2022 · 2022
Closest in time.