Fetching the paper…
Reading the bibliography…
For natural language processing `text-to-text' tasks, the prevailing approaches heavily rely on pretraining large self-supervised models on increasingly larger `task-external' data.
Learning and Evaluating General Linguistic Intelligence
Yogatama, D., de Masson d’Autume, C., Connor, J., Kociský, T., Chrzanowski, M., Kong, L., Lazaridou, A., Ling, W., Yu, L., Dyer, C., and Blunsom, P · 1901
Earlier work this paper cites.
Chang, W.-C., Yu, H.-F., Zhong, K., Yang, Y., and Dhillon, I · 1905
Earlier work this paper cites.
What Do Compressed Deep Neural Networks Forget?, 2020a
Hooker, S., Courville, A., Clark, G., Dauphin, Y., and Frome, A · 1911
Earlier work this paper cites.
Fine-Tuning Pretrained Language Models: Weight Initializations, Data Orders, and Early Stopping
Dodge, J., Ilharco, G., Schwartz, R., Farhadi, A., Hajishirzi, H., and Smith, N. A · 2002
Earlier work this paper cites.
A Survey on Contextual Embeddings
Liu, Q., Kusner, M. J., and Blunsom, P · 2003
Earlier work this paper cites.
The relationship between Precision-Recall and ROC curves
Davis, J. and Goadrich, M · 2006
Earlier work this paper cites.
Amnesic Probing: Behavioral Explanation with Amnesic Counterfactuals, 2020
Elazar, Y., Ravfogel, S., Jacovi, A., and Goldberg, Y · 2006
Earlier work this paper cites.
Hooker, S · 2009
Earlier work this paper cites.
It’s Not Just Size That Matters: Small Language Models Are Also Few-Shot Learners
Schick, T. and Schütze, H · 2009
Earlier work this paper cites.
Characterizing and Mitigating Bias in Compact Models
Hooker, S., Moorosi, N., Clark, G., Bengio, S., and Denton, E · 2010
Earlier work this paper cites.
A fast and simple algorithm for training neural probabilistic language models
Mnih, A. and Teh, Y. W · 2012
Earlier work this paper cites.
Tagging Your Tweets: A Probabilistic Modeling of Hashtag Annotation in Twitter
Ma, Z., Sun, A., Yuan, Q., and Cong, G · 2014
Earlier work this paper cites.
Enriching Word Vectors with Subword Information
Bojanowski, P., Grave, E., Joulin, A., and Mikolov, T · 2017
Earlier work this paper cites.
Deep Learning for Extreme Multi-label Text Classification
Liu, J., Chang, W., Wu, Y., and Yang, Y · 2017
Earlier work this paper cites.
Multi-task learning of pairwise sequence classification tasks over disparate label spaces
Augenstein, I., Ruder, S., and Søgaard, A · 2018
Earlier work this paper cites.
Learning from Imbalanced Data Sets
Fernández, A., García, S., Galar, M., Prati, R. C., Krawczyk, B., and Herrera, F · 2018
Cited alongside, same era.
Noise Contrastive Estimation and Negative Sampling for Conditional Models: Consistency and Statistical Efficiency
Ma, Z. and Collins, M · 2018
Cited alongside, same era.
Representation learning with contrastive predictive coding
van den Oord, A., Li, Y., and Vinyals, O · 2018
Cited alongside, same era.
Multi-Task Label Embedding for Text Classification
Zhang, H., Xiao, L., Chen, W., Wang, Y., and Jin, Y · 2018
Cited alongside, same era.
Multi-Task Label Embedding for Text Classification
Zhang, H., Xiao, L., Chen, W., Wang, Y., and Jin, Y · 2018
Cited alongside, same era.
MultiFC: A Real-World Multi-Domain Dataset for Evidence-Based Fact Checking of Claims
A Simple Framework for Contrastive Learning of Visual Representations
Chen, T., Kornblith, S., Norouzi, M., and Hinton, G. E · 2020
Closest in time.
Task-Aware Representation of Sentences for Generic Text Classification
Halder, K., Akbik, A., Krapac, J., and Vollgraf, R · 2020
Closest in time.
Multi-pretraining for Large-scale Text Classification
Kim, K.-M., Hyeon, B., Kim, Y., Park, J.-H., and Lee, S · 2020
Closest in time.
Understanding the Difficulty of Training Transformers
Liu, L., Liu, X., Gao, J., Chen, W., and Han, J · 2020
Closest in time.
Diversity and Inclusion Metrics in Subset Selection
Mitchell, M., Baker, D., Moorosi, N., Denton, E., Hutchinson, B., Hanna, A., Gebru, T., and Morgenstern, J · 2020
Closest in time.
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Augenstein, I., Lioma, C., Wang, D., Chaves Lima, L., Hansen, C., Hansen, C., and Simonsen, J. G · 2019
Cited alongside, same era.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Devlin, J., Chang, M., Lee, K., and Toutanova, K · 2019
Cited alongside, same era.
Learning deep representations by mutual information estimation and maximization
Hjelm, R. D., Fedorov, A., Lavoie-Marchildon, S., Grewal, K., Bachman, P., Trischler, A., and Bengio, Y · 2019
Cited alongside, same era.
Transferable Contrastive Network for Generalized Zero-Shot Learning
Jiang, H., Wang, R., Shan, S., and Chen, X · 2019
Cited alongside, same era.
Large-Scale Long-Tailed Recognition in an Open World
Liu, Z., Miao, Z., Zhan, X., Wang, J., Gong, B., and Yu, S. X · 2019
Cited alongside, same era.
GILE: A Generalized Input-Label Embedding for Text Classification
Pappas, N. and Henderson, J · 2019
Cited alongside, same era.
A theoretical analysis of contrastive unsupervised representation learning
Saunshi, N., Plevrakis, O., Arora, S., Khodak, M., and Khandeparkar, H · 2019
Cited alongside, same era.
Closest in time.
EffiCare: Better Prognostic Models via Resource-Efficient Health Embeddings
Şerbetci, O. N., Möller, S., Roller, R., and Rethmeier, N · 2020
Closest in time.
oLMpics-On What Language Model Pre-training Captures
Talmor, A., Elazar, Y., Goldberg, Y., and Berant, J · 2020
Closest in time.
To Pretrain or Not to Pretrain: Examining the Benefits of Pretrainng on Resource Rich Tasks
Wang, S., Khabsa, M., and Ma, H · 2020
Closest in time.
Disembodied Machine Learning: On the Illusion of Objectivity in NLP
Waseem, Z., Lulz, S., Bingel, J., and Augenstein, I · 2020
Closest in time.
On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?
Bender, E. M., Gebru, T., McMillan-Major, A., and Shmitchell, S · 2021
Closest in time.
Moving beyond “algorithmic bias is a data problem”
Hooker, S · 2021
Closest in time.
Learning Transferable Visual Models From Natural Language Supervision, 2021
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., Krueger, G., and Sutskever, I · 2021
Closest in time.
A Primer on Contrastive Pretraining in Language Processing: Methods, Lessons Learned and Perspectives, 2021
Rethmeier, N. and Augenstein, I · 2021
Closest in time.