Fetching the paper…
Reading the bibliography…
Toxic comment detection on social media has proven to be essential for content moderation.
ALBERT: A lite BERT for self-supervised learning of language representations
Lan, Z., M. Chen, S. Goodman, K. Gimpel, P. Sharma, and R. Soricut (2019) · 1909
Earlier work this paper cites.
Distilling the knowledge in a neural network
Hinton, G., O. Vinyals, and J. Dean (2014) · 2014
Earlier work this paper cites.
GloVe: Global vectors for word representation
Pennington, J., R. Socher, and C. Manning (2014) · 2014
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., X. Zhang, S. Ren, and J. Sun (2016) · 2016
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Sennrich, R., B. Haddow, and A. Birch (2016) · 2016
Earlier work this paper cites.
Focal loss for dense object detection
Lin, T., P. Goyal, R. Girshick, K. He, and P. Dollar (2017) · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. u. Kaiser, and I. Polosukhin (2017) · 2017
Earlier work this paper cites.
Nuanced metrics for measuring unintended bias with real data for text classification
Borkan, D., L. Dixon, J. Sorensen, N. Thain, and L. Vasserman (2019) · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., M.-W. Chang, K. Lee, and K. Toutanova (2019) · 2019
Cited alongside, same era.
NULI at SemEval-2019 task 6: Transfer learning for offensive language detection using bidirectional transformers
Liu, P., W. Li, and L. Zou (2019) · 2019
Cited alongside, same era.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Sanh, V., L. Debut, J. Chaumond, and T. Wolf (2019) · 2019
Cited alongside, same era.
Studying generalisability across abusive language detection datasets
Swamy, S. D., A. Jamatia, and B. Gambäck (2019) · 2019
Cited alongside, same era.
Using transfer-based language models to detect hateful and offensive language online
Isaksen, V. and B. Gambäck (2020) · 2020
Later among the works it cites.
BERTweet: A pre-trained language model for English tweets
Nguyen, D. Q., T. Vu, and A. Tuan Nguyen (2020) · 2020
Later among the works it cites.
HateBERT: Retraining BERT for abusive language detection in English
Caselli, T., V. Basile, J. Mitrović, and M. Granitzer (2021) · 2021
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby (2021) · 2021
Later among the works it cites.
Escaping the big data paradigm with compact transformers
Hassani, A., S. Walton, N. Shah, A. Abuduweili, J. Li, and H. Shi (2021) · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yang, Z., Z. Dai, Y. Yang, J. Carbonell, R. Salakhutdinov, and Q. V. Le (2019) · 2019
Cited alongside, same era.
Unsupervised cross-lingual representation learning at scale
Conneau, A., K. Khandelwal, N. Goyal, V. Chaudhary, G. Wenzek, F. Guzmán, E. Grave, M. Ott, L. Zettlemoyer, and V. Stoyanov (2020) · 2020
Cited alongside, same era.
Zhuang, L., L. Wayne, S. Ya, and Z. Jun (2021) · 2021
Later among the works it cites.