2020

COVID-Twitter-BERT: A Natural Language Processing Model to Analyse COVID-19 Content on Twitter

Müller, Martin, Salathé, Marcel, Kummervold, Per E

Understand

In this work, we release COVID-Twitter-BERT (CT-BERT), a transformer-based model, pretrained on a large corpus of Twitter messages on the topic of COVID-19.

  • Our model shows a 10-30% marginal improvement compared to its base model, BERT-Large, on five different classification datasets.
  • The largest improvements are on the target domain.
  • Pretrained transformer models, such as CT-BERT, are trained on a specific target domain and can be used for a wide variety of natural language processing tasks, including classification, question-answering and chatbots.

Built on

  • Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales

    Bo Pang and Lillian Lee · 2005

    Earlier work this paper cites.

  • Recursive deep models for semantic compositionality over a sentiment treebank

    Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Y Ng, and Christopher Potts · 2013

    Earlier work this paper cites.

  • Attention is all you need

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017

    Earlier work this paper cites.

  • spacy 2: Natural language understanding with bloom embeddings, convolutional neural networks and incremental parsing

    Matthew Honnibal and Ines Montani · 2017

    Earlier work this paper cites.

Similar

Then

  • Scibert: Pretrained contextualized embeddings for scientific text

    Original

    Iz Beltagy, Arman Cohan, and Kyle Lo · 2019

    Later among the works it cites.

  • Crowdbreaks: Tracking health trends using public social media data and crowdsourcing

    Martin M Müller and Marcel Salathé · 2019

    Later among the works it cites.

  • Semeval-2016 task 4: Sentiment analysis in twitter

    Original

    Preslav Nakov, Alan Ritter, Sara Rosenthal, Fabrizio Sebastiani, and Veselin Stoyanov · 2019

    Later among the works it cites.

  • Biobert: a pre-trained biomedical language representation model for biomedical text mining

    Jinhyuk Lee, Wonjin Yoon, Sungdong Kim, Donghyeon Kim, Sunkyu Kim, Chan Ho So, and Jaewoo Kang · 2020

    Closest in time.

Beyond the bibliography

alphaXiv searches the wider corpus for related work and actual follow-ups.

Open on alphaXiv

alphaXiv is searching for related work…