Fetching the paper…
Reading the bibliography…
In recent years, pre-trained language models (PLMs) achieve the best performance on a wide range of natural language processing (NLP) tasks.
Roberta: A robustly optimized BERT pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
The unified medical language system (umls): Integrating biomedical terminology
Olivier Bodenreider. 2004 · 2004
Earlier work this paper cites.
The OPUS corpus - parallel and free: http://logos.uio.no/opus
Jörg Tiedemann and Lars Nygaard. 2004 · 2004
Earlier work this paper cites.
LEGAL-BERT: the muppets straight out of law school
Ilias Chalkidis, Manos Fergadiotis, Prodromos Malakasiotis, Nikolaos Aletras, and Ion Androutsopoulos. 2020 · 2010
Earlier work this paper cites.
The quaero french medical corpus : A ressource for medical entity recognition and normalization
Aurélie Névéol, Cyril Grouin, Jérémy Leixa, Sophie Rosset, and Pierre Zweigenbaum. 2014 · 2014
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Earlier work this paper cites.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V. Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, Jeff Klingner, Apurva Shah, Melvin Johnson, Xiaobing Liu, Łukasz Kaiser, Stephan Gouws, Yoshikiyo Kato, Taku Kudo, Hideto Kazawa, Keith Stevens, George Kurian, Nishant Patil, Wei Wang, Cliff Young, Jason Smith, Jason Riesa, Alex Rudnick, Oriol Vinyals, Greg Corrado, Macduff Hughes, and Jeffrey Dean. 2016 · 2016
Earlier work this paper cites.
Paulius Micikevicius, Sharan Narang, Jonah Alben, Gregory F. Diamos, Erich Elsen, David García, Boris Ginsburg, Michael Houston, Oleksii Kuchaiev, Ganesh Venkatesh, and Hao Wu. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Earlier work this paper cites.
CAS: French Corpus with Clinical Cases
Natalia Grabar, Vincent Claveau, and Clément Dalloux. 2018 · 2018
Earlier work this paper cites.
Universal Language Model Fine-tuning for Text Classification
Jeremy Howard and Sebastian Ruder. 2018 · 2018
Earlier work this paper cites.
Taku Kudo and John Richardson. 2018 · 2018
Cited alongside, same era.
Scibert: Pretrained language model for scientific text
Iz Beltagy, Kyle Lo, and Arman Cohan. 2019 · 2019
Cited alongside, same era.
Clinicalbert: Modeling clinical notes and predicting hospital readmission
Kexin Huang, Jaan Altosaar, and Rajesh Ranganath. 2019 · 2019
Cited alongside, same era.
BioBERT: a pre-trained biomedical language representation model for biomedical text mining
Jinhyuk Lee, Wonjin Yoon, Sungdong Kim, Donghyeon Kim, Sunkyu Kim, Chan Ho So, and Jaewoo Kang. 2019 · 2019
Cited alongside, same era.
Transfer learning in biomedical natural language processing: An evaluation of bert and elmo on ten benchmarking datasets
Yifan Peng, Shankai Yan, and Zhiyong Lu. 2019 · 2019
Cited alongside, same era.
Supervised learning for the detection of negation and of its scope in French and Brazilian Portuguese biomedical corpora
Clément Dalloux, Vincent Claveau, Natalia Grabar, Lucas Emanuel Silva Oliveira, Claudia Maria Cabral Moro, Yohan Bonescki Gumiel, and Deborah Ribeiro Carvalho. 2021 · 2021
Later among the works it cites.
Domain-Specific Language Model Pretraining for Biomedical Natural Language Processing
Yu Gu, Robert Tinn, Hao Cheng, Michael Lucas, Naoto Usuyama, Xiaodong Liu, Tristan Naumann, Jianfeng Gao, and Hoifung Poon. 2021 · 2021
Later among the works it cites.
Societal biases in language generation: Progress and challenges
Emily Sheng, Kai-Wei Chang, Premkumar Natarajan, and Nanyun Peng. 2021 · 2021
Later among the works it cites.
Development of a language model for medical domain
Manjil Shrestha. 2021 · 2021
Later among the works it cites.
Travelbert: Pre-training language model incorporating domain-specific heterogeneous knowledge into a unified representation
Hongyin Zhu, Hao Peng, Zhiheng Lyu, Lei Hou, Juanzi Li, and Jinghui Xiao. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Huggingface’s transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush. 2019 · 2019
Cited alongside, same era.
CamemBERT: a Tasty French Language Model
Louis Martin, Benjamin Muller, Pedro Javier Ortiz Suá rez, Yoann Dupont, Laurent Romary, Éric de la Clergerie, Djamé Seddah, and Benoît Sagot. 2020 · 2020
Cited alongside, same era.
BioBERTpt - a Portuguese neural language model for clinical named entity recognition
Elisa Terumi Rubel Schneider, João Vitor Andrioli de Souza, Julien Knafou, Lucas Emanuel Silva e Oliveira, Jenny Copara, Yohan Bonescki Gumiel, Lucas Ferro Antunes de Oliveira, Emerson Cabrera Paraiso, Douglas Teodoro, and Cláudia Maria Cabral Moro Barra. 2020 · 2020
Cited alongside, same era.
Finbert: A pretrained language model for financial communications
Yi Yang, Mark Christopher Siy UY, and Allen Huang. 2020 · 2020
Cited alongside, same era.
When do you need billions of words of pretraining data?
Yian Zhang, Alex Warstadt, Haau-Sing Li, and Samuel R. Bowman. 2020 · 2020
Cited alongside, same era.
On the dangers of stochastic parrots: Can language models be too big?
Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021 · 2021
Cited alongside, same era.
Biomedical and clinical language models for spanish: On the benefits of domain-specific pretraining in a mid-resource scenario
Casimiro Pio Carrino, Jordi Armengol-Estapé, Asier Gutiérrez-Fandiño, Joan Llop-Palao, Marc Pàmies, Aitor Gonzalez-Agirre, and Marta Villegas. 2021 · 2021
Cited alongside, same era.
Re-train or Train from Scratch? Comparing Pre-training Strategies of BERT in the Medical Domain
Hicham El Boukkouri, Olivier Ferret, Thomas Lavergne, and Pierre Zweigenbaum. 2022 · 2022
Later among the works it cites.
Massive: A 1m-example multilingual natural language understanding dataset with 51 typologically-diverse languages
Jack FitzGerald, Christopher Hench, Charith Peris, Scott Mackie, Kay Rottmann, Ana Sanchez, Aaron Nash, Liam Urbach, Vishesh Kakarala, Richa Singh, Swetha Ranganath, Laurie Crist, Misha Britan, Wouter Leeuwis, Gokhan Tur, and Prem Natarajan. 2022 · 2022
Later among the works it cites.
FrenchMedMCQA: A French Multiple-Choice Question Answering Dataset for Medical domain
Yanis Labrak, Adrien Bazoge, Richard Dufour, Béatrice Daille, Pierre-Antoine Gourraud, Emmanuel Morin, and Mickael Rouvier. 2022 · 2022
Later among the works it cites.
Estimating the carbon footprint of bloom, a 176b parameter language model
Alexandra Sasha Luccioni, Sylvain Viguier, and Anne-Laure Ligozat. 2022 · 2022
Later among the works it cites.
Bioberturk: Exploring turkish biomedical language model development strategies in low resource setting
Hazal Türkmen, Oğuz Dikenelli, Cenk Eraslan, and Mehmet Callı. 2022 · 2022
Later among the works it cites.
Downstream task performance of BERT models pre-trained using automatically de-identified clinical data
Thomas Vakili, Anastasios Lamproudis, Aron Henriksson, and Hercules Dalianis. 2022 · 2022
Later among the works it cites.
A large language model for electronic health records
Xi Yang, Aokun Chen, Nima PourNejatian, Hoo Chang Shin, Kaleb E. Smith, Christopher Parisien, Colin Compas, Cheryl Martin, Anthony B. Costa, Mona G. Flores, Ying Zhang, Tanja Magoc, Christopher A. Harle, Gloria Lipori, Duane A. Mitchell, William R. Hogan, Elizabeth A. Shenkman, Jiang Bian, and Yonghui Wu. 2022 · 2022
Later among the works it cites.