Fetching the paper…
Reading the bibliography…
Contextual embedding-based language models trained on large data sets, such as BERT and RoBERTa, provide strong performance across a wide range of tasks and are ubiquitous in modern NLP.
Clinicalbert: Modeling clinical notes and predicting hospital readmission
Kexin Huang, Jaan Altosaar, and Rajesh Ranganath. 2019 · 1904
Earlier work this paper cites.
Defending against neural fake news
Rowan Zellers, Ari Holtzman, Hannah Rashkin, Yonatan Bisk, Ali Farhadi, Franziska Roesner, and Yejin Choi. 2019 · 1905
Earlier work this paper cites.
Xlnet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime G. Carbonell, Ruslan Salakhutdinov, and Quoc V. Le. 2019 · 1906
Earlier work this paper cites.
Roberta: A robustly optimized BERT pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Roy Schwartz, Jesse Dodge, Noah A. Smith, and Oren Etzioni. 2019 · 1907
Earlier work this paper cites.
Social differentiation in the use of english vocabulary: some analyses of the conversational component of the british national corpus
Paul Rayson, G. Leech, and Mary Hodges. 1997 · 1997
Earlier work this paper cites.
Longformer: The long-document transformer
Iz Beltagy, Matthew E. Peters, and Arman Cohan. 2020 · 2004
Earlier work this paper cites.
Fightin’ words: Lexical feature selection and evaluation for identifying the content of political conflict
B. L. Monroe, Michael Colaresi, and K. Quinn. 2008 · 2008
Earlier work this paper cites.
Detecting multiple facets of an event using graph-based unsupervised methods
Pradeep Muthukrishnan, Joshua Gerrish, and Dragomir R. Radev. 2008 · 2008
Earlier work this paper cites.
Probing pretrained language models for lexical semantics
Ivan Vulić, E. Ponti, Robert Litschko, Goran Glavas, and Anna Korhonen. 2020 · 2010
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts. 2011 · 2011
Earlier work this paper cites.
Gensim–python framework for vector space modelling
Radim Rehurek and Petr Sojka. 2011 · 2011
Earlier work this paper cites.
Japanese and korean voice search
Mike Schuster and Kaisuke Nakajima. 2012 · 2012
Earlier work this paper cites.
Personality, gender, and age in the language of social media: The open-vocabulary approach
H. Andrew Schwartz, Johannes C. Eichstaedt, Margaret L. Kern, Lukasz Dziurzynski, Stephanie M. Ramones, Megha Agrawal, Achal Shah, Michal Kosinski, David Stillwell, Martin E. P. Seligman, and Lyle H. Ungar. 2013 · 2013
Earlier work this paper cites.
Ups and downs: Modeling the visual evolution of fashion trends with one-class collaborative filtering
Ruining He and Julian McAuley. 2016 · 2016
Cited alongside, same era.
ChemProt-3.0: a global chemical biology diseases mapping
Jens Kringelum, Sonny Kim Kjaerulff, Søren Brunak, Ole Lund, Tudor I. Oprea, and Olivier Taboureau. 2016 · 2016
Cited alongside, same era.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Cited alongside, same era.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V. Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, Jeff Klingner, Apurva Shah, Melvin Johnson, Xiaobing Liu, Lukasz Kaiser, Stephan Gouws, Yoshikiyo Kato, Taku Kudo, Hideto Kazawa, Keith Stevens, George Kurian, Nishant Patil, Wei Wang, Cliff Young, Jason Smith, Jason Riesa, Alex Rudnick, Oriol Vinyals, Greg Corrado, Macduff Hughes, and Jeffrey Dean. 2016 · 2016
Cited alongside, same era.
PubMed 200k RCT: a dataset for sequential sentence classification in medical abstracts
BioBERT: a pre-trained biomedical language representation model for biomedical text mining
Jinhyuk Lee, Wonjin Yoon, Sungdong Kim, Donghyeon Kim, Sunkyu Kim, Chan Ho So, and Jaewoo Kang. 2019 · 2019
Later among the works it cites.
Energy and policy considerations for deep learning in NLP
Emma Strubell, Ananya Ganesh, and Andrew McCallum. 2019 · 2019
Later among the works it cites.
Language models are few-shot learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 2020
Later among the works it cites.
Efficient intent detection with dual sentence encoders
Iñigo Casanueva, Tadas Temčinas, Daniela Gerz, Matthew Henderson, and Ivan Vulić. 2020 · 2020
Later among the works it cites.
ELECTRA: Pre-training text encoders as discriminators rather than generators
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Franck Dernoncourt and Ji Young Lee. 2017 · 2017
Cited alongside, same era.
Scattertext: a browser-based tool for visualizing how corpora differ
Jason Kessler. 2017 · 2017
Cited alongside, same era.
Measuring the evolution of a scientific field through citation frames
David Jurgens, Srijan Kumar, Raine Hoover, Dan McFarland, and Dan Jurafsky. 2018 · 2018
Cited alongside, same era.
Multi-task identification of entities, relations, and coreference for scientific knowledge graph construction
Yi Luan, Luheng He, Mari Ostendorf, and Hannaneh Hajishirzi. 2018 · 2018
Cited alongside, same era.
Publicly available clinical BERT embeddings
Emily Alsentzer, John Murphy, William Boag, Wei-Hung Weng, Di Jindi, Tristan Naumann, and Matthew McDermott. 2019 · 2019
Cited alongside, same era.
SciBERT: A pretrained language model for scientific text
Iz Beltagy, Kyle Lo, and Arman Cohan. 2019 · 2019
Cited alongside, same era.
IMHO fine-tuning improves claim detection
Tuhin Chakrabarty, Christopher Hidey, and Kathy McKeown. 2019 · 2019
Cited alongside, same era.
Pretrained language models for sequential sentence classification
Arman Cohan, Iz Beltagy, Daniel King, Bhavana Dalvi, and Dan Weld. 2019 · 2019
Cited alongside, same era.
Kevin Clark, Minh-Thang Luong, Quoc V. Le, and Christopher D. Manning. 2020 · 2020
Later among the works it cites.
Domain-specific language model pretraining for biomedical natural language processing
Yu Gu, Robert Tinn, Hao Cheng, Michael Lucas, Naoto Usuyama, Xiaodong Liu, Tristan Naumann, Jianfeng Gao, and Hoifung Poon. 2020 · 2020
Later among the works it cites.
Don’t stop pretraining: Adapt language models to domains and tasks
Suchin Gururangan, Ana Marasović, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A. Smith. 2020 · 2020
Later among the works it cites.
DagoBERT: Generating derivational morphology with a pretrained language model
Valentin Hofmann, Janet Pierrehumbert, and Hinrich Schütze. 2020 · 2020
Later among the works it cites.
S2ORC: The semantic scholar open research corpus
Kyle Lo, Lucy Lu Wang, Mark Neumann, Rodney Kinney, and Daniel Weld. 2020 · 2020
Later among the works it cites.
Inexpensive domain adaptation of pretrained language models: Case studies on biomedical NER and covid-19 QA
Nina Poerner, Ulli Waltinger, and Hinrich Schütze. 2020 · 2020
Later among the works it cites.
Adapt or get left behind: Domain adaptation through BERT language model finetuning for aspect-target sentiment classification
Alexander Rietzler, Sebastian Stabinger, Paul Opitz, and Stefan Engl. 2020 · 2020
Later among the works it cites.
exBERT: Extending pre-trained models with domain-specific vocabulary under constrained training resources
Wen Tai, H. T. Kung, Xin Dong, Marcus Comiter, and Chang-Fu Kuo. 2020 · 2020
Later among the works it cites.
Multi-stage pre-training for low-resource domain adaptation
Rong Zhang, Revanth Gangi Reddy, Md Arafat Sultan, Vittorio Castelli, Anthony Ferritto, Radu Florian, Efsun Sarioglu Kayi, Salim Roukos, Avi Sil, and Todd Ward. 2020 · 2020
Later among the works it cites.
Valentin Hofmann, Janet B. Pierrehumbert, and Hinrich Schütze. 2021 · 2021
Closest in time.