Fetching the paper…
Reading the bibliography…
Building pretrained language models is considered expensive and data-intensive, but must we increase dataset size to achieve better performance? We propose an alternative to larger training sets by automatically identifying smaller yet domain-representative subsets.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Multi-criteria-based active learning for named entity recognition
Dan Shen, Jie Zhang, Jian Su, Guodong Zhou, and Chew-Lim Tan. 2004 · 2004
Earlier work this paper cites.
SemEval-2010 task 8: Multi-way classification of semantic relations between pairs of nominals
Iris Hendrickx, Su Nam Kim, Zornitsa Kozareva, Preslav Nakov, Diarmuid Ó Séaghdha, Sebastian Padó, Marco Pennacchiotti, Lorenza Romano, and Stan Szpakowicz. 2010 · 2010
Earlier work this paper cites.
Intelligent selection of language model training data
Robert C. Moore and William Lewis. 2010 · 2010
Earlier work this paper cites.
Domain adaptation via pseudo in-domain data selection
Amittai Axelrod, Xiaodong He, and Jianfeng Gao. 2011 · 2011
Earlier work this paper cites.
GLISTER: generalization based data subset selection for efficient and robust learning
KrishnaTeja Killamsetty, Durga Sivasubramanian, Ganesh Ramakrishnan, and Rishabh K. Iyer. 2020 · 2012
Earlier work this paper cites.
CoNLL-2012 shared task: Modeling multilingual unrestricted coreference in OntoNotes
Sameer Pradhan, Alessandro Moschitti, Nianwen Xue, Olga Uryupina, and Yuchen Zhang. 2012 · 2012
Earlier work this paper cites.
Xenc: An open-source tool for data selection in natural language processing
Anthony Rousseau. 2013 · 2013
Earlier work this paper cites.
OntoNotes Release 5.0
Ralph Weischedel, Martha Palmer, Mitchell Marcus, Eduard Hovy, Sameer Pradhan, Lance Ramshaw, Nianwen Xue, Ann Taylor, Jeff Kaufman, Michelle Franchini, Mohammed El-Bachouti, Robert Belvin, and Ann Houston. 2013 · 2013
Earlier work this paper cites.
Submodularity for data selection in machine translation
Katrin Kirchhoff and Jeff Bilmes. 2014 · 2014
Earlier work this paper cites.
A gold standard dependency corpus for English
Natalia Silveira, Timothy Dozat, Marie-Catherine de Marneffe, Samuel Bowman, Miriam Connor, John Bauer, and Chris Manning. 2014 · 2014
Earlier work this paper cites.
Survey of data-selection methods in statistical machine translation
Sauleh Eetemadi, William Lewis, Kristina Toutanova, and Hayder Radha. 2015 · 2015
Earlier work this paper cites.
Cynical selection of language model training data
Amittai Axelrod. 2017 · 2017
Earlier work this paper cites.
Learning to select data for transfer learning with bayesian optimization
Sebastian Ruder and Barbara Plank. 2017 · 2017
Earlier work this paper cites.
Dynamic data selection for neural machine translation
Marlies van der Wees, Arianna Bisazza, and Christof Monz. 2017 · 2017
Cited alongside, same era.
Zipporah: a fast and scalable data cleaning system for noisy web-crawled parallel corpora
Hainan Xu and Philipp Koehn. 2017 · 2017
Cited alongside, same era.
How transferable are the datasets collected by active learners?
David Lowell, Zachary Chase Lipton, and Byron C. Wallace. 2018 · 2018
Cited alongside, same era.
Neural-Davidsonian semantic proto-role labeling
Rachel Rudinger, Adam Teichert, Ryan Culkin, Sheng Zhang, and Benjamin Van Durme. 2018 · 2018
Cited alongside, same era.
SciBERT: A pretrained language model for scientific text
Iz Beltagy, Kyle Lo, and Arman Cohan. 2019 · 2019
Cited alongside, same era.
What does bert look at? an analysis of bert’s attention
How Can We Know What Language Models Know?
Zhengbao Jiang, Frank F. Xu, Jun Araki, and Graham Neubig. 2020 · 2020
Later among the works it cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020 · 2020
Later among the works it cites.
BERTweet: A pre-trained language model for English tweets
Dat Quoc Nguyen, Thanh Vu, and Anh Tuan Nguyen. 2020 · 2020
Later among the works it cites.
Code and named entity recognition in StackOverflow
Jeniya Tabassum, Mounica Maddela, Wei Xu, and Alan Ritter. 2020 · 2020
Later among the works it cites.
Pre-train or annotate? domain adaptation with a constrained budget
Fan Bai, Alan Ritter, and Wei Xu. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D. Manning. 2019 · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Practical, efficient, and customizable active learning for named entity recognition in the digital humanities
Alexander Erdmann, David Joseph Wrisley, Benjamin Allen, Christopher Brown, Sophie Cohen-Bodénès, Micha Elsner, Yukun Feng, Brian Joseph, Béatrice Joyeux-Prunel, and Marie-Catherine de Marneffe. 2019 · 2019
Cited alongside, same era.
Variational pretraining for semi-supervised text classification
Suchin Gururangan, Tam Dang, Dallas Card, and Noah A. Smith. 2019 · 2019
Cited alongside, same era.
Learning to selectively transfer: Reinforced transfer learning for deep text matching
Chen Qu, Feng Ji, Minghui Qiu, Liu Yang, Zhiyu Min, Haiqing Chen, Jun Huang, and W. Bruce Croft. 2019 · 2019
Cited alongside, same era.
Active learning with deep pre-trained models for sequence tagging of clinical and biomedical texts
Artem Shelmanov, Vadim Liventsev, Danil Kireev, Nikita Khromov, Alexander Panchenko, Irina Fedulova, and Dmitry V. Dylov. 2019 · 2019
Cited alongside, same era.
LEGAL-BERT: The muppets straight out of law school
Ilias Chalkidis, Manos Fergadiotis, Prodromos Malakasiotis, Nikolaos Aletras, and Ion Androutsopoulos. 2020 · 2020
Cited alongside, same era.
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott Gray, Nick Ryder, Mikhail Pavlov, Alethea Power, Lukasz Kaiser, Mohammad Bavarian, Clemens Winter, Philippe Tillet, Felipe Petroski Such, Dave Cummings, Matthias Plappert, Fotios Chantzis, Elizabeth Barnes, Ariel Herbert-Voss, William Hebgen Guss, Alex Nichol, Alex Paino, Nikolas Tezak, Jie Tang, Igor Babuschkin, Suchir Balaji, Shantanu Jain, William Saunders, Christopher Hesse, Andrew N. Carr, Jan Leike, Josh Achiam, Vedant Misra, Evan Morikawa, Alec Radford, Matthew Knight, Miles Brundage, Mira Murati, Katie Mayer, Peter Welinder, Bob McGrew, Dario Amodei, Sam McCandlish, Ilya Sutskever, and Wojciech Zaremba. 2021 · 2021
Later among the works it cites.
Should we be pre-training? an argument for end-task aware training as an alternative
Lucio M. Dery, Paul Michel, Ameet Talwalkar, and Graham Neubig. 2021 · 2021
Later among the works it cites.
The pile: An 800gb dataset of diverse text for language modeling
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, and Connor Leahy. 2021 · 2021
Later among the works it cites.
Autosampling: Search for effective data sampling schedules
Ming Sun, Haoxuan Dou, Baopu Li, Junjie Yan, Wanli Ouyang, and Lei Cui. 2021 · 2021
Later among the works it cites.
When do you need billions of words of pretraining data?
Yian Zhang, Alex Warstadt, Xiaocheng Li, and Samuel R. Bowman. 2021 · 2021
Later among the works it cites.
Quantifying memorization across neural language models
Nicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee, Florian Tramèr, and Chiyuan Zhang. 2022 · 2022
Closest in time.
Quantifying adaptability in pre-trained language models with 500 tasks
Belinda Li, Jane Yu, Madian Khabsa, Luke Zettlemoyer, Alon Halevy, and Jacob Andreas. 2022 · 2022
Closest in time.
On the importance of effectively adapting pretrained language models for active learning
Katerina Margatina, Loic Barrault, and Nikolaos Aletras. 2022 · 2022
Closest in time.
Towards computationally feasible deep active learning
Akim Tsvigun, Artem Shelmanov, Gleb Kuzmin, Leonid Sanochkin, Daniil Larionov, Gleb Gusev, Manvel Avetisian, and Leonid Zhukov. 2022 · 2022
Closest in time.