Roberta: A robustly optimized bert pretraining approach
Original
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019c · 1907
Earlier work this paper cites.
Semi-supervised sequence modeling with cross-view training
Kevin Clark, Minh-Thang Luong, Christopher D. Manning, and Quoc Le. 2018 · 1925
Earlier work this paper cites.
éléments de syntaxe structurale
Lucien Tesnière. 1959 · 1959
Earlier work this paper cites.
Grammatical category disambiguation by statistical optimization
Steven J. DeRose. 1988 · 1988
Earlier work this paper cites.
Designing neural networks using genetic algorithms
Geoffrey Miller, Peter Todd, and Shailesh Hegde. 1989 · 1989
Earlier work this paper cites.
A comparative analysis of selection schemes used in genetic algorithms
David E Goldberg and Kalyanmoy Deb. 1991 · 1991
Earlier work this paper cites.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto. 1992 · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams. 1992 · 1992
Earlier work this paper cites.
An evolutionary algorithm that constructs recurrent neural networks
Peter J Angeline, Gregory M Saunders, and Jordan B Pollack. 1994 · 1994
Earlier work this paper cites.
Named entity task definition, version 2.1
Beth M. Sundheim. 1995 · 1995
Earlier work this paper cites.
Introduction to the CoNLL-2000 shared task chunking
Erik F. Tjong Kim Sang and Sabine Buchholz. 2000 · 2000
Earlier work this paper cites.
Evolving neural networks through augmenting topologies
Kenneth O Stanley and Risto Miikkulainen. 2002 · 2002
Earlier work this paper cites.
Introduction to the CoNLL-2002 shared task: Language-independent named entity recognition
Erik F. Tjong Kim Sang. 2002 · 2002
Earlier work this paper cites.
Introduction to the CoNLL-2003 shared task: Language-independent named entity recognition
Erik F. Tjong Kim Sang and Fien De Meulder. 2003 · 2003
Earlier work this paper cites.
Mining and summarizing customer reviews
Minqing Hu and Bing Liu. 2004 · 2004
Earlier work this paper cites.
Non-projective dependency parsing using spanning tree algorithms
Ryan McDonald, Fernando Pereira, Kiril Ribarov, and Jan Hajič. 2005 · 2005
Earlier work this paper cites.
CoNLL-X shared task on multilingual dependency parsing
Sabine Buchholz and Erwin Marsi. 2006 · 2006
Earlier work this paper cites.
Neuroevolution: from architectures to learning
Dario Floreano, Peter Dürr, and Claudio Mattiussi. 2008 · 2008
Earlier work this paper cites.
Autotrans: Automating transformer design via reinforced architecture search
Original
Wei Zhu, Xiaoling Wang, Xipeng Qiu, Yuan Ni, and Guotong Xie. 2020 · 2009
Earlier work this paper cites.
Part-of-speech tagging for Twitter: Annotation, features, and experiments
Kevin Gimpel, Nathan Schneider, Brendan O’Connor, Dipanjan Das, Daniel Mills, Jacob Eisenstein, Michael Heilman, Dani Yogatama, Jeffrey Flanigan, and Noah A. Smith. 2011 · 2011
Earlier work this paper cites.
Named entity recognition in tweets: An experimental study
Alan Ritter, Sam Clark, Mausam, and Oren Etzioni. 2011 · 2011
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean. 2013 · 2013
Earlier work this paper cites.
Improved part-of-speech tagging for online conversational text with word clusters
Olutobi Owoputi, Brendan O’Connor, Chris Dyer, Kevin Gimpel, Nathan Schneider, and Noah A. Smith. 2013 · 2013
Earlier work this paper cites.
Semeval 2014 task 8: Broad-coverage semantic dependency parsing
Stephan Oepen, Marco Kuhlmann, Yusuke Miyao, Daniel Zeman, Dan Flickinger, Jan Hajic, Angelina Ivanova, and Yi Zhang. 2014 · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher Manning. 2014 · 2014
Earlier work this paper cites.
SemEval-2014 task 4: Aspect based sentiment analysis
Maria Pontiki, Dimitris Galanis, John Pavlopoulos, Harris Papageorgiou, Ion Androutsopoulos, and Suresh Manandhar. 2014 · 2014
Earlier work this paper cites.
Learning character-level representations for part-of-speech tagging
Cicero D Santos and Bianca Zadrozny. 2014 · 2014
Earlier work this paper cites.
An empirical exploration of recurrent network architectures
Rafal Jozefowicz, Wojciech Zaremba, and Ilya Sutskever. 2015 · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba. 2015 · 2015
Earlier work this paper cites.
Semeval 2015 task 18: Broad-coverage semantic dependency parsing
Stephan Oepen, Marco Kuhlmann, Yusuke Miyao, Daniel Zeman, Silvie Cinková, Dan Flickinger, Jan Hajic, and Zdenka Uresova. 2015 · 2015
Earlier work this paper cites.
SemEval-2015 task 12: Aspect based sentiment analysis
Maria Pontiki, Dimitris Galanis, Haris Papageorgiou, Suresh Manandhar, and Ion Androutsopoulos. 2015 · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Earlier work this paper cites.