Fetching the paper…
Reading the bibliography…
Much recent work suggests that incorporating syntax information from dependency trees can improve task-specific transformer models.
RoBERTa: A robustly optimized BERT pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Dropout: A simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. 2014 · 1958
Earlier work this paper cites.
The viterbi algorithm
G. D. Forney. 1973 · 1973
Earlier work this paper cites.
Brown corpus manual
W. N. Francis and H. Kucera. 1979 · 1979
Earlier work this paper cites.
Conditional random fields: Probabilistic models for segmenting and labeling sequence data
John D. Lafferty, Andrew McCallum, and Fernando C. N. Pereira. 2001 · 2001
Earlier work this paper cites.
Introduction to the CoNLL-2005 shared task: Semantic role labeling
Xavier Carreras and Lluís Màrquez. 2005 · 2005
Earlier work this paper cites.
Semantic role labeling using different syntactic views
Sameer Pradhan, Wayne Ward, Kadri Hacioglu, James Martin, and Daniel Jurafsky. 2005 · 2005
Earlier work this paper cites.
RelEx – Relation extraction using dependency parse trees
Katrin Fundel, Robert Küffner, and Ralf Zimmer. 2006 · 2006
Earlier work this paper cites.
The Stanford typed dependencies representation
Marie-Catherine de Marneffe and Christopher D. Manning. 2008 · 2008
Earlier work this paper cites.
Extracting complex biological events with rich graph-based feature sets
Jari Björne, Juho Heimonen, Filip Ginter, Antti Airola, Tapio Pahikkala, and Tapio Salakoski. 2009 · 2009
Earlier work this paper cites.
Joint parsing and named entity recognition
Jenny Rose Finkel and Christopher D. Manning. 2009 · 2009
Earlier work this paper cites.
Open-domain semantic role labeling by modeling word spans
Fei Huang and Alexander Yates. 2010 · 2010
Earlier work this paper cites.
CoNLL-2012 shared task: Modeling multilingual unrestricted coreference in OntoNotes
Sameer Pradhan, Alessandro Moschitti, Nianwen Xue, Olga Uryupina, and Yuchen Zhang. 2012 · 2012
Earlier work this paper cites.
The Stanford CoreNLP natural language processing toolkit
Christopher Manning, Mihai Surdeanu, John Bauer, Jenny Finkel, Steven Bethard, and David McClosky. 2014 · 2014
Earlier work this paper cites.
GloVe: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher Manning. 2014 · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba. 2015 · 2015
Cited alongside, same era.
A dependency-based neural network for relation classification
Yang Liu, Furu Wei, Sujian Li, Heng Ji, Ming Zhou, and Houfeng Wang. 2015 · 2015
Cited alongside, same era.
Training very deep networks
Rupesh K Srivastava, Klaus Greff, and Jürgen Schmidhuber. 2015 · 2015
Cited alongside, same era.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. 2016 · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Cited alongside, same era.
Bridging nonlinearities and stochastic regularizers with gaussian error linear units
Relational inductive biases, deep learning, and graph networks
Peter Battaglia, Jessica Blake Chandler Hamrick, Victor Bapst, Alvaro Sanchez, Vinicius Zambaldi, Mateusz Malinowski, Andrea Tacchetti, David Raposo, Adam Santoro, Ryan Faulkner, Caglar Gulcehre, Francis Song, Andy Ballard, Justin Gilmer, George E. Dahl, Ashish Vaswani, Kelsey Allen, Charles Nash, Victoria Jayne Langston, Chris Dyer, Nicolas Heess, Daan Wierstra, Pushmeet Kohli, Matt Botvinick, Oriol Vinyals, Yujia Li, and Razvan Pascanu. 2018 · 2018
Later among the works it cites.
Deep contextualized word representations
Matthew Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018 · 2018
Later among the works it cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018 · 2018
Later among the works it cites.
Linguistically-informed self-attention for semantic role labeling
Emma Strubell, Patrick Verga, Daniel Andor, David Weiss, and Andrew McCallum. 2018 · 2018
Later among the works it cites.
Graph attention networks
Petar Veličković, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Lio, and Yoshua Bengio. 2018 · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dan Hendrycks and Kevin Gimpel. 2016 · 2016
Cited alongside, same era.
End-to-end relation extraction using LSTMs on sequences and tree structures
Makoto Miwa and Mohit Bansal. 2016 · 2016
Cited alongside, same era.
Neural semantic role labeling with dependency path embeddings
Michael Roth and Mirella Lapata. 2016 · 2016
Cited alongside, same era.
Deep biaffine attention for neural dependency parsing
Timothy Dozat and Christopher D Manning. 2017 · 2017
Cited alongside, same era.
Representation learning on graphs: Methods and applications
William L. Hamilton, Rex Ying, and Jure Leskovec. 2017 · 2017
Cited alongside, same era.
Efficient dependency-guided named entity recognition
Zhanming Jie, Aldrian Obaja Muis, and Wei Lu. 2017 · 2017
Cited alongside, same era.
Encoding sentences with graph convolutional networks for semantic role labeling
Diego Marcheggiani and Ivan Titov. 2017 · 2017
Cited alongside, same era.
Later among the works it cites.
Graph convolution over pruned dependency trees improves relation extraction
Yuhao Zhang, Peng Qi, and Christopher D. Manning. 2018 · 2018
Later among the works it cites.
What does BERT look at? An analysis of BERT’s attention
Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D. Manning. 2019 · 2019
Later among the works it cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
A structural probe for finding syntax in word representations
John Hewitt and Christopher D. Manning. 2019 · 2019
Later among the works it cites.
Dependency-guided LSTM-CRF for named entity recognition
Zhanming Jie and Wei Lu. 2019 · 2019
Later among the works it cites.
BERT rediscovers the classical NLP pipeline
Ian Tenney, Dipanjan Das, and Ellie Pavlick. 2019 · 2019
Later among the works it cites.
TACRED revisited: A thorough evaluation of the TACRED relation extraction task
Christoph Alt, Aleksandra Gabryszak, and Leonhard Hennig. 2020 · 2020
Closest in time.
Syntactic structure distillation pretraining for bidirectional encoders
Adhiguna Kuncoro, Lingpeng Kong, Daniel Fried, Dani Yogatama, Laura Rimell, Chris Dyer, and Phil Blunsom. 2020 · 2020
Closest in time.
Universal Dependencies v2: An evergrowing multilingual treebank collection
Joakim Nivre, Marie-Catherine de Marneffe, Filip Ginter, Jan Hajič, Christopher D. Manning, Sampo Pyysalo, Sebastian Schuster, Francis Tyers, and Daniel Zeman. 2020 · 2020
Closest in time.
Stanza: A Python natural language processing toolkit for many human languages
Peng Qi, Yuhao Zhang, Yuhui Zhang, Jason Bolton, and Christopher D. Manning. 2020 · 2020
Closest in time.