Fetching the paper…
Reading the bibliography…
We consider retrofitting structure-aware Transformer-based language model for facilitating end tasks by proposing to exploit syntactic distance to encode both the phrasal constituency and dependency connection into the language model.
Assessing bert’s syntactic abilities
Yoav Goldberg. 2019 · 1901
Earlier work this paper cites.
Unsupervised latent tree induction with deep inside-outside recursive autoencoders
Andrew Drozdov, Patrick Verga, Mohit Yadav, Mohit Iyyer, and Andrew McCallum. 2019 · 1904
Earlier work this paper cites.
Analyzing the structure of attention in a transformer language model
Jesse Vig and Yonatan Belinkov. 2019 · 1906
Earlier work this paper cites.
Xlnet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Ruslan Salakhutdinov, and Quoc V Le. 2019 · 1906
Earlier work this paper cites.
Revealing the dark secrets of BERT
Olga Kovaleva, Alexey Romanov, Anna Rogers, and Anna Rumshisky. 2019 · 1908
Earlier work this paper cites.
Building a large annotated corpus of english: The penn treebank
Mitchell P. Marcus, Beatrice Santorini, and Mary Ann Marcinkiewicz. 1993 · 1993
Earlier work this paper cites.
Introduction to the CoNLL-2003 shared task: Language-independent named entity recognition
Erik F. Tjong Kim Sang and Fien De Meulder. 2003 · 2003
Earlier work this paper cites.
The CoNLL-2009 shared task: Syntactic and semantic dependencies in multiple languages
Jan Hajič, Massimiliano Ciaramita, Richard Johansson, Daisuke Kawahara, Maria Antònia Martí, Lluís Màrquez, Adam Meyers, Joakim Nivre, Sebastian Padó, Jan Štěpánek, Pavel Straňák, Mihai Surdeanu, Nianwen Xue, and Yi Zhang. 2009 · 2009
Earlier work this paper cites.
SemEval-2010 task 8: Multi-way classification of semantic relations between pairs of nominals
Iris Hendrickx, Su Nam Kim, Zornitsa Kozareva, Preslav Nakov, Diarmuid Ó Séaghdha, Sebastian Padó, Marco Pennacchiotti, Lorenza Romano, and Stan Szpakowicz. 2010 · 2010
Earlier work this paper cites.
Learning continuous phrase representations and syntactic parsing with recursive neural networks
Richard Socher, Christopher D. Manning, and Andrew Y. Ng. 2010 · 2010
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Tomas Mikolov, Ilya Sutskever, Kai Chen, Gregory S. Corrado, and Jeffrey Dean. 2013 · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts. 2013 · 2013
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher Manning. 2014 · 2014
Earlier work this paper cites.
Improved semantic representations from tree-structured long short-term memory networks
Kai Sheng Tai, Richard Socher, and Christopher D. Manning. 2015 · 2015
Earlier work this paper cites.
Incremental recurrent neural network dependency parser with search-based discriminative training
Majid Yazdani and James Henderson. 2015 · 2015
Cited alongside, same era.
Recurrent neural network grammars
Chris Dyer, Adhiguna Kuncoro, Miguel Ballesteros, and Noah A. Smith. 2016 · 2016
Cited alongside, same era.
Towards string-to-tree neural machine translation
Roee Aharoni and Yoav Goldberg. 2017 · 2017
Cited alongside, same era.
Tree-structured decoding with doubly-recurrent neural networks
David Alvarez-Melis and Tommi S. Jaakkola. 2017 · 2017
Cited alongside, same era.
Deep biaffine attention for neural dependency parsing
Timothy Dozat and Christopher D. Manning. 2017 · 2017
Cited alongside, same era.
Learning to parse and translate improves neural machine translation
LSTMs can learn syntax-sensitive dependencies well, but modeling structure makes them better
Adhiguna Kuncoro, Chris Dyer, John Hale, Dani Yogatama, Stephen Clark, and Phil Blunsom. 2018 · 2018
Later among the works it cites.
Deep contextualized word representations
Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018 · 2018
Later among the works it cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018 · 2018
Later among the works it cites.
A tree-based decoder for neural machine translation
Xinyi Wang, Hieu Pham, Pengcheng Yin, and Graham Neubig. 2018 · 2018
Later among the works it cites.
You only need attention to traverse trees
Mahtab Ahmed, Muhammad Rifayat Samee, and Robert E. Mercer. 2019 · 2019
Later among the works it cites.
Transformer-XL: Attentive language models beyond a fixed-length context
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Akiko Eriguchi, Yoshimasa Tsuruoka, and Kyunghyun Cho. 2017 · 2017
Cited alongside, same era.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. 2017 · 2017
Cited alongside, same era.
Abstract syntax networks for code generation and semantic parsing
Maxim Rabinovich, Mitchell Stern, and Dan Klein. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Learning to parse from a semantic objective: It works. is it syntax?
Adina Williams, Andrew Drozdov, and Samuel R. Bowman. 2017 · 2017
Cited alongside, same era.
What you can cram into a single $&!#* vector: Probing sentence embeddings for linguistic properties
Alexis Conneau, German Kruszewski, Guillaume Lample, Loïc Barrault, and Marco Baroni. 2018 · 2018
Cited alongside, same era.
Top-down tree structured decoding with syntactic connections for neural machine translation and parsing
Jetic Gū, Hassan S. Shavarani, and Anoop Sarkar. 2018 · 2018
Cited alongside, same era.
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc Le, and Ruslan Salakhutdinov. 2019 · 2019
Later among the works it cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
Visualizing and understanding the effectiveness of BERT
Yaru Hao, Li Dong, Furu Wei, and Ke Xu. 2019 · 2019
Later among the works it cites.
What does BERT learn about the structure of language?
Ganesh Jawahar, Benoît Sagot, and Djamé Seddah. 2019 · 2019
Later among the works it cites.
Improving neural language models by segmenting, attending, and predicting the future
Hongyin Luo, Lan Jiang, Yonatan Belinkov, and James Glass. 2019 · 2019
Later among the works it cites.
What do you learn from context? probing for sentence structure in contextualized word representations
Ian Tenney, Patrick Xia, Berlin Chen, Alex Wang, Adam Poliak, R. Thomas McCoy, Najoung Kim, Benjamin Van Durme, Samuel R. Bowman, Dipanjan Das, and Ellie Pavlick. 2019 · 2019
Later among the works it cites.
Tree transformer: Integrating tree structures into self-attention
Yaushian Wang, Hung-Yi Lee, and Yun-Nung Chen. 2019 · 2019
Later among the works it cites.
Syntax-aware neural semantic role labeling
Qingrong Xia, Zhenghua Li, Min Zhang, Meishan Zhang, Guohong Fu, Rui Wang, and Luo Si. 2019 · 2019
Later among the works it cites.
Cross-lingual semantic role labeling with model transfer
Hao Fei, Meishan Zhang, Fei Li, and Donghong Ji. 2020 · 2020
Closest in time.