Fetching the paper…
Reading the bibliography…
Human language is known to exhibit a nested, hierarchical structure, allowing us to form complex sentences out of smaller pieces.
Pay less attention with lightweight and dynamic convolutions
Felix Wu, Angela Fan, Alexei Baevski, Yann N. Dauphin, and Michael Auli. 2019 · 1901
Earlier work this paper cites.
Unsupervised latent tree induction with deep inside-outside recursive autoencoders
Andrew Drozdov, Pat Verga, Mohit Yadav, Mohit Iyyer, and Andrew McCallum. 2019 · 1904
Earlier work this paper cites.
fairseq: A fast, extensible toolkit for sequence modeling
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli. 2019 · 1904
Earlier work this paper cites.
Open sesame: Getting inside bert’s linguistic knowledge
Yongjie Lin, Yi Chern Tan, and Robert Frank. 2019 · 1906
Earlier work this paper cites.
Tree-transformer: A transformer-based method for correction of tree-structured data
Jacob Harer, Chris Reale, and Peter Chin. 2019 · 1908
Earlier work this paper cites.
Albert: A lite bert for self-supervised learning of language representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2019 · 1909
Earlier work this paper cites.
Tree transformer: Integrating tree structures into self-attention
Yau-Shian Wang, Hung-Yi Lee, and Yun-Nung Chen. 2019 · 1909
Earlier work this paper cites.
Three models for the description of language
N. Chomsky. 1956 · 1956
Earlier work this paper cites.
An efficient recognition and syntax-analysis algorithm for context-free languages
Tadao Kasami. 1965 · 1965
Earlier work this paper cites.
Context-free language processing in time n3
Daniel H. Younger. 1966 · 1966
Earlier work this paper cites.
Programming Languages and Their Compilers: Preliminary Notes
John Cocke. 1969 · 1969
Earlier work this paper cites.
Universal grammar
Richard Montague. 1970 · 1970
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
English gigaword
David Graff, Junbo Kong, Ke Chen, and Kazuaki Maeda. 2003 · 2003
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases
William B. Dolan and Chris Brockett. 2005 · 2005
Earlier work this paper cites.
Moses: Open source toolkit for statistical machine translation
Philipp Koehn, Hieu Hoang, Alexandra Birch, Chris Callison-Burch, Marcello Federico, Nicola Bertoldi, Brooke Cowan, Wade Shen, Christine Moran, Richard Zens, Chris Dyer, Ondřej Bojar, Alexandra Constantin, and Evan Herbst. 2007 · 2007
Earlier work this paper cites.
WIT3: Web inventory of transcribed and translated talks
Mauro Cettolo, Christian Girardi, and Marcello Federico. 2012 · 2012
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts. 2013a · 2013
Cited alongside, same era.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts. 2013b · 2013
Cited alongside, same era.
A large annotated corpus for learning natural language inference
Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015 · 2015
Cited alongside, same era.
Compositional distributional semantics with long short term memory
Phong Le and Willem Zuidema. 2015 · 2015
Cited alongside, same era.
A neural attention model for abstractive sentence summarization
Alexander M. Rush, Sumit Chopra, and Jason Weston. 2015 · 2015
Cited alongside, same era.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel Bowman. 2018 · 2018
Later among the works it cites.
You only need attention to traverse trees
Mahtab Ahmed, Muhammad Rifayat Samee, and Robert E. Mercer. 2019 · 2019
Later among the works it cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. 2019 · 2019
Later among the works it cites.
Novel positional encodings to enable tree-based transformers
Vighnesh Shiv and Chris Quirk. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2015 · 2015
Cited alongside, same era.
Improved semantic representations from tree-structured long short-term memory networks
Kai Sheng Tai, Richard Socher, and Christopher D. Manning. 2015 · 2015
Cited alongside, same era.
Recurrent neural network grammars
Chris Dyer, Adhiguna Kuncoro, Miguel Ballesteros, and Noah A. Smith. 2016 · 2016
Cited alongside, same era.
Towards neural machine translation with latent tree attention
James Bradbury and Richard Socher. 2017 · 2017
Cited alongside, same era.
SemEval-2017 task 1: Semantic textual similarity multilingual and crosslingual focused evaluation
Daniel Cer, Mona Diab, Eneko Agirre, Iñigo Lopez-Gazpio, and Lucia Specia. 2017 · 2017
Cited alongside, same era.
Learning to compose task-specific tree structures
Jihun Choi, Kang Min Yoo, and Sang goo Lee. 2017 · 2017
Cited alongside, same era.
Jointly learning sentence embeddings and syntax with unsupervised tree-lstms
Jean Maillard, Stephen Clark, and Dani Yogatama. 2017 · 2017
Cited alongside, same era.
Hao Zheng and Mirella Lapata. 2022 · 2019
Later among the works it cites.
COGS: A compositional generalization challenge based on semantic interpretation
Najoung Kim and Tal Linzen. 2020 · 2020
Later among the works it cites.
Heads-up! unsupervised constituency parsing via self-attention heads
Bowen Li, Taeuk Kim, Reinald Kim Amplayo, and Frank Keller. 2020 · 2020
Later among the works it cites.
Tree-structured attention with hierarchical accumulation
Xuan-Phi Nguyen, Shafiq Joty, Steven Hoi, and Richard Socher. 2020 · 2020
Later among the works it cites.
A primer in bertology: What we know about how bert works
Anna Rogers, Olga Kovaleva, and Anna Rumshisky. 2020 · 2020
Later among the works it cites.
The devil is in the detail: Simple tricks improve systematic generalization of transformers
R. Csordás, Kazuki Irie, and J. Schmidhuber. 2021 · 2021
Later among the works it cites.
R2d2: Recursive transformer based on differentiable tree for interpretable hierarchical language modeling
Xiang Hu, Haitao Mi, Zujie Wen, Yafang Wang, Yi Su, Jing Zheng, and Gerard de Melo. 2021 · 2021
Later among the works it cites.
Beyond reptile: Meta-learned dot-product maximization between gradients for improved single-task regularization
Akhil Kedia, Sai Chetan Chinthakindi, and Wonho Ryu. 2021 · 2021
Later among the works it cites.
On compositional generalization of neural machine translation
Yafu Li, Yongjing Yin, Yulong Chen, and Yue Zhang. 2021 · 2021
Later among the works it cites.
Recursive tree-structured self-attention for answer sentence selection
Khalil Mrini, Emilia Farcas, and Ndapa Nakashole. 2021 · 2021
Later among the works it cites.
Improving compositional generalization with latent structure and data augmentation
Linlu Qiu, Peter Shaw, Panupong Pasupat, Pawel Krzysztof Nowak, Tal Linzen, Fei Sha, and Kristina Toutanova. 2021 · 2021
Later among the works it cites.
Transformer grammars: Augmenting transformer language models with syntactic inductive biases at scale
Laurent Sartran, Samuel Barrett, Adhiguna Kuncoro, Miloš Stanojević, Phil Blunsom, and Chris Dyer. 2022 · 2022
Closest in time.
Learning program representations with a tree-structured transformer
Wenhan Wang, Kechi Zhang, Ge Li, Shangqing Liu, Anran Li, Zhi Jin, and Yang Liu. 2022 · 2022
Closest in time.