Fetching the paper…
Reading the bibliography…
We investigate the ability of transformer models to approximate the CKY algorithm, using them to directly predict a sentence's parse and thus avoid the CKY algorithm's cubic dependence on sentence length.
Analysing mathematical reasoning abilities of neural models
David Saxton, Edward Grefenstette, Felix Hill, and Pushmeet Kohli. 2019 · 1904
Earlier work this paper cites.
On the shortest arborescence of a directed graph
Y. J. Chu and T. H. Liu. 1965 · 1965
Earlier work this paper cites.
An efficient recognition and syntax-analysis algorithm for context-free languages
Tadao Kasami. 1966 · 1966
Earlier work this paper cites.
Recognition and parsing of context-free languages in time n 3
Daniel H Younger. 1967 · 1967
Earlier work this paper cites.
Trainable grammars for speech recognition
James K Baker. 1979 · 1979
Earlier work this paper cites.
The estimation of stochastic context-free grammars using the inside-outside algorithm
K. Lari and S. J. Young. 1990 · 1990
Earlier work this paper cites.
Two experiments on learning probabilistic dependency grammars from corpora
Glenn Carroll and Eugene Charniak. 1992 · 1992
Earlier work this paper cites.
Building a large annotated corpus of english: The penn treebank
Mitchell P. Marcus, Mary Ann Marcinkiewicz, and Beatrice Santorini. 1993 · 1993
Earlier work this paper cites.
The penn chinese treebank: Phrase structure annotation of a large corpus
Naiwen Xue, Fei Xia, Fu-dong Chiou, and Marta Palmer. 2005 · 2005
Earlier work this paper cites.
Transition-based parsing of the Chinese treebank using a global discriminative model
Yue Zhang and Stephen Clark. 2009 · 2009
Earlier work this paper cites.
Overview of the SPMRL 2013 shared task: A cross-framework evaluation of parsing morphologically rich languages
Djamé Seddah, Reut Tsarfaty, Sandra Kübler, Marie Candito, Jinho D. Choi, Richárd Farkas, Jennifer Foster, Iakes Goenaga, Koldo Gojenola Galletebeitia, Yoav Goldberg, Spence Green, Nizar Habash, Marco Kuhlmann, Wolfgang Maier, Joakim Nivre, Adam Przepiórkowski, Ryan Roth, Wolfgang Seeker, Yannick Versley, Veronika Vincze, Marcin Woliński, Alina Wróblewska, and Eric Villemonte de la Clergerie. 2013 · 2013
Earlier work this paper cites.
Alex Graves, Greg Wayne, and Ivo Danihelka. 2014 · 2014
Earlier work this paper cites.
Auto-Encoding Variational Bayes
Diederik P. Kingma and Max Welling. 2014 · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba. 2015 · 2015
Earlier work this paper cites.
Grammar as a foreign language
Oriol Vinyals, Łukasz Kaiser, Terry Koo, Slav Petrov, Ilya Sutskever, and Geoffrey Hinton. 2015 · 2015
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. 2016 · 2016
Earlier work this paper cites.
Span-based constituency parsing with a structure-label system and provably optimal dynamic oracles
James Cross and Liang Huang. 2016 · 2016
Earlier work this paper cites.
Inside-outside and forward-backward algorithms are just backprop (tutorial paper)
Jason Eisner. 2016 · 2016
Earlier work this paper cites.
Gaussian error linear units (gelus)
Dan Hendrycks and Kevin Gimpel. 2016 · 2016
Earlier work this paper cites.
Perceptual losses for real-time style transfer and super-resolution
Justin Johnson, Alexandre Alahi, and Li Fei-Fei. 2016 · 2016
Earlier work this paper cites.
Lukasz Kaiser and Ilya Sutskever. 2016 · 2016
Earlier work this paper cites.
Efficient structured inference for transition-based parsing with neural networks and error states
Ashish Vaswani and Kenji Sagae. 2016 · 2016
Earlier work this paper cites.
Supervised attention for sequence-to-sequence constituency parsing
Hidetaka Kamigaito, Katsuhiko Hayashi, Tsutomu Hirao, Hiroya Takamura, Manabu Okumura, and Masaaki Nagata. 2017 · 2017
Earlier work this paper cites.
A minimal span-based neural constituency parser
Mitchell Stern, Jacob Andreas, and Dan Klein. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Universal transformers
Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, and Lukasz Kaiser. 2018 · 2018
Cited alongside, same era.
What’s going on in neural constituency parsers? an analysis
David Gaddy, Mitchell Stern, and Dan Klein. 2018 · 2018
Cited alongside, same era.
Constituency parsing with a self-attentive encoder
Nikita Kitaev and Dan Klein. 2018 · 2018
Cited alongside, same era.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter. 2018 · 2018
Cited alongside, same era.
Emergent linguistic structure in artificial neural networks trained by self-supervision
Christopher D. Manning, Kevin Clark, John Hewitt, Urvashi Khandelwal, and Omer Levy. 2020 · 2020
Later among the works it cites.
Rethinking self-attention: Towards interpretability in neural parsing
Khalil Mrini, Franck Dernoncourt, Quan Hung Tran, Trung Bui, Walter Chang, and Ndapa Nakashole. 2020 · 2020
Later among the works it cites.
Torch-struct: Deep structured prediction library
Alexander Rush. 2020 · 2020
Later among the works it cites.
Improving constituency parsing with span attention
Yuanhe Tian, Yan Song, Fei Xia, and Tong Zhang. 2020 · 2020
Later among the works it cites.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
An empirical study of building a strong baseline for constituency parsing
Jun Suzuki, Sho Takase, Hidetaka Kamigaito, Makoto Morishita, and Masaaki Nagata. 2018 · 2018
Cited alongside, same era.
Learning approximate inference networks for structured prediction
Lifu Tu and Kevin Gimpel. 2018 · 2018
Cited alongside, same era.
Transition-based parsing with lighter feed-forward networks
David Vilares and Carlos Gómez-Rodríguez. 2018 · 2018
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
A structural probe for finding syntax in word representations
John Hewitt and Christopher D. Manning. 2019 · 2019
Cited alongside, same era.
What does BERT learn about the structure of language?
Ganesh Jawahar, Benoît Sagot, and Djamé Seddah. 2019 · 2019
Cited alongside, same era.
Can transformers jump around right in natural language? assessing performance transfer from SCAN
Rahma Chaabouni, Roberto Dessì, and Eugene Kharitonov. 2021 · 2021
Later among the works it cites.
Decision transformer: Reinforcement learning via sequence modeling
Lili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee, Aditya Grover, Misha Laskin, Pieter Abbeel, Aravind Srinivas, and Igor Mordatch. 2021 · 2021
Later among the works it cites.
Consistent unsupervised estimators for anchored PCFGs
Alexander Clark and Nathanaël Fijalkow. 2021 · 2021
Later among the works it cites.
The neural data router: Adaptive control flow in transformers improves systematic generalization
Róbert Csordás, Kazuki Irie, and Jürgen Schmidhuber. 2021 · 2021
Later among the works it cites.
Do syntactic probes probe syntax? experiments with jabberwocky probing
Rowan Hall Maudslay and Ryan Cotterell. 2021 · 2021
Later among the works it cites.
A conditional splitting framework for efficient constituency parsing
Thanh-Tung Nguyen, Xuan-Phi Nguyen, Shafiq Joty, and Xiaoli Li. 2021 · 2021
Later among the works it cites.
FlashAttention: Fast and memory-efficient exact attention with IO-awareness
Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, and Christopher Ré. 2022 · 2022
Later among the works it cites.
Neural networks and the chomsky hierarchy
Grégoire Delétang, Anian Ruoss, Jordi Grau-Moya, Tim Genewein, Li Kevin Wenliang, Elliot Catt, Marcus Hutter, Shane Legg, and Pedro A Ortega. 2022 · 2022
Later among the works it cites.
Graph neural networks are dynamic programmers
Andrew Dudzik and Petar Veličković. 2022 · 2022
Later among the works it cites.
Probing for incremental parse states in autoregressive language models
Tiwalayo Eisape, Vineet Gangireddy, Roger Levy, and Yoon Kim. 2022 · 2022
Later among the works it cites.
Cramming: Training a language model on a single gpu in one day
Jonas Geiping and Tom Goldstein. 2022 · 2022
Later among the works it cites.
A generalist neural algorithmic learner
Borja Ibarz, Vitaly Kurin, George Papamakarios, Kyriacos Nikiforou, Mehdi Bennani, Róbert Csordás, Andrew Dudzik, Matko Bošnjak, Alex Vitvitskyi, Yulia Rubanova, et al. 2022 · 2022
Later among the works it cites.
Characterizing intrinsic compositionality in transformers with tree projections
Shikhar Murty, Pratyusha Sharma, Jacob Andreas, and Christopher D Manning. 2022 · 2022
Later among the works it cites.
The clrs algorithmic reasoning benchmark
Petar Veličković, Adrià Puigdomènech Badia, David Budden, Razvan Pascanu, Andrea Banino, Misha Dashevskiy, Raia Hadsell, and Charles Blundell. 2022 · 2022
Later among the works it cites.
Bottom-up constituency parsing and nested named entity recognition with pointer networks
Songlin Yang and Kewei Tu. 2022 · 2022
Later among the works it cites.
Transformers learn shortcuts to automata
Bingbin Liu, Jordan T. Ash, Surbhi Goel, Akshay Krishnamurthy, and Cyril Zhang. 2023 · 2023
Closest in time.
Don’t parse, choose spans! continuous and discontinuous constituency parsing via autoregressive span selection
Songlin Yang and Kewei Tu. 2023 · 2023
Closest in time.
Do transformers parse while predicting the masked word?
Haoyu Zhao, Abhishek Panigrahi, Rong Ge, and Sanjeev Arora. 2023 · 2023
Closest in time.