Fetching the paper…
Reading the bibliography…
When trained on language data, do transformers learn some arbitrary computation that utilizes the full capacity of the architecture or do they learn a simpler, tree-like computation, hypothesized to underlie compositional meaning systems like human languages? There is an apparent tension between compositional accounts of human language understanding, which are based on a restricted bottom-up computational process, and the enormous success of neural models like transformers, which can route information arbitrarily between different parts of their input.
The compositionality of neural networks: integrating symbolism and connectionism
Dieuwke Hupkes, Verna Dankers, Mathijs Mul, and Elia Bruni · 1908
Earlier work this paper cites.
Syntactic Structures
Noam Chomsky · 1957
Earlier work this paper cites.
Universal grammar
Richard Montague · 1970
Earlier work this paper cites.
Structure dependence in grammar formation
Stephen Crain and Mineharu Nakayama · 1987
Earlier work this paper cites.
Tensor product variable binding and the representation of symbolic structures in connectionist systems
Paul Smolensky · 1990
Earlier work this paper cites.
A procedure for quantitatively comparing the syntactic coverage of English grammars
E. Black, S. Abney, D. Flickenger, C. Gdaniec, R. Grishman, P. Harrison, D. Hindle, R. Ingria, F. Jelinek, J. Klavans, M. Liberman, M. Marcus, S. Roukos, B. Santorini, and T. Strzalkowski · 1991
Earlier work this paper cites.
Learning to parse database queries using inductive logic programming
M. Zelle and R. J. Mooney · 1996
Earlier work this paper cites.
Generalizing from several related classification tasks to a new unlabeled sample
Gilles Blanchard, Gyemin Lee, and Clayton Scott · 2011
Earlier work this paper cites.
Cortical representation of the constituent structure of sentences
Christophe Pallier, Anne-Dominique Devauchelle, and Stanislas Dehaene · 2011
Earlier work this paper cites.
Domain generalization via invariant feature representation
Krikamol Muandet, David Balduzzi, and Bernhard Schölkopf · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Y Wu, Jason Chuang, Christopher D Manning, Andrew Y Ng, and Christopher Potts · 2013
Earlier work this paper cites.
Improved semantic representations from tree-structured long short-term memory networks
Kai Shen Tai, Richard Socher, and Christopher D. Manning · 2015
Cited alongside, same era.
Neural module networks
Jacob Andreas, Marcus Rohrbach, Trevor Darrell, and Dan Klein · 2016
Cited alongside, same era.
Recurrent neural network grammars
Chris Dyer, Adhiguna Kuncoro, Miguel Ballesteros, and Noah A Smith · 2016
Cited alongside, same era.
Domain-adversarial training of neural networks
Yaroslav Ganin, Evgeniya Ustinova, Hana Ajakan, Pascal Germain, Hugo Larochelle, Francois Laviolette, Mario March, and Victor Lempitsky · 2016
Cited alongside, same era.
Assessing the ability of LSTMs to learn syntax-sensitive dependencies
Tal Linzen, Emmanuel Dupoux, and Yoav Goldberg · 2016
Cited alongside, same era.
A minimal span-based neural constituency parser
Mitchell Stern, Jacob Andreas, and Dan Klein · 2017
Cited alongside, same era.
Targeted syntactic evaluation of language models
Rebecca Marvin and Tal Linzen · 2018
Later among the works it cites.
Measuring compositionality in representation learning
Jacob Andreas · 2019
Later among the works it cites.
Systematic generalization: What is required and can it be learned?
Dzmitry Bahdanau, Shikhar Murty, Michael Noukhovitch, Thien Huu Nguyen, Harm de Vries, and Aaron Courville · 2019
Later among the works it cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Later among the works it cites.
Rnns implicitly implement tensor product representations
R. Thomas McCoy, Tal Linzen, Ewan Dunbar, and Paul Smolensky · 2019
Later among the works it cites.
Ordered neurons: Integrating tree structures into recurrent neural networks
Yikang Shen, Shawn Tan, Alessandro Sordoni, and Aaron Courville · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Improving text-to-SQL evaluation methodology
Catherine Finegan-Dollak, Jonathan K. Kummerfeld, Li Zhang, Karthik Ramanathan, Sesh Sadasivam, Rui Zhang, and Dragomir Radev · 2018
Cited alongside, same era.
Finding syntax in human encephalography with beam search
John Hale, Chris Dyer, Adhiguna Kuncoro, and Jonathan Brennan · 2018
Cited alongside, same era.
Generalization without systematicity: On the compositional skills of sequence-to-sequence recurrent networks
Brenden Lake and Marco Baroni · 2018
Cited alongside, same era.
Later among the works it cites.
COGS: A compositional generalization challenge based on semantic interpretation
Najoung Kim and Tal Linzen · 2020
Later among the works it cites.
Beyond accuracy: Behavioral testing of NLP models with CheckList
Marco Tulio Ribeiro, Tongshuang Wu, Carlos Guestrin, and Sameer Singh · 2020
Later among the works it cites.
Assessing phrasal representation and composition in transformers
Lang Yu and Allyson Ettinger · 2020
Later among the works it cites.
Revisiting the compositional generalization abilities of neural sequence models
Arkil Patel, Satwik Bhattamishra, Phil Blunsom, and Navin Goyal · 2022
Closest in time.