Fetching the paper…
Reading the bibliography…
Deep learning models generalize well to in-distribution data but struggle to generalize compositionally, i.e., to combine a set of learned primitives to solve more complex tasks.
Good-enough compositional data augmentation
Jacob Andreas. 2019 · 1904
Earlier work this paper cites.
Xingxing Zhang, Furu Wei, and Ming Zhou. 2019 · 1905
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2019 · 1910
Earlier work this paper cites.
Measuring compositional generalization: A comprehensive method on realistic data
Daniel Keysers, Nathanael Schärli, Nathan Scales, Hylke Buisman, Daniel Furrer, Sergii Kashubin, Nikola Momchev, Danila Sinopalnikov, Lukasz Stafiniak, Tibor Tihon, et al. 2019 · 1912
Earlier work this paper cites.
Are transformers universal approximators of sequence-to-sequence functions?
Chulhee Yun, Srinadh Bhojanapalli, Ankit Singh Rawat, Sashank J Reddi, and Sanjiv Kumar. 2019 · 1912
Earlier work this paper cites.
Language, thought and compositionality
Jerry Fodor. 2001 · 2001
Earlier work this paper cites.
Incorporating bert into neural machine translation
Jinhua Zhu, Yingce Xia, Lijun Wu, Di He, Tao Qin, Wengang Zhou, Houqiang Li, and Tie-Yan Liu. 2020 · 2002
Earlier work this paper cites.
Etc: Encoding long and structured inputs in transformers
Joshua Ainslie, Santiago Ontanon, Chris Alberti, Vaclav Cvicek, Zachary Fisher, Philip Pham, Anirudh Ravula, Sumit Sanghai, Qifan Wang, and Li Yang. 2020 · 2004
Earlier work this paper cites.
Compositional generalization in semantic parsing: Pre-training vs. specialized architectures
Daniel Furrer, Marc van Zee, Nathan Scales, and Nathanael Schärli. 2020 · 2007
Earlier work this paper cites.
The eos decision and length extrapolation
Benjamin Newman, John Hewitt, Percy Liang, and Christopher D Manning. 2020 · 2010
Cited alongside, same era.
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton. 2015 · 2015
Cited alongside, same era.
Pointer networks
Vinyals Oriol, Fortunato Meire, Jaitly Navdeep, C Cortes, ND Lawrence, DD Lee, M Sugiyama, and R Garnett. 2015 · 2015
Cited alongside, same era.
Deep learning
Ian Goodfellow, Yoshua Bengio, and Aaron Courville. 2016 · 2016
Cited alongside, same era.
Hybrid computing using a neural network with dynamic external memory
Alex Graves, Greg Wayne, Malcolm Reynolds, Tim Harley, Ivo Danihelka, Agnieszka Grabska-Barwińska, Sergio Gómez Colmenarejo, Edward Grefenstette, Tiago Ramalho, John Agapiou, et al. 2016 · 2016
Cited alongside, same era.
Incorporating copying mechanism in sequence-to-sequence learning
Generalization without systematicity: On the compositional skills of sequence-to-sequence recurrent networks
Brenden Lake and Marco Baroni. 2018 · 2018
Later among the works it cites.
Memorize or generalize? searching for a compositional rnn in a haystack
Adam Liška, Germán Kruszewski, and Marco Baroni. 2018 · 2018
Later among the works it cites.
Listops: A diagnostic dataset for latent tree learning
Nikita Nangia and Samuel R Bowman. 2018 · 2018
Later among the works it cites.
Self-attention with relative position representations
Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani. 2018 · 2018
Later among the works it cites.
Compositionality decomposed: how do neural networks generalise?
Dieuwke Hupkes, Verna Dankers, Mathijs Mul, and Elia Bruni. 2020 · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jiatao Gu, Zhengdong Lu, Hang Li, and Victor OK Li. 2016 · 2016
Cited alongside, same era.
Caglar Gulcehre, Sungjin Ahn, Ramesh Nallapati, Bowen Zhou, and Yoshua Bengio. 2016 · 2016
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Cited alongside, same era.
Later among the works it cites.
Andrea Banino, Jan Balaguer, and Charles Blundell. 2021 · 2021
Closest in time.
The devil is in the detail: Simple tricks improve systematic generalization of transformers
Róbert Csordás, Kazuki Irie, and Jürgen Schmidhuber. 2021 · 2021
Closest in time.
Improving compositional generalization in classification tasks via structure annotations
Juyong Kim, Pradeep Ravikumar, Joshua Ainslie, and Santiago Ontañón. 2021 · 2021
Closest in time.
Making transformers solve compositional tasks
Santiago Ontañón, Joshua Ainslie, Vaclav Cvicek, and Zachary Fisher. 2021 · 2021
Closest in time.