Fetching the paper…
Reading the bibliography…
Despite their practical success, modern seq2seq architectures are unable to generalize systematically on several SCAN tasks.
Compositional generalization in a deep seq2seq model by separating syntax and semantics
Jake Russin, Jason Jo, Randall C O’Reilly, and Yoshua Bengio. 2019 · 1904
Earlier work this paper cites.
Compositional generalization for primitive substitutions
Yuanpeng Li, Liang Zhao, Jianyu Wang, and Joel Hestness. 2019 · 1910
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2019 · 1910
Earlier work this paper cites.
Towards a human-like open-domain chatbot
Daniel Adiwardana, Minh-Thang Luong, David R So, Jamie Hall, Noah Fiedel, Romal Thoppilan, Zi Yang, Apoorv Kulshreshtha, Gaurav Nemade, Yifeng Lu, et al. 2020 · 2001
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
Ilya Sutskever, James Martens, George Dahl, and Geoffrey Hinton. 2013 · 2013
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba. 2014 · 2014
Earlier work this paper cites.
The Stanford CoreNLP natural language processing toolkit
Christopher Manning, Mihai Surdeanu, John Bauer, Jenny Finkel, Steven Bethard, and David McClosky. 2014 · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V Le. 2014 · 2014
Earlier work this paper cites.
Large-scale evidence of dependency length minimization in 37 languages
Richard Futrell, Kyle Mahowald, and Edward Gibson. 2015 · 2015
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. 2016 · 2016
Earlier work this paper cites.
Weight normalization: A simple reparameterization to accelerate training of deep neural networks
Tim Salimans and Durk P Kingma. 2016 · 2016
Earlier work this paper cites.
Language modeling with gated convolutional networks
Yann N Dauphin, Angela Fan, Michael Auli, and David Grangier. 2017 · 2017
Cited alongside, same era.
Convolutional sequence to sequence learning
Jonas Gehring, Michael Auli, David Grangier, Denis Yarats, and Yann N. Dauphin. 2017 · 2017
Cited alongside, same era.
Jump to better conclusions: Scan both left and right
Jasmijn Bastings, Marco Baroni, Jason Weston, Kyunghyun Cho, and Douwe Kiela. 2018 · 2018
Cited alongside, same era.
Learning compositionally through attentive guidance
Dieuwke Hupkes, Anand Singh, Kris Korrel, German Kruszewski, and Elia Bruni. 2018 · 2018
Cited alongside, same era.
Generalization without systematicity: On the compositional skills of sequence-to-sequence recurrent networks
Brenden Lake and Marco Baroni. 2018 · 2018
Cited alongside, same era.
Two new evaluation datasets for low-resource machine translation: Nepali-english and sinhala-english
Francisco Guzmán, Peng-Jen Chen, Myle Ott, Juan Pino, Guillaume Lample, Philipp Koehn, Vishrav Chaudhary, and Marc’Aurelio Ranzato. 2019 · 2019
Later among the works it cites.
Compositional generalization through meta sequence-to-sequence learning
Brenden Lake. 2019 · 2019
Later among the works it cites.
fairseq: A fast, extensible toolkit for sequence modeling
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli. 2019 · 2019
Later among the works it cites.
Compositional generalization via neural-symbolic stack machines
Xinyun Chen, Chen Liang, Adams Wei Yu, Dawn Song, and Denny Zhou. 2020 · 2020
Later among the works it cites.
On the relationship between self-attention and convolutional layers
Jean-Baptiste Cordonnier, Andreas Loukas, and Martin Jaggi. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Joao Loula, Marco Baroni, and Brenden M Lake. 2018 · 2018
Cited alongside, same era.
Self-attention with relative position representations
Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani. 2018 · 2018
Cited alongside, same era.
Modeling localness for self-attention networks
Baosong Yang, Zhaopeng Tu, Derek F Wong, Fandong Meng, Lidia S Chao, and Tong Zhang. 2018 · 2018
Cited alongside, same era.
Linguistic generalization and compositionality in modern artificial neural networks
Marco Baroni. 2019 · 2019
Cited alongside, same era.
CNNs found to jump around more skillfully than RNNs: Compositional generalization in seq2seq convolutional networks
Roberto Dessì and Marco Baroni. 2019 · 2019
Cited alongside, same era.
Permutation equivariant models for compositional generalization in language
Jonathan Gordon, David Lopez-Paz, Marco Baroni, and Diane Bouchacourt. 2019 · 2019
Cited alongside, same era.
Adaptive attention span in transformers
Sainbayar Sukhbaatar, Édouard Grave, Piotr Bojanowski, and Armand Joulin
Cited in the paper.
Compositionality decomposed: How do neural networks generalise?
Dieuwke Hupkes, Verna Dankers, Mathijs Mul, and Elia Bruni. 2020 · 2020
Later among the works it cites.
Does syntax need to grow on trees? sources of hierarchical inductive bias in sequence-to-sequence networks
R Thomas McCoy, Robert Frank, and Tal Linzen. 2020 · 2020
Later among the works it cites.
Stabilizing transformers for reinforcement learning
Emilio Parisotto, Francis Song, Jack Rae, Razvan Pascanu, Caglar Gulcehre, Siddhant Jayakumar, Max Jaderberg, Raphael Lopez Kaufman, Aidan Clark, Seb Noury, et al. 2020 · 2020
Later among the works it cites.
Do transformers need deep long-range memory?
Jack Rae and Ali Razavi. 2020 · 2020
Later among the works it cites.
The paradox of the compositionality of natural language: a neural machine translation case study
Verna Dankers, Elia Bruni, and Dieuwke Hupkes. 2021 · 2021
Closest in time.
What they do when in doubt: a study of inductive biases in seq2seq learners
Eugene Kharitonov and Rahma Chaabouni. 2021 · 2021
Closest in time.