Fetching the paper…
Reading the bibliography…
Extrapolation to unseen sequence lengths is a challenge for neural generative models of language.
The foundations of arithmetic: A logico-mathematical enquiry into the concept of number, trans. jl austin
Gottlob Frege. 1953 · 1953
Earlier work this paper cites.
Syntactic Structures
Noam Chomsky. 1957 · 1957
Earlier work this paper cites.
Universal grammar
Richard Montague. 1970 · 1970
Earlier work this paper cites.
A tutorial on hidden markov models and selected applications in speech recognition
Lawrence R Rabiner. 1989 · 1989
Earlier work this paper cites.
Head-driven statistical models for natural language processing
Michael Collins. 1999 · 1999
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Learning compositional rules via neural program synthesis
Maxwell Nye, A. Solar-Lezama, J. Tenenbaum, and B. Lake. 2020 · 2003
Earlier work this paper cites.
A benchmark for systematic generalization in grounded language understanding
Laura Ruis, Jacob Andreas, Marco Baroni, Diane Bouchacourt, and Brenden M Lake. 2020 · 2003
Earlier work this paper cites.
On the linguistic capacity of real-time counter automata
William Merrill. 2020 · 2004
Earlier work this paper cites.
A study of compositional generalization in neural models
Tim Klinger, Dhaval Adjodah, Vincent Marois, Josh Joseph, Matthew Riemer, Alex ‘Sandy’ Pentland, and Murray Campbell. 2020 · 2006
Earlier work this paper cites.
Alex Graves, Greg Wayne, and Ivo Danihelka. 2014 · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba. 2015 · 2015
Cited alongside, same era.
Globally coherent text generation with neural checklist models
Chloé Kiddon, Luke Zettlemoyer, and Yejin Choi. 2016 · 2016
Cited alongside, same era.
Why neural translations are the right length
Xing Shi, Kevin Knight, and Deniz Yuret. 2016 · 2016
Cited alongside, same era.
OpenNMT: Open-source toolkit for neural machine translation
Guillaume Klein, Yoon Kim, Yuntian Deng, Jean Senellart, and Alexander Rush. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
What you can cram into a single $&!#* vector: Probing sentence embeddings for linguistic properties
Correcting length bias in neural machine translation
Kenton Murray and David Chiang. 2018 · 2018
Later among the works it cites.
A call for clarity in reporting BLEU scores
Matt Post. 2018 · 2018
Later among the works it cites.
On the practical computational power of finite precision RNNs for language recognition
Gail Weiss, Yoav Goldberg, and Eran Yahav. 2018 · 2018
Later among the works it cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
Permutation equivariant models for compositional generalization in language
Jonathan Gordon, David Lopez-Paz, Marco Baroni, and Diane Bouchacourt. 2019 · 2019
Later among the works it cites.
RNNs can generate bounded hierarchical languages with optimal memory
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Alexis Conneau, German Kruszewski, Guillaume Lample, Loïc Barrault, and Marco Baroni. 2018 · 2018
Cited alongside, same era.
Assessing composition in sentence vector representations
Allyson Ettinger, Ahmed Elgohary, Colin Phillips, and Philip Resnik. 2018 · 2018
Cited alongside, same era.
Brenden Lake and Marco Baroni. 2018 · 2018
Cited alongside, same era.
Rearranging the familiar: Testing compositional generalization in recurrent networks
João Loula, Marco Baroni, and Brenden Lake. 2018 · 2018
Cited alongside, same era.
Deep learning: A critical appraisal
Gary Marcus. 2018 · 2018
Cited alongside, same era.
Jeff Mitchell, Pasquale Minervini, Pontus Stenetorp, and Sebastian Riedel. 2018 · 2018
Cited alongside, same era.
LSTM networks can perform dynamic counting
Mirac Suzgun, Yonatan Belinkov, Stuart Shieber, and Sebastian Gehrmann. 2019a
Cited in the paper.
John Hewitt, Michael Hahn, Surya Ganguli, Percy Liang, and Christopher D. Manning. 2020 · 2019
Later among the works it cites.
Compositional generalization through meta sequence-to-sequence learning
Brenden M Lake. 2019 · 2019
Later among the works it cites.
On evaluating the generalization of LSTM models in formal languages
Mirac Suzgun, Yonatan Belinkov, and Stuart M. Shieber. 2019b · 2019
Later among the works it cites.
Location Attention for Extrapolation to Longer Sequences
Yann Dubois, Gautier Dagan, Dieuwke Hupkes, and Elia Bruni. 2020 · 2020
Closest in time.
Compositionality decomposed: How do neural networks generalise?
Dieuwke Hupkes, Verna Dankers, Mathijs Mul, and Elia Bruni. 2020 · 2020
Closest in time.