Fetching the paper…
Reading the bibliography…
Sequence-to-sequence (seq2seq) learners are widely used, but we still have only limited knowledge about what inductive biases shape the way they generalize.
A formal theory of inductive inference. part i
Ray J Solomonoff · 1964
Earlier work this paper cites.
Aspects of the Theory of Syntax , volume 11
Noam Chomsky · 1965
Earlier work this paper cites.
Modeling by shortest data description
Jorma Rissanen · 1978
Earlier work this paper cites.
Rules and representations: behavioral and brain sciences, 1980
Noam Chomsky · 1980
Earlier work this paper cites.
Handwritten digit recognition with a back-propagation network
Yann LeCun, Bernhard E Boser, John S Denker, Donnie Henderson, Richard E Howard, Wayne E Hubbard, and Lawrence D Jackel · 1990
Earlier work this paper cites.
On the computational power of neural nets
Hava T. Siegelmann and Eduardo D. Sontag · 1992
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Rethinking innateness: A connectionist perspective on development , volume 10
Jeffrey L Elman, Elizabeth A Bates, Mark H Johnson, Annette Karmiloff-Smith, Kim Plunkett, and Domenico Parisi · 1998
Earlier work this paper cites.
Information theory, inference and learning algorithms
David JC MacKay · 2003
Earlier work this paper cites.
A tutorial introduction to the minimum description length principle
Peter Grunwald · 2004
Earlier work this paper cites.
The learnability of abstract syntactic principles
Amy Perfors, Joshua B Tenenbaum, and Terry Regier · 2011
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V Le · 2014
Earlier work this paper cites.
Deep learning
Ian Goodfellow, Yoshua Bengio, and Aaron Courville · 2016
Earlier work this paper cites.
A closer look at memorization in deep networks
Devansh Arpit, Stanisław Jastrzębski, Nicolas Ballas, David Krueger, Emmanuel Bengio, Maxinder S Kanwal, Tegan Maharaj, Asja Fischer, Aaron Courville, Yoshua Bengio, et al · 2017
Cited alongside, same era.
Language modeling with gated convolutional networks
Yann N Dauphin, Angela Fan, Michael Auli, and David Grangier · 2017
Cited alongside, same era.
Convolutional sequence to sequence learning
Jonas Gehring, Michael Auli, David Grangier, Denis Yarats, and Yann N Dauphin · 2017
Cited alongside, same era.
Brenden M Lake and Marco Baroni · 2017
Cited alongside, same era.
Cognitive psychology for deep neural networks: A shape bias case study
Samuel Ritter, David GT Barrett, Adam Santoro, and Matt M Botvinick · 2017
Cited alongside, same era.
On the practical computational power of finite precision rnns for language recognition
Gail Weiss, Yoav Goldberg, and Eran Yahav · 2018
Later among the works it cites.
Word-order biases in deep-agent emergent communication
Rahma Chaabouni, Eugene Kharitonov, Alessandro Lazaric, Emmanuel Dupoux, and Marco Baroni · 2019
Later among the works it cites.
Cnns found to jump around more skillfully than rnns: Compositional generalization in seq2seq convolutional networks
Roberto Dessì and Marco Baroni · 2019
Later among the works it cites.
Joint source-target self attention with locality constraints
José AR Fonollosa, Noe Casas, and Marta R Costa-jussà · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
The marginal value of adaptive gradient methods in machine learning
Ashia C Wilson, Rebecca Roelofs, Mitchell Stern, Nati Srebro, and Benjamin Recht · 2017
Cited alongside, same era.
Jump to better conclusions: SCAN both left and right
Joost Bastings, Marco Baroni, Jason Weston, Kyunghyun Cho, and Douwe Kiela · 2018
Cited alongside, same era.
The description length of deep learning models
Léonard Blier and Yann Ollivier · 2018
Cited alongside, same era.
Optimization methods for large-scale machine learning
Léon Bottou, Frank E Curtis, and Jorge Nocedal · 2018
Cited alongside, same era.
Cognitive science in the era of artificial intelligence: A roadmap for reverse-engineering the infant language-learner
Emmanuel Dupoux · 2018
Cited alongside, same era.
Hierarchical neural story generation
Angela Fan, Mike Lewis, and Yann Dauphin · 2018
Cited alongside, same era.
Brenden M Lake, Tal Linzen, and Marco Baroni · 2019
Later among the works it cites.
fairseq: A fast, extensible toolkit for sequence modeling
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Later among the works it cites.
Lstm networks can perform dynamic counting
Mirac Suzgun, Sebastian Gehrmann, Yonatan Belinkov, and Stuart M Shieber · 2019
Later among the works it cites.
Identity crisis: Memorization and generalization under extreme overparameterization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Michael C Mozer, and Yoram Singer · 2019
Later among the works it cites.
Towards a human-like open-domain chatbot
Daniel Adiwardana, Minh-Thang Luong, David R So, Jamie Hall, Noah Fiedel, Romal Thoppilan, Zi Yang, Apoorv Kulshreshtha, Gaurav Nemade, Yifeng Lu, et al · 2020
Closest in time.
Language models are few-shot learners
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Closest in time.
Compositionality decomposed: How do neural networks generalise?
Dieuwke Hupkes, Verna Dankers, Mathijs Mul, and Elia Bruni · 2020
Closest in time.
R Thomas McCoy, Robert Frank, and Tal Linzen · 2020
Closest in time.
A formal hierarchy of rnn architectures, 2020
William Merrill, Gail Weiss, Yoav Goldberg, Roy Schwartz, Noah A. Smith, and Eran Yahav · 2020
Closest in time.