Fetching the paper…
Reading the bibliography…
Out-of-distribution generalization capabilities of sequence-to-sequence models can be studied from the lens of two crucial forms of generalization: length generalization -- the ability to generalize to longer sequences than ones seen during training, and compositional generalization: the ability to generalize to token combinations not seen during training.
Pragmatics and intensional logic
Richard Montague · 1970
Earlier work this paper cites.
Connectionism and cognitive architecture: A critical analysis
Jerry A Fodor and Zenon W Pylyshyn · 1988
Earlier work this paper cites.
Mapping part-whole hierarchies into connectionist networks
Geoffrey E Hinton · 1990
Earlier work this paper cites.
Holographic reduced representations: Convolution algebra for compositional distributed representations
Tony Plate et al · 1991
Earlier work this paper cites.
Neural nets as systems models and controllers
Eduardo D Sontag · 1992
Earlier work this paper cites.
Probability and measure theory
Robert B Ash and Catherine A Doléans-Dade · 2000
Earlier work this paper cites.
Covariate shift adaptation by importance weighted cross validation
Masashi Sugiyama, Matthias Krauledat, and Klaus-Robert Müller · 2007
Earlier work this paper cites.
A new learning paradigm: Learning using privileged information
Vladimir Vapnik and Akshay Vashist · 2009
Earlier work this paper cites.
Impossibility theorems for domain adaptation
Shai Ben David, Tyler Lu, Teresa Luu, and Dávid Pál · 2010
Earlier work this paper cites.
Domain adaptation–can quantity compensate for quality?
Shai Ben-David and Ruth Urner · 2014
Earlier work this paper cites.
Geometric measure theory
Herbert Federer · 2014
Earlier work this paper cites.
Understanding machine learning: From theory to algorithms
Shai Shalev-Shwartz and Shai Ben-David · 2014
Earlier work this paper cites.
The zero set of a real analytic function
Boris Mityagin · 2015
Earlier work this paper cites.
Collaborative pac learning
Avrim Blum, Nika Haghtalab, Ariel D Procaccia, and Mingda Qiao · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Deep sets
Manzil Zaheer, Satwik Kottur, Siamak Ravanbakhsh, Barnabas Poczos, Russ R Salakhutdinov, and Alexander J Smola · 2017
Earlier work this paper cites.
Generalization without systematicity: On the compositional skills of sequence-to-sequence recurrent networks
Brenden Lake and Marco Baroni · 2018
Earlier work this paper cites.
Rearranging the familiar: Testing compositional generalization in recurrent networks
Joao Loula, Marco Baroni, and Brenden M Lake · 2018
Earlier work this paper cites.
Invariant models for causal transfer learning
Mateo Rojas-Carulla, Bernhard Schölkopf, Richard Turner, and Jonas Peters · 2018
Earlier work this paper cites.
Self-attention with relative position representations
Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani · 2018
Earlier work this paper cites.
Martin Arjovsky, Léon Bottou, Ishaan Gulrajani, and David Lopez-Paz · 2019
Earlier work this paper cites.
Permutation equivariant models for compositional generalization in language
Jonathan Gordon, David Lopez-Paz, Marco Baroni, and Diane Bouchacourt · 2019
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2019
Cited alongside, same era.
Shiori Sagawa, Pang Wei Koh, Tatsunori B Hashimoto, and Percy Liang · 2019
Cited alongside, same era.
Expressivity of deep neural networks
Ingo Gühring, Mones Raslan, and Gitta Kutyniok · 2020
Cited alongside, same era.
Compositionality decomposed: How do neural networks generalise?
Dieuwke Hupkes, Verna Dankers, Mathijs Mul, and Elia Bruni · 2020
Cited alongside, same era.
Variational autoencoders and nonlinear ica: A unifying framework
Ilyes Khemakhem, Diederik Kingma, Ricardo Monti, and Aapo Hyvarinen · 2020
Cited alongside, same era.
Suriya Gunasekar, Yi Zhang, Jyoti Aneja, Caio César Teodoro Mendes, Allie Del Giorno, Sivakanth Gopi, Mojan Javaheripi, Piero Kauffmann, Gustavo de Rosa, Olli Saarikivi, et al · 2023
Later among the works it cites.
Length generalization in arithmetic transformers
Samy Jelassi, Stéphane d’Ascoli, Carles Domingo-Enrich, Yuhuai Wu, Yuanzhi Li, and François Charton · 2023
Later among the works it cites.
Additive decoders for latent variables identification and cartesian-product extrapolation
Sébastien Lachapelle, Divyat Mahajan, Ioannis Mitliagkas, and Simon Lacoste-Julien · 2023
Later among the works it cites.
A survey on compositional generalization in applications
Baihan Lin, Djallel Bouneffouf, and Irina Rish · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cogs: A compositional generalization challenge based on semantic interpretation
Najoung Kim and Tal Linzen · 2020
Cited alongside, same era.
Invariance principle meets information bottleneck for out-of-distribution generalization
Kartik Ahuja, Ethan Caballero, Dinghuai Zhang, Jean-Christophe Gagnon-Audet, Yoshua Bengio, Ioannis Mitliagkas, and Irina Rish · 2021
Cited alongside, same era.
Efficiently modeling long sequences with structured state spaces
Albert Gu, Karan Goel, and Christopher Ré · 2021
Cited alongside, same era.
On linear identifiability of learned representations
Geoffrey Roeder, Luke Metz, and Durk Kingma · 2021
Cited alongside, same era.
Exploring length generalization in large language models
Cem Anil, Yuhuai Wu, Anders Andreassen, Aitor Lewkowycz, Vedant Misra, Vinay Ramasesh, Ambrose Slone, Guy Gur-Ari, Ethan Dyer, and Behnam Neyshabur · 2022
Cited alongside, same era.
Neural networks and the chomsky hierarchy
Grégoire Delétang, Anian Ruoss, Jordi Grau-Moya, Tim Genewein, Li Kevin Wenliang, Elliot Catt, Chris Cundy, Marcus Hutter, Shane Legg, Joel Veness, et al · 2022
Cited alongside, same era.
First steps toward understanding the extrapolation of nonlinear models to unseen domains
Kefan Dong and Tengyu Ma · 2022
Cited alongside, same era.
Aviv Netanyahu, Abhishek Gupta, Max Simchowitz, Kaiqing Zhang, and Pulkit Agrawal · 2023
Later among the works it cites.
Discovering modular solutions that generalize compositionally
Simon Schug, Seijin Kobayashi, Yassir Akram, Maciej Wołczyk, Alexandra Proca, Johannes Von Oswald, Razvan Pascanu, João Sacramento, and Angelika Steger · 2023
Later among the works it cites.
A study on relu and softmax in transformer
Kai Shen, Junliang Guo, Xu Tan, Siliang Tang, Rui Wang, and Jiang Bian · 2023
Later among the works it cites.
Engression: Extrapolation for nonlinear regression?
Xinwei Shen and Nicolai Meinshausen · 2023
Later among the works it cites.
Gpt-4 doesn’t know it’s wrong: An analysis of iterative prompting for reasoning problems
Kaya Stechly, Matthew Marquez, and Subbarao Kambhampati · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Later among the works it cites.
Can large language models really improve by self-critiquing their own plans?
Karthik Valmeekam, Matthew Marquez, and Subbarao Kambhampati · 2023
Later among the works it cites.
Replacing softmax with relu in vision transformers
Mitchell Wortsman, Jaehoon Lee, Justin Gilmer, and Simon Kornblith · 2023
Later among the works it cites.
Conditions for length generalization in learning reasoning skills
Changnan Xiao and Bing Liu · 2023
Later among the works it cites.
What algorithms can transformers learn? a study in length generalization
Hattie Zhou, Arwen Bradley, Etai Littwin, Noam Razin, Omid Saremi, Josh Susskind, Samy Bengio, and Preetum Nakkiran · 2023
Later among the works it cites.
Generalization on the unseen, logic reasoning and degree curriculum
Emmanuel Abbe, Samy Bengio, Aryo Lotfi, and Kevin Rizk · 2024
Closest in time.
Faith and fate: Limits of transformers on compositionality
Nouha Dziri, Ximing Lu, Melanie Sclar, Xiang Lorraine Li, Liwei Jiang, Bill Yuchen Lin, Sean Welleck, Peter West, Chandra Bhagavatula, Ronan Le Bras, et al · 2024
Closest in time.
Spuriosity didn’t kill the classifier: Using invariant predictions to harness spurious features
Cian Eastwood, Shashank Singh, Andrei L Nicolicioiu, Marin Vlastelica Pogančić, Julius von Kügelgen, and Bernhard Schölkopf · 2024
Closest in time.
Universal length generalization with turing programs
Kaiying Hou, David Brandfonbrener, Sham Kakade, Samy Jelassi, and Eran Malach · 2024
Closest in time.
A formal framework for understanding length generalization in transformers
Xinting Huang, Andy Yang, Satwik Bhattamishra, Yash Sarrof, Andreas Krebs, Hattie Zhou, Preetum Nakkiran, and Michael Hahn · 2024
Closest in time.
The impact of positional encoding on length generalization in transformers
Amirhossein Kazemnejad, Inkit Padhi, Karthikeyan Natesan Ramamurthy, Payel Das, and Siva Reddy · 2024
Closest in time.
A survey on compositional learning of ai models: Theoretical and experimetnal practices
Sania Sinha, Tanawan Premsri, and Parisa Kordjamshidi · 2024
Closest in time.