Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have revolutionized the field of Natural Language Processing thanks to their ability to reuse knowledge acquired on massive text corpora on a wide variety of downstream tasks, with minimal (if any) tuning steps.
Randomized positional encodings boost length generalization of transformers
Anian Ruoss, Grégoire Delétang, Tim Genewein, Jordi Grau-Moya, Róbert Csordás, Mehdi Bennani, Shane Legg, and Joel Veness. 2023 · 1903
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Generalization without systematicity: On the compositional skills of sequence-to-sequence recurrent networks
Brenden M. Lake and Marco Baroni. 2018 · 2018
Earlier work this paper cites.
Memorize or generalize? searching for a compositional RNN in a haystack
Adam Liska, Germán Kruszewski, and Marco Baroni. 2018 · 2018
Earlier work this paper cites.
Listops: A diagnostic dataset for latent tree learning
Nikita Nangia and Samuel R. Bowman. 2018 · 2018
Earlier work this paper cites.
Language models are few-shot learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 2020
Earlier work this paper cites.
Compositionality decomposed: How do neural networks generalise?
Dieuwke Hupkes, Verna Dankers, Mathijs Mul, and Elia Bruni. 2020 · 2020
Earlier work this paper cites.
COGS: A compositional generalization challenge based on semantic interpretation
Najoung Kim and Tal Linzen. 2020 · 2020
Earlier work this paper cites.
The devil is in the detail: Simple tricks improve systematic generalization of transformers
Róbert Csordás, Kazuki Irie, and Jürgen Schmidhuber. 2021 · 2021
Earlier work this paper cites.
Show your work: Scratchpads for intermediate computation with language models
Maxwell I. Nye, Anders Johan Andreassen, Guy Gur-Ari, Henryk Michalewski, Jacob Austin, David Bieber, David Dohan, Aitor Lewkowycz, Maarten Bosma, David Luan, Charles Sutton, and Augustus Odena. 2021 · 2021
Earlier work this paper cites.
Exploring length generalization in large language models
Cem Anil, Yuhuai Wu, Anders Andreassen, Aitor Lewkowycz, Vedant Misra, Vinay V. Ramasesh, Ambrose Slone, Guy Gur-Ari, Ethan Dyer, and Behnam Neyshabur. 2022 · 2022
Cited alongside, same era.
CTL++: evaluating generalization on never-seen compositional patterns of known functions, and compatibility of neural representations
Róbert Csordás, Kazuki Irie, and Jürgen Schmidhuber. 2022a · 2022
Cited alongside, same era.
The neural data router: Adaptive control flow in transformers improves systematic generalization
Róbert Csordás, Kazuki Irie, and Jürgen Schmidhuber. 2022b · 2022
Cited alongside, same era.
Large language models are zero-shot reasoners
Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022 · 2022
Cited alongside, same era.
Systematic generalization and emergent structures in transformers trained on structured tasks
Yuxuan Li and James L. McClelland. 2022 · 2022
Star: Bootstrapping reasoning with reasoning
Eric Zelikman, Yuhuai Wu, Jesse Mu, and Noah D. Goodman. 2022 · 2022
Later among the works it cites.
Teaching algorithmic reasoning via in-context learning
Hattie Zhou, Azade Nova, Hugo Larochelle, Aaron C. Courville, Behnam Neyshabur, and Hanie Sedghi. 2022 · 2022
Later among the works it cites.
Length generalization in arithmetic transformers
Samy Jelassi, Stéphane d’Ascoli, Carles Domingo-Enrich, Yuhuai Wu, Yuanzhi Li, and François Charton. 2023 · 2023
Later among the works it cites.
The impact of positional encoding on length generalization in transformers
Amirhossein Kazemnejad, Inkit Padhi, Karthikeyan Natesan Ramamurthy, Payel Das, and Siva Reddy. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Out-of-distribution generalization in algorithmic reasoning through curriculum learning
Andrew Joohun Nam, Mustafa Abdool, Trevor Maxfield, and James L. McClelland. 2022 · 2022
Cited alongside, same era.
Making transformers solve compositional tasks
Santiago Ontañón, Joshua Ainslie, Zachary Fisher, and Vaclav Cvicek. 2022 · 2022
Cited alongside, same era.
Revisiting the compositional generalization abilities of neural sequence models
Arkil Patel, Satwik Bhattamishra, Phil Blunsom, and Navin Goyal. 2022 · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V. Le, and Denny Zhou. 2022 · 2022
Cited alongside, same era.
Aobo Kong, Shiwan Zhao, Hao Chen, Qicheng Li, Yong Qin, Ruiqi Sun, and Xin Zhou. 2023 · 2023
Later among the works it cites.
Learning to reason and memorize with self-notes
Jack Lanchantin, Shubham Toshniwal, Jason Weston, Arthur Szlam, and Sainbayar Sukhbaatar. 2023 · 2023
Later among the works it cites.
Improving mathematics tutoring with A code scratchpad
Shriyash Upadhyay, Etan Ginsberg, and Chris Callison-Burch. 2023 · 2023
Later among the works it cites.
Self-consistency improves chain of thought reasoning in language models
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le, Ed H. Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2023 · 2023
Later among the works it cites.
Least-to-most prompting enables complex reasoning in large language models
Denny Zhou, Nathanael Schärli, Le Hou, Jason Wei, Nathan Scales, Xuezhi Wang, Dale Schuurmans, Claire Cui, Olivier Bousquet, Quoc V. Le, and Ed H. Chi. 2023 · 2023
Later among the works it cites.