Fetching the paper…
Reading the bibliography…
Can Transformers predict new syllogisms by composing established ones? More generally, what type of targets can be learned by such models from scratch? Recent works show that Transformers can be Turing-complete in terms of expressivity, but this does not address the learnability objective.
Efficient noise-tolerant learning from statistical queries
Michael Kearns · 1998
Earlier work this paper cites.
Curriculum learning
Yoshua Bengio, Jérôme Louradour, Ronan Collobert, and Jason Weston · 2009
Earlier work this paper cites.
Algorithms for learning kernels based on centered alignment
Corinna Cortes, Mehryar Mohri, and Afshin Rostamizadeh · 2012
Earlier work this paper cites.
Understanding Machine Learning: From Theory to Algorithms
Shai Shalev-Shwartz and Shai Ben-David · 2014
Earlier work this paper cites.
Analysis of Boolean Functions
Ryan O’Donnell · 2014
Earlier work this paper cites.
Wojciech Zaremba and Ilya Sutskever · 2014
Earlier work this paper cites.
A general characterization of the statistical query complexity
Vitaly Feldman · 2016
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization, 2016
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2016
Earlier work this paper cites.
Adaptive computation time for recurrent neural networks
Alex Graves · 2016
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
Justin Johnson, Bharath Hariharan, Laurens Van Der Maaten, Li Fei-Fei, C Lawrence Zitnick, and Ross Girshick · 2017
Earlier work this paper cites.
Fixing weight decay regularization in adam
Ilya Loshchilov and Frank Hutter · 2017
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Earlier work this paper cites.
Generalization without systematicity: On the compositional skills of sequence-to-sequence recurrent networks
Brenden Lake and Marco Baroni · 2018
Earlier work this paper cites.
Self-attention with relative position representations
Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani · 2018
Earlier work this paper cites.
Listops: A diagnostic dataset for latent tree learning
Nikita Nangia and Samuel R. Bowman · 2018
Earlier work this paper cites.
Analysing mathematical reasoning abilities of neural models
David Saxton, Edward Grefenstette, Felix Hill, and Pushmeet Kohli · 2019
Earlier work this paper cites.
Phyre: A new benchmark for physical reasoning
Anton Bakhtin, Laurens van der Maaten, Justin Johnson, Laura Gustafson, and Ross Girshick · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners, 2019
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Earlier work this paper cites.
Transformer-xl: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime G. Carbonell, Quoc V. Le, and Ruslan Salakhutdinov · 2019
Earlier work this paper cites.
Universal transformers, 2019
Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, and Łukasz Kaiser · 2019
Earlier work this paper cites.
Drop: A reading comprehension benchmark requiring discrete reasoning over paragraphs, 2019
Dheeru Dua, Yizhong Wang, Pradeep Dasigi, Gabriel Stanovsky, Sameer Singh, and Matt Gardner · 2019
Earlier work this paper cites.
Clutrr: A diagnostic benchmark for inductive reasoning from text, 2019
Koustuv Sinha, Shagun Sodhani, Jin Dong, Joelle Pineau, and William L. Hamilton · 2019
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Earlier work this paper cites.
fairseq: A fast, extensible toolkit for sequence modeling
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli · 2019
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Earlier work this paper cites.
On the universality of deep learning
Emmanuel Abbe and Colin Sandon · 2020
Earlier work this paper cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Earlier work this paper cites.
What neural networks memorize and why: Discovering the long tail via influence estimation
Vitaly Feldman and Chiyuan Zhang · 2020
Earlier work this paper cites.
Compositionality decomposed: How do neural networks generalise?
Dieuwke Hupkes, Verna Dankers, Mathijs Mul, and Elia Bruni · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu · 2020
Earlier work this paper cites.
Logiqa: A challenge dataset for machine reading comprehension with logical reasoning, 2020
Jian Liu, Leyang Cui, Hanmeng Liu, Dandan Huang, Yile Wang, and Yue Zhang · 2020
Earlier work this paper cites.
Chiyuan Zhang, Maithra Raghu, Jon M. Kleinberg, and Samy Bengio · 2021
Earlier work this paper cites.
Online stochastic gradient descent on non-convex losses from high-dimensional inference
Gerard Ben Arous, Reza Gheissari, and Aukosh Jagannath · 2021
Cited alongside, same era.
On the power of differentiable learning versus PAC and SQ learning
Emmanuel Abbe, Pritish Kamath, Eran Malach, Colin Sandon, and Nathan Srebro · 2021
Cited alongside, same era.
Show your work: Scratchpads for intermediate computation with language models, 2021
Maxwell Nye, Anders Johan Andreassen, Guy Gur-Ari, Henryk Michalewski, Jacob Austin, David Bieber, David Dohan, Aitor Lewkowycz, Maarten Bosma, David Luan, Charles Sutton, and Augustus Odena · 2021
Cited alongside, same era.
The devil is in the detail: Simple tricks improve systematic generalization of transformers
Róbert Csordás, Kazuki Irie, and Jürgen Schmidhuber · 2021
Cited alongside, same era.
Star: A benchmark for situated reasoning in real-world videos
Bo Wu, Shoubin Yu, Zhenfang Chen, Joshua B Tenenbaum, and Chuang Gan · 2021
Cited alongside, same era.
On Single Index Models beyond Gaussian Data
Joan Bruna, Loucas Pillaud-Vivien, and Aaron Zweig · 2023
Later among the works it cites.
Sgd learning on neural networks: leap complexity and saddle-to-saddle dynamics
Emmanuel Abbe, Enric Boix Adsera, and Theodor Misiakiewicz · 2023
Later among the works it cites.
The expressive power of transformers with chain of thought
William Merrill and Ashish Sabharwal · 2023
Later among the works it cites.
Auto-Regressive Next-Token Predictors are Universal Learners
Eran Malach · 2023
Later among the works it cites.
The impact of positional encoding on length generalization in transformers
Amirhossein Kazemnejad, Inkit Padhi, Karthikeyan Natesan Ramamurthy, Payel Das, and Siva Reddy · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Do transformers really perform bad for graph representation?, 2021
Chengxuan Ying, Tianle Cai, Shengjie Luo, Shuxin Zheng, Guolin Ke, Di He, Yanming Shen, and Tie-Yan Liu · 2021
Cited alongside, same era.
Knowledge graph based synthetic corpus generation for knowledge-enhanced language model pre-training, 2021
Oshin Agarwal, Heming Ge, Siamak Shakeri, and Rami Al-Rfou · 2021
Cited alongside, same era.
Thinking like transformers
Gail Weiss, Yoav Goldberg, and Eran Yahav · 2021
Cited alongside, same era.
Quantifying the benefit of using differentiable learning over tangent kernels
Eran Malach, Pritish Kamath, Emmanuel Abbe, and Nathan Srebro · 2021
Cited alongside, same era.
Revisiting neural scaling laws in language and vision
Ibrahim Alabdulmohsin, Behnam Neyshabur, and Xiaohua Zhai · 2022
Cited alongside, same era.
Solving quantitative reasoning problems with language models
Aitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer, Henryk Michalewski, Vinay Ramasesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, et al · 2022
Cited alongside, same era.
The clrs algorithmic reasoning benchmark
Petar Veličković, Adrià Puigdomènech Badia, David Budden, Razvan Pascanu, Andrea Banino, Misha Dashevskiy, Raia Hadsell, and Charles Blundell · 2022
Cited alongside, same era.
Later among the works it cites.
What algorithms can transformers learn? a study in length generalization
Hattie Zhou, Arwen Bradley, Etai Littwin, Noam Razin, Omid Saremi, Josh Susskind, Samy Bengio, and Preetum Nakkiran · 2023
Later among the works it cites.
Length generalization in arithmetic transformers
Samy Jelassi, Stéphane d’Ascoli, Carles Domingo-Enrich, Yuhuai Wu, Yuanzhi Li, and François Charton · 2023
Later among the works it cites.
Generalization on the unseen, logic reasoning and degree curriculum
Emmanuel Abbe, Samy Bengio, Aryo Lotfi, and Kevin Rizk · 2023
Later among the works it cites.
Looped transformers as programmable computers
Angeliki Giannou, Shashank Rajput, Jy-Yong Sohn, Kangwook Lee, Jason D. Lee, and Dimitris Papailiopoulos · 2023
Later among the works it cites.
Chain-of-thought prompting elicits reasoning in large language models, 2023
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, and Denny Zhou · 2023
Later among the works it cites.
Large language models are zero-shot reasoners, 2023
Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa · 2023
Later among the works it cites.
Functional interpolation for relative positions improves long context transformers
Shanda Li, Chong You, Guru Guruganesh, Joshua Ainslie, Santiago Ontanon, Manzil Zaheer, Sumit Sanghai, Yiming Yang, Sanjiv Kumar, and Srinadh Bhojanapalli · 2023
Later among the works it cites.
Randomized positional encodings boost length generalization of transformers
Anian Ruoss, Grégoire Delétang, Tim Genewein, Jordi Grau-Moya, Róbert Csordás, Mehdi Bennani, Shane Legg, and Joel Veness · 2023
Later among the works it cites.
Talk like a graph: Encoding graphs for large language models, 2023
Bahare Fatemi, Jonathan Halcrow, and Bryan Perozzi · 2023
Later among the works it cites.
Neural compositional rule learning for knowledge graph reasoning, 2023
Kewei Cheng, Nesreen K. Ahmed, and Yizhou Sun · 2023
Later among the works it cites.
Linc: A neurosymbolic approach for logical reasoning by combining language models with first-order logic provers
Theo Olausson, Alex Gu, Ben Lipkin, Cedegao Zhang, Armando Solar-Lezama, Joshua Tenenbaum, and Roger Levy · 2023
Later among the works it cites.
Patton: Language model pretraining on text-rich networks, 2023
Bowen Jin, Wentao Zhang, Yu Zhang, Yu Meng, Xinyang Zhang, Qi Zhu, and Jiawei Han · 2023
Later among the works it cites.
A mathematical model for curriculum learning
Elisabetta Cornacchia and Elchanan Mossel · 2023
Later among the works it cites.
Why are sensitive functions hard for transformers?
Michael Hahn and Mark Rofin · 2024
Closest in time.
Computational-Statistical Gaps in Gaussian Single-Index Models
Alex Damian, Loucas Pillaud-Vivien, Jason D. Lee, and Joan Bruna · 2024
Closest in time.
Transformers can achieve length generalization but not robustly
Yongchao Zhou, Uri Alon, Xinyun Chen, Xuezhi Wang, Rishabh Agarwal, and Denny Zhou · 2024
Closest in time.
Memorization capacity of multi-head attention in transformers
Sadegh Mahdavi, Renjie Liao, and Christos Thrampoulidis · 2024
Closest in time.
Teaching arithmetic to small transformers
Nayoung Lee, Kartik Sreenivasan, Jason D. Lee, Kangwook Lee, and Dimitris Papailiopoulos · 2024
Closest in time.
Roformer: Enhanced transformer with rotary position embedding
Jianlin Su, Murtadha Ahmed, Yu Lu, Shengfeng Pan, Wen Bo, and Yunfeng Liu · 2024
Closest in time.
Learning to reason and memorize with self-notes
Jack Lanchantin, Shubham Toshniwal, Jason Weston, Sainbayar Sukhbaatar, et al · 2024
Closest in time.
Think before you speak: Training language models with pause tokens
Sachin Goyal, Ziwei Ji, Ankit Singh Rawat, Aditya Krishna Menon, Sanjiv Kumar, and Vaishnavh Nagarajan · 2024
Closest in time.
Enhancing chain-of-thoughts prompting with iterative bootstrapping in large language models, 2024
Jiashuo Sun, Yi Luo, Yeyun Gong, Chen Lin, Yelong Shen, Jian Guo, and Nan Duan · 2024
Closest in time.
Towards revealing the mystery behind chain of thought: a theoretical perspective
Guhao Feng, Bohang Zhang, Yuntian Gu, Haotian Ye, Di He, and Liwei Wang · 2024
Closest in time.
Iteration head: A mechanistic study of chain-of-thought
Vivien Cabannes, Charles Arnal, Wassim Bouaziz, Alice Yang, Francois Charton, and Julia Kempe · 2024
Closest in time.
Can language models solve graph problems in natural language?, 2024
Heng Wang, Shangbin Feng, Tianxing He, Zhaoxuan Tan, Xiaochuang Han, and Yulia Tsvetkov · 2024
Closest in time.
Understanding transformer reasoning capabilities via graph algorithms
Clayton Sanford, Bahare Fatemi, Ethan Hall, Anton Tsitsulin, Mehran Kazemi, Jonathan Halcrow, Bryan Perozzi, and Vahab Mirrokni · 2024
Closest in time.
Provable advantage of curriculum learning on parity targets with mixed inputs
Emmanuel Abbe, Elisabetta Cornacchia, and Aryo Lotfi · 2024
Closest in time.
Keyformer: Kv cache reduction through key tokens selection for efficient generative inference
Muhammad Adnan, Akhil Arunkumar, Gaurav Jain, Prashant Nair, Ilya Soloveychik, and Purushotham Kamath · 2024
Closest in time.