Fetching the paper…
Reading the bibliography…
Relational reasoning is a central component of generally intelligent systems, enabling robust and data-efficient inductive generalization.
“What Does BERT Look At? An Analysis of BERT’s Attention”, 2019
Kevin Clark, Urvashi Khandelwal, Omer Levy and Christopher. Manning · 1906
Earlier work this paper cites.
“Do Attention Heads in BERT Track Syntactic Dependencies?”, 2019
Phu Htut, Jason Phang, Shikha Bordia and Samuel. Bowman · 1911
Earlier work this paper cites.
“Representation of a preference ordering by a numerical function”
Gerard Debreu · 1954
Earlier work this paper cites.
“Physical symbol systems”
Allen Newell · 1980
Earlier work this paper cites.
“Structure-mapping: A theoretical framework for analogy”
Dedre Gentner · 1983
Earlier work this paper cites.
“The topography of ability and learning correlations”
Richard Snow, Patrick Kyllonen and Brachia Marshalek · 1984
Earlier work this paper cites.
“Connectionism and cognitive architecture: A critical analysis”
Jerry Fodor and Zenon Pylyshyn · 1988
Earlier work this paper cites.
“The symbol grounding problem”
Stevan Harnad · 1990
Earlier work this paper cites.
“No free lunch theorems for search”, 1995
David Wolpert and William Macready · 1995
Earlier work this paper cites.
“A model of inductive bias learning”
Jonathan Baxter · 2000
Earlier work this paper cites.
“Scaling Laws for Neural Language Models”, 2020
Jared Kaplan et al · 2001
Earlier work this paper cites.
“GLU Variants Improve Transformer”
Noam Shazeer · 2002
Earlier work this paper cites.
“The algebraic mind: Integrating connectionism and cognitive science”
Gary Marcus · 2003
Earlier work this paper cites.
“Language Models are Few-Shot Learners”, 2020
Tom. Brown et al · 2005
Earlier work this paper cites.
“Object-Centric Learning with Slot Attention”
Francesco Locatello et al · 2006
Earlier work this paper cites.
“The discovery of structural form”
Charles Kemp and Joshua Tenenbaum · 2008
Earlier work this paper cites.
“Learning multiple layers of features from tiny images”, 2009
Alex Krizhevsky · 2009
Earlier work this paper cites.
“Analogy and relational reasoning”
Keith Holyoak · 2012
Earlier work this paper cites.
“Emergent Symbols through Binding in External Memory”
Taylor. Webb, Ishan Sinha and Jonathan. Cohen · 2012
Earlier work this paper cites.
“Gaussian Error Linear Units (GELUs)”, 2016
Dan Hendrycks and Kevin Gimpel · 2016
Earlier work this paper cites.
Jimmy Ba, Jamie Kiros and Geoffrey. Hinton · 2016
Earlier work this paper cites.
“Building machines that learn and think like people”
Brenden Lake, Tomer Ullman, Joshua Tenenbaum and Samuel Gershman · 2017
Earlier work this paper cites.
“A Simple Neural Network Module for Relational Reasoning”
Adam Santoro et al · 2017
Earlier work this paper cites.
“Attention Is All You Need”
Ashish Vaswani et al · 2017
Earlier work this paper cites.
“CLEVR: A Diagnostic Dataset for Compositional Language and Elementary Visual Reasoning”
Justin Johnson et al · 2017
Cited alongside, same era.
“Language modeling with gated convolutional networks”
Yann Dauphin, Angela Fan, Michael Auli and David Grangier · 2017
Cited alongside, same era.
“Generalization without systematicity: On the compositional skills of sequence-to-sequence recurrent networks”
Brenden Lake and Marco Baroni · 2018
Cited alongside, same era.
“Measuring abstract reasoning in neural networks”
David Barrett et al · 2018
Cited alongside, same era.
“Relational Recurrent Neural Networks”
Adam Santoro et al · 2018
Cited alongside, same era.
“mixup: Beyond Empirical Risk Minimization”
Hongyi Zhang, Moustapha Cisse, Yann. Dauphin and David Lopez-Paz · 2018
Cited alongside, same era.
“Training Compute-Optimal Large Language Models”, 2022
Jordan Hoffmann et al · 2022
Later among the works it cites.
“Scaling vision transformers”
Xiaohua Zhai, Alexander Kolesnikov, Neil Houlsby and Lucas Beyer · 2022
Later among the works it cites.
“A benchmark for compositional visual reasoning”
Aimen Zerroug et al · 2022
Later among the works it cites.
“Flashattention: Fast and memory-efficient exact attention with IO-awareness”
Tri Dao et al · 2022
Later among the works it cites.
“In-context Learning and Induction Heads” https://transformer-circuits.pub/2022/in-context-learning-and-induction-heads/index.html
Catherine Olsson et al · 2022
Later among the works it cites.
“Towards Reasoning in Large Language Models: A Survey”, 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“The promise of artificial intelligence: reckoning and judgment”
Brian Smith · 2019
Cited alongside, same era.
“Learning to Make Analogies by Contrasting Abstract Relational Structure”
Felix Hill et al · 2019
Cited alongside, same era.
“BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding”
Jacob Devlin, Ming-Wei Chang, Kenton Lee and Kristina Toutanova · 2019
Cited alongside, same era.
“Analysing Mathematical Reasoning Abilities of Neural Models”
David Saxton, Edward Grefenstette, Felix Hill and Pushmeet Kohli · 2019
Cited alongside, same era.
“Cutmix: Regularization strategy to train strong classifiers with localizable features”
Sangdoo Yun et al · 2019
Cited alongside, same era.
“Root mean square layer normalization”
Biao Zhang and Rico Sennrich · 2019
Cited alongside, same era.
Jie Huang and Kevin Chen-Chuan Chang · 2023
Later among the works it cites.
“Emergent analogical reasoning in large language models”
Taylor Webb, Keith Holyoak and Hongjing Lu · 2023
Later among the works it cites.
“RoFormer: Enhanced Transformer with Rotary Position Embedding”
Jianlin Su et al · 2023
Later among the works it cites.
“Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 Small”
Kevin Wang et al · 2023
Later among the works it cites.
“Llama 2: Open Foundation and Fine-Tuned Chat Models”
Hugo Touvron et al · 2023
Later among the works it cites.
“GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints”
Joshua Ainslie et al · 2023
Later among the works it cites.
“TinyStories: How Small Can Language Models Be and Still Speak Coherent English?”
Ronen Eldan and Yuanzhi Li · 2023
Later among the works it cites.
“Abstractors and relational cross-attention: An inductive bias for explicit relational reasoning in Transformers”
Awni Altabaa, Taylor Webb, Jonathan. Cohen and John Lafferty · 2024
Closest in time.
“Learning Hierarchical Relational Representations through Relational Convolutions”
Awni Altabaa and John Lafferty · 2024
Closest in time.
“The Relational Bottleneck as an Inductive Bias for Efficient Abstraction”
Taylor. Webb et al · 2024
Closest in time.
“Large Language Models Cannot Self-Correct Reasoning Yet”, 2024
Jie Huang et al · 2024
Closest in time.
Iman Mirzadeh et al · 2024
Closest in time.
“The Llama 3 Herd of Models”, 2024
Aaron Grattafiori et al · 2024
Closest in time.
Bingchen Zhao, Yongshuo Zong, Letian Zhang and Timothy Hospedales · 2024
Closest in time.
“FineWeb-Edu”, 2024
Anton Lozhkov, Loubna Ben, Leandro von Werra and Thomas Wolf · 2024
Closest in time.
“Approximation of Relation Functions and Attention Mechanisms”
Awni Altabaa and John Lafferty · 2024
Closest in time.
“The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale”, 2024
Guilherme Penedo et al · 2024
Closest in time.
“Self-Attention with Relative Position Representations”
Peter Shaw, Jakob Uszkoreit and Ashish Vaswani · 2074
Closest in time.