Towards understanding grokking: An effective theory of representation learning
Z. Liu, O. Kitouni, N. S. Nolte, E. Michaud, M. Tegmark, and M. Williams · 2022
Later among the works it cites.
Locating and editing factual associations in gpt
K. Meng, D. Bau, A. Andonian, and Y. Belinkov · 2022
Later among the works it cites.
Saturated transformers are constant-depth threshold circuits
W. Merrill, A. Sabharwal, and N. A. Smith · 2022
Later among the works it cites.
Rethinking the role of demonstrations: What makes in-context learning work?
S. Min, X. Lyu, A. Holtzman, M. Artetxe, M. Lewis, H. Hajishirzi, and L. Zettlemoyer · 2022
Later among the works it cites.
In-context learning and induction heads
C. Olsson, N. Elhage, N. Nanda, N. Joseph, N. DasSarma, T. Henighan, B. Mann, A. Askell, Y. Bai, A. Chen, T. Conerly, D. Drain, D. Ganguli, Z. Hatfield-Dodds, D. Hernandez, S. Johnston, A. Jones, J. Kernion, L. Lovitt, K. Ndousse, D. Amodei, T. Brown, J. Clark, J. Kaplan, S. McCandlish, and C. Olah · 2022
Later among the works it cites.
Impact of pretraining term frequencies on few-shot reasoning
Y. Razeghi, R. L. Logan IV, M. Gardner, and S. Singh · 2022
Later among the works it cites.
On the effect of pretraining corpora on in-context learning by a large-scale language model
S. Shin, S.-W. Lee, H. Ahn, S. Kim, H. Kim, B. Kim, K. Cho, G. Lee, W. Park, J.-W. Ha, et al · 2022
Later among the works it cites.
An explanation of in-context learning as implicit bayesian inference
S. M. Xie, A. Raghunathan, P. Liang, and T. Ma · 2022
Later among the works it cites.
Unveiling transformers with lego: a synthetic reasoning task
Original
Y. Zhang, A. Backurs, S. Bubeck, R. Eldan, S. Gunasekar, and T. Wagner · 2022
Later among the works it cites.
What learning algorithm is in-context learning? investigations with linear models
E. Akyürek, D. Schuurmans, J. Andreas, T. Ma, and D. Zhou · 2023
Closest in time.
Backward feature correction: How deep learning performs deep learning
Z. Allen-Zhu and Y. Li · 2023
Closest in time.
Towards understanding ensemble, knowledge distillation and self-distillation in deep learning
Z. Allen-Zhu and Y. Li · 2023
Closest in time.
Scaling laws for associative memories
Original
V. Cabannes, E. Dohmatob, and A. Bietti · 2023
Closest in time.
Dissecting recall of factual associations in auto-regressive language models
Original
M. Geva, J. Bastings, K. Filippova, and A. Globerson · 2023
Closest in time.
How do transformers learn topic structure: Towards a mechanistic understanding
Y. Li, Y. Li, and A. Risteski · 2023
Closest in time.
Transformers learn shortcuts to automata
B. Liu, J. T. Ash, S. Goel, A. Krishnamurthy, and C. Zhang · 2023
Closest in time.
Progress measures for grokking via mechanistic interpretability
N. Nanda, L. Chan, T. Liberum, J. Smith, and J. Steinhardt · 2023
Closest in time.
Representational strengths and limitations of transformers
C. Sanford, D. Hsu, and M. Telgarsky · 2023
Closest in time.
Scan and snap: Understanding training dynamics and token composition in 1-layer transformer
Y. Tian, Y. Wang, B. Chen, and S. Du · 2023
Closest in time.
Transformers learn in-context by gradient descent
J. Von Oswald, E. Niklasson, E. Randazzo, J. Sacramento, A. Mordvintsev, A. Zhmoginov, and M. Vladymyrov · 2023
Closest in time.
Interpretability in the wild: a circuit for indirect object identification in gpt-2 small
K. Wang, A. Variengien, A. Conmy, B. Shlegeris, and J. Steinhardt · 2023
Closest in time.