Fetching the paper…
Reading the bibliography…
Hierarchical Reasoning Model (HRM) is a novel approach using two small neural networks recursing at different frequencies.
Beyond regression: New tools for prediction and analysis in the behavioral sciences
Werbos, P · 1974
Earlier work this paper cites.
Une procedure d’apprentissage ponr reseau a seuil asymetrique
LeCun, Y · 1985
Earlier work this paper cites.
Learning internal representations by error propagation
Rumelhart, D. E., Hinton, G. E., and Williams, R. J · 1985
Earlier work this paper cites.
Generalization of backpropagation with application to a recurrent gas market model
Werbos, P. J · 1988
Earlier work this paper cites.
Finding structure in time
Elman, J. L · 1990
Earlier work this paper cites.
The implicit function theorem: history, theory, and applications
Krantz, S. G. and Parks, H. R · 2002
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Root mean square layer normalization
Zhang, B. and Sennrich, R · 2014
Earlier work this paper cites.
Gaussian error linear units (gelus)
Hendrycks, D. and Gimpel, K · 2016
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2017
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Shazeer, N., Mirhoseini, A., Maziarz, K., Davis, A., Le, Q., Hinton, G., and Dean, J · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Large scale gan training for high fidelity natural image synthesis
Brock, A., Donahue, J., and Simonyan, K · 2018
Earlier work this paper cites.
Recurrent relational networks
Palm, R., Paquet, U., and Winther, O · 2018
Cited alongside, same era.
Can convolutional neural networks crack sudoku puzzles?
Park, K · 2018
Cited alongside, same era.
Deep equilibrium models
Bai, S., Kolter, J. Z., and Koltun, V · 2019
Cited alongside, same era.
On the measure of intelligence
Chollet, F · 2019
Cited alongside, same era.
Backpropagation through time and the brain
Lillicrap, T. P. and Santoro, A · 2019
Cited alongside, same era.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Hierarchical graph generation with k2-trees
Jang, Y., Kim, D., and Ahn, S · 2023
Later among the works it cites.
The conceptarc benchmark: Evaluating understanding and generalization in the arc domain
Moskvichev, A., Odouard, V. V., and Mitchell, M · 2023
Later among the works it cites.
Fixed point diffusion models
Bai, X. and Melas-Kyriazi, L · 2024
Later among the works it cites.
Beyond a*: Better planning with transformers via search dynamics bootstrapping
Lehnert, L., Sukhbaatar, S., Su, D., Zheng, Q., Mcvay, P., Rabbat, M., and Tian, Y · 2024
Later among the works it cites.
Scaling llm test-time compute optimally can be more effective than scaling model parameters
Snell, C., Lee, J., Xu, K., and Kumar, A · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Glu variants improve transformer
Shazeer, N · 2020
Cited alongside, same era.
Improved techniques for training score-based generative models
Song, Y. and Ermon, S · 2020
Cited alongside, same era.
Mlp-mixer: An all-mlp architecture for vision
Tolstikhin, I. O., Houlsby, N., Kolesnikov, A., Beyer, L., Zhai, X., Unterthiner, T., Yung, J., Steiner, A., Keysers, D., Uszkoreit, J., et al · 2021
Cited alongside, same era.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
Fedus, W., Zoph, B., and Shazeer, N · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al · 2022
Cited alongside, same era.
Palm: Scaling language modeling with pathways
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al · 2023
Cited alongside, same era.
Roformer: Enhanced transformer with rotary position embedding
Su, J., Ahmed, M., Lu, Y., Pan, S., Bo, W., and Liu, Y · 2024
Later among the works it cites.
The Hidden Drivers of HRM’s Performance on ARC-AGI
ARC Prize Foundation · 2025
Closest in time.
ARC-AGI Leaderboard
ARC Prize Foundation · 2025
Closest in time.
Arc-agi-2: A new challenge for frontier ai reasoning systems
Chollet, F., Knoop, M., Kamradt, G., Landers, B., and Pinkard, H · 2025
Closest in time.
Tdoku: A fast sudoku solver and generator
Dillion, T · 2025
Closest in time.
Grokking at the edge of numerical stability
Prieto, L., Barsbey, M., Mediano, P. A., and Birdal, T · 2025
Closest in time.
Wang, G., Li, J., Sun, Y., Chen, X., Liu, C., Wu, Y., Lu, M., Song, S., and Yadkori, Y. A · 2025
Closest in time.