Fetching the paper…
Reading the bibliography…
We present an approach to modifying Transformer architectures by integrating graph-aware relational reasoning into the attention mechanism, merging concepts from graph neural networks and language modeling.
Hierarchical graph attention network for semi-supervised node classification (2019)
Zhang, M., Cui, Z., Neumann, M. & Chen, Y · 1902
Earlier work this paper cites.
Yun, S., Jeong, M., Kim, R., Kang, J. & Kim, H. J · 1911
Earlier work this paper cites.
Group Extensions and Homology
Eilenberg, S. & MacLane, S · 1942
Earlier work this paper cites.
General theory of natural equivalences
Eilenberg, S. & Mac Lane, S · 1945
Earlier work this paper cites.
Neural networks and physical systems with emergent collective computational abilities
Hopfield, J. J · 1982
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. & Schmidhuber, J · 1997
Earlier work this paper cites.
The information bottleneck method
Tishby, N., Pereira, F. C. & Bialek, W · 2000
Earlier work this paper cites.
Principal neighborhood aggregation for graph nets
Corso, G., Cavalleri, L., Beaini, D., Lio, P. & Velickovic, P · 2004
Earlier work this paper cites.
Reoccurring Patterns in Hierarchical Protein Materials and Music: The Power of Analogies
Giesa, T., Spivak, D. & Buehler, M · 2011
Earlier work this paper cites.
Graph sparsification by effective resistances
Spielman, D. A. & Srivastava, N · 2011
Earlier work this paper cites.
Category theoretic analysis of hierarchical protein materials and social networks
Spivak, D., Giesa, T., Wood, E. & Buehler, M · 2011
Earlier work this paper cites.
Materials by design: Merging proteins and music
Wong, J. et al · 2012
Earlier work this paper cites.
Category theory based solution for the building block replacement problem in materials design
Giesa, T., Spivak, D. & Buehler, M · 2012
Earlier work this paper cites.
Biomateriomics (Springer Netherlands, 2012)
Cranford, S. W. & Buehler, M. J · 2012
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Bahdanau, D., Cho, K. & Bengio, Y · 2014
Earlier work this paper cites.
Graves, A., Wayne, G. & Danihelka, I · 2014
Earlier work this paper cites.
Structure-function-property-design interplay in biopolymers: Spider silk (2014)
Tokareva, O., Jacobsen, M., Buehler, M., Wong, J. & Kaplan, D. L · 2014
Earlier work this paper cites.
Attention is All you Need (2017)
Vaswani, A. et al · 2017
Earlier work this paper cites.
Graph attention networks (2018)
Veličković, P. et al · 2018
Earlier work this paper cites.
Representation learning on graphs: Methods and applications (2018)
Hamilton, W. L., Ying, R. & Leskovec, J · 2018
Earlier work this paper cites.
Multiscale modeling of keratin, collagen, elastin and related human diseases: Perspectives from atomistic to coarse-grained molecular dynamics simulations
Yeo, J. et al · 2018
Earlier work this paper cites.
How powerful are graph neural networks?
Xu, K., Hu, W., Leskovec, J. & Jegelka, S · 2019
Earlier work this paper cites.
On the information bottleneck theory of deep learning
Saxe, A. M. et al · 2019
Earlier work this paper cites.
Coarse-grained model of tropoelastin self-assembly into nascent fibrils
Tarakanova, A., Ozsvar, J., Weiss, A. S. & Buehler, M. J · 2019
Earlier work this paper cites.
Machine Learning of Coarse-Grained Molecular Dynamics Force Fields
Wang, J. et al · 2019
Cited alongside, same era.
Language Models are Few-Shot Learners (2020)
Brown, T. B. et al · 2020
Cited alongside, same era.
A semi-supervised approach to architected materials design using graph neural networks
Guo, K. & Buehler, M · 2020
Cited alongside, same era.
ByT5: Towards a token-free future with pre-trained byte-to-byte models
Xue, L. et al · 2021
Cited alongside, same era.
RoFormer: Enhanced transformer with rotary position embedding
Su, J. et al · 2021
Cited alongside, same era.
Buehler, M. J · 2024
Later among the works it cites.
SciAgents: Automating scientific discovery through bioinspired multi-agent intelligent graph reasoning (2024)
Ghafarollahi, A. & Buehler, M. J · 2024
Later among the works it cites.
Ask, and it shall be given: Turing completeness of prompting (2024)
Qiu, R., Xu, Z., Bao, W. & Tong, H · 2024
Later among the works it cites.
Scaling monosemanticity: Extracting interpretable features from claude 3 sonnet
Templeton, A. et al · 2024
Later among the works it cites.
Interplm: Discovering interpretable features in protein language models via sparse autoencoders
Simon, E. & Zou, J · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hu, E. J. et al · 2021
Cited alongside, same era.
Graph neural networks for materials science and chemistry
Reiser, P. et al · 2022
Cited alongside, same era.
Computational Design and Manufacturing of Sustainable Materials through First-Principles and Materiomics
Shen, S. C. et al · 2022
Cited alongside, same era.
Rapid prediction of protein natural frequencies using graph neural networks
Guo, K. & Buehler, M. J · 2022
Cited alongside, same era.
Hierarchically structured bioinspired nanocomposites
Nepal, D. et al · 2022
Cited alongside, same era.
Star: Bootstrapping reasoning with reasoning (2022)
Zelikman, E., Wu, Y., Mu, J. & Goodman, N. D · 2022
Cited alongside, same era.
General-purpose, long-context autoregressive modeling with perceiver ar (2022)
Hawthorne, C. et al · 2022
Cited alongside, same era.
Contextual feature extraction hierarchies converge in large language models and the brain
Mischler, G., Li, Y. A., Bickel, S. et al · 2024
Later among the works it cites.
Flashattention on a napkin: A diagrammatic approach to deep learning io-awareness (2024)
Abbott, V. & Zardini, G · 2024
Later among the works it cites.
Buehler, E. L. & Buehler, M. J · 2024
Later among the works it cites.
Unveiling the hidden structure of self-attention via kernel principal component analysis (2024)
Teo, R. S. Y. & Nguyen, T. M · 2024
Later among the works it cites.
Your Transformer is Secretly Linear (2024)
Razzhigaev, A. et al · 2024
Later among the works it cites.
Transformers are Multi-state RNNs (2024)
Oren, M., Hassid, M., Yarden, N., Adi, Y. & Schwartz, R · 2024
Later among the works it cites.
Let’s Think Dot by Dot: Hidden Computation in Transformer Language Models (2024)
Pfau, J., Merrill, W. & Bowman, S. R · 2024
Later among the works it cites.
Understanding transformer reasoning capabilities via graph algorithms (2024)
Sanford, C. et al · 2024
Later among the works it cites.
Chain of thought empowers transformers to solve inherently serial problems (2024)
Li, Z., Liu, H., Zhou, D. & Ma, T · 2024
Later among the works it cites.
Zhang, R., Yu, Q., Zang, M., Eickhoff, C. & Pavlick, E · 2024
Later among the works it cites.
Amortized planning with large-scale transformers: A case study on chess (2024)
Ruoss, A. et al · 2024
Later among the works it cites.
Cephalo: Multi-modal vision-language models for bio-inspired materials analysis and design (2024)
Buehler, M. J · 2024
Later among the works it cites.
OpenAI o1 System Card
OpenAI · 2024
Later among the works it cites.
Reasoning with large language models, a survey (2024)
Plaat, A. et al · 2024
Later among the works it cites.
Quiet-STaR: Language models can teach themselves to think before speaking (2024)
Zelikman, E. et al · 2024
Later among the works it cites.
Byte latent transformer: Patches scale better than tokens (2024)
Pagnoni, A. et al · 2024
Later among the works it cites.
Orca-math: Unlocking the potential of SLMs in grade school math (2024)
Mitra, A., Khanpour, H., Rosset, C. & Awadallah, A · 2024
Later among the works it cites.
A survey on LoRA of large language models
Mao, Y. et al · 2025
Closest in time.
Comparing cooperative geometric puzzle solving in ants versus humans
Dreyer, T. et al · 2025
Closest in time.