Fetching the paper…
Reading the bibliography…
A better understanding of the emergent computation and problem-solving capabilities of recent large language models is of paramount importance to further improve them and broaden their applicability.
Keskar N. S., Mudigere D., Nocedal J., Smelyanskiy M., Tang P. T. P., On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima, 5th International Conference on Learning Representations
2017
Earlier work this paper cites.
Vaswani A., Shazeer N., Parmar N., Uszkoreit J., Jones L., Gomez A., Kaiser L., Polosukhin I., Attention is All you Need, Advances in Neural Information Processing Systems (NIPS)
2017
Earlier work this paper cites.
Liu P., Saleh M., Pot E., Goodrich B., Sepassi R., Kaiser L., Shazeer N., Generating Wikipedia by Summarizing Long Sequences, 6th International Conference on Learning Representations (ICLR)
2018
Earlier work this paper cites.
Schober P., Boer C., Schwarte L., Correlation Coefficients: Appropriate Use and Interpretation, Anesthesia & Analgesia
2018
Earlier work this paper cites.
Hewitt J., Manning C., A Structural Probe for Finding Syntax in Word Representations, Conference of the North American Chapter of the Association for Computational Linguistics
2019
Earlier work this paper cites.
Naik A., Ravichander A., Rose C., Hovy E., Exploring Numeracy in Word Embeddings, 57th Annual Meeting of The Association for Computational Linguistics
2019
Earlier work this paper cites.
Wallace E., Wang Y., Li S., Singh S., Gardner M., Do NLP Models Know Numbers? Probing Numeracy in Embeddings, arXiv: 1909.07940
2019
Earlier work this paper cites.
Kaplan J., McCandlish S., Henighan T., Brown T. B., Chess B., Child R., Gray S., Radford A., Wu J., Amodei D., Scaling Laws for Neural Language Models, arXiv: 2001.08361
2020
Earlier work this paper cites.
Ravfogel S., Elazar Y., Gonen H., Twiton M., Goldberg Y., Null It Out: Guarding Protected Attributes by Iterative Nullspace Projection, 58th Annual Meeting of the Association for Computational Linguistics
2020
Earlier work this paper cites.
Sundararaman D., Si S., Subramanian V., Wang G., Hazarika D., Carin L., Methods for Numeracy-Preserving Word Embeddings, Conference on Empirical Methods in Natural Language Processing (EMNLP)
2020
Earlier work this paper cites.
Elazar Y., Ravfogel S., Jacovi A., Goldberg Y., Amnesic Probing: Behavioral Explanation with Amnesic Counterfactuals, Transactions of The Association for Computational Linguisticss
2021
Cited alongside, same era.
Elhage N., Nanda N., Olsson C., Henighan T., Joseph N., Mann B., Askell A., Bai Y., Chen A., Conerly T., DasSarma N., Drain D., Ganguli D., Hatfield-Dodds Z., Hernandez D., Jones A., Kernion J., Lovitt L., Ndousse K., Amodei D., Brown T., Clark J., Kaplan J., McCandlish S., Olah C., A Mathematical Framework for Transformer Circuits, Transformer Circuits Thread - https://transformer-circuits.pub/2021/framework/index.html
2021
Cited alongside, same era.
Geiger A., Lu H., Icard T. F., Potts C., Causal Abstractions of Neural Networks, 35th Conference on Neural Information Processing Systems (NeurIPS 2021)
2021
Cited alongside, same era.
Nogueira R., Jiang Z., Lin J., Investigating the Limitations of Transformers with Simple Arithmetic Tasks, 1st Mathematical Reasoning in General Artificial Intelligence Workshop @ (ICLR)
Bubeck S., Chandrasekaran V., Eldan R., Gehrke J., Horvitz E., Kamar E., Lee P., Lee Y., Li Y., Lundberg S., Nori H., Palangi H., Ribeiro M., Zhang Y., Sparks of Artificial General Intelligence: Early experiments with GPT-4, arXiv: 2303.12712
2023
Closest in time.
Feng Y., Zhang W., Tu Y., Activity–weight duality in feed-forward neural networks reveals two co-determinants for generalization Nature Machine Intelligence
2023
Closest in time.
Lee N., Sreenivasan K., Lee J., Lee K., Papailiopoulos D., Teaching Arithmetic to Small Transformers, arXiv: 2307.03381
2023
Closest in time.
Liu T., Low B. K. H., Goat: Fine-tuned LLaMA Outperforms GPT-4 on Arithmetic Tasks. arXiv: 2305.14201
2023
Closest in time.
Muffo M., Cocco A., Bertino E., Evaluating Transformer Language Models on Arithmetic Operations Using Number Decomposition, 13th Conference on Language Resources and Evaluation (LREC)
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2021
Cited alongside, same era.
White J., Pimentel T., Saphra N., Cotterell R., A Non-Linear Structural Probe, Conference of the North American Chapter of the Association for Computational Linguistics
2021
Cited alongside, same era.
Belinkov Y., Probing Classifiers: Promises, Shortcomings, and Advances, Computational Linguistics
2022
Cited alongside, same era.
Karpathy A., nanoGPT: a lightweight implementation of medium-sized GPTs, https://github.com/karpathy/nanoGPT
2022
Cited alongside, same era.
Lasri K., Pimentel T., Lenci A., Poibeau T., Cotterell R., Probing for the Usage of Grammatical Number, 60th Annual Meeting of The Association for Computational Linguistics
2022
Cited alongside, same era.
Wei J., Tay Y., Bommasani R., Raffel C., Zoph B., Borgeaud S., Yogatama D., Bosma M., Zhou D., Metzler D., Chi E., Hashimoto T., Vinyals O., Liang P., Dean J., Fedus W., Emergent Abilities of Large Language Models, Transactions on Machine Learning Research (TMLR)
2022
Cited alongside, same era.
Wei J., Wang X., Schuurmans D., Bosma M., Ichter B., Xia F., Chi E., Le Q., Zhou D., Chain-of-Thought Prompting Elicits Reasoning in Large Language Models, arXiv: 2201.11903
2022
Cited alongside, same era.
2023
Closest in time.
Nanda N., Chan L., Lieberum T., Smith J., Steinhardt J., Progress measures for grokking via mechanistic interpretability, arXiv: 2301.05217
2023
Closest in time.
Räuker T., Ho A., Casper S., Hadfield-Menell D., Toward Transparent AI: A Survey on Interpreting the Inner Structures of Deep Neural Networks, arXiv: 2207.13243
2023
Closest in time.
Stolfo A., Belinkov Y., Sachan M., A Mechanistic Interpretation of Arithmetic Reasoning in Language Models using Causal Mediation Analysis, arXiv: 2305.15054
2023
Closest in time.
Yuan Z., Yuan H., Tan C., Wang W., Huang S., How well do Large Language Models perform in Arithmetic tasks?. arXiv: 2304.02015
2023
Closest in time.