Fetching the paper…
Reading the bibliography…
Despite the impressive performance in a variety of complex tasks, modern large language models (LLMs) still have trouble dealing with some math problems that are simple and intuitive for humans, such as addition.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I · 2019
Earlier work this paper cites.
Language models are few-shot learners, 2020
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Earlier work this paper cites.
Are transformers universal approximators of sequence-to-sequence functions?, 2020
Yun, C., Bhojanapalli, S., Rawat, A. S., Reddi, S. J., and Kumar, S · 2020
Earlier work this paper cites.
Show your work: Scratchpads for intermediate computation with language models, 2021
Nye, M., Andreassen, A. J., Gur-Ari, G., Michalewski, H., Austin, J., Bieber, D., Dohan, D., Lewkowycz, A., Bosma, M., Luan, D., Sutton, C., and Odena, A · 2021
Earlier work this paper cites.
Exploring length generalization in large language models, 2022
Anil, C., Wu, Y., Andreassen, A., Lewkowycz, A., Misra, V., Ramasesh, V., Slone, A., Gur-Ari, G., Dyer, E., and Neyshabur, B · 2022
Earlier work this paper cites.
Towards understanding grokking: An effective theory of representation learning, 2022
Liu, Z., Kitouni, O., Nolte, N., Michaud, E. J., Tegmark, M., and Williams, M · 2022
Earlier work this paper cites.
Introducing chatgpt, 2022
OpenAI · 2022
Earlier work this paper cites.
Grokking: Generalization beyond overfitting on small algorithmic datasets, 2022
Power, A., Burda, Y., Edwards, H., Babuschkin, I., and Misra, V · 2022
Earlier work this paper cites.
Limitations of language models in arithmetic and symbolic induction, 2022
Qian, J., Wang, H., Li, Z., Li, S., and Yan, X · 2022
Earlier work this paper cites.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Srivastava, A., Rastogi, A., Rao, A., Shoeb, A. A. M., Abid, A., Fisch, A., Brown, A. R., Santoro, A., Gupta, A., Garriga-Alonso, A., et al · 2022
Earlier work this paper cites.
Lamda: Language models for dialog applications
Thoppilan, R., De Freitas, D., Hall, J., Shazeer, N., Kulshreshtha, A., Cheng, H.-T., Jin, A., Bos, T., Baker, L., Du, Y., et al · 2022
Earlier work this paper cites.
Emergent abilities of large language models, 2022
Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S., Yogatama, D., Bosma, M., Zhou, D., Metzler, D., Chi, E. H., Hashimoto, T., Vinyals, O., Liang, P., Dean, J., and Fedus, W · 2022
Earlier work this paper cites.
Teaching algorithmic reasoning via in-context learning, 2022
Zhou, H., Nova, A., Larochelle, H., Courville, A., Neyshabur, B., and Sedghi, H · 2022
Earlier work this paper cites.
Generalization on the unseen, logic reasoning and degree curriculum, 2023
Abbe, E., Bengio, S., Lotfi, A., and Rizk, K · 2023
Earlier work this paper cites.
What learning algorithm is in-context learning? investigations with linear models, 2023
Akyürek, E., Schuurmans, D., Andreas, J., Ma, T., and Zhou, D · 2023
Earlier work this paper cites.
Chatgpt is a knowledgeable but inexperienced solver: An investigation of commonsense problem in large language models, 2023
Bian, N., Han, X., Sun, L., Lin, H., Lu, Y., and He, B · 2023
Cited alongside, same era.
Sparks of artificial general intelligence: Early experiments with gpt-4, 2023
Bubeck, S., Chandrasekaran, V., Eldan, R., Gehrke, J., Horvitz, E., Kamar, E., Lee, P., Lee, Y. T., Li, Y., Lundberg, S., Nori, H., Palangi, H., Ribeiro, M. T., and Zhang, Y · 2023
Cited alongside, same era.
Palm: Scaling language modeling with pathways
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al · 2023
Cited alongside, same era.
Why can gpt learn in-context? language models implicitly perform gradient descent as meta-optimizers, 2023
Dai, D., Sun, Y., Dong, L., Hao, Y., Ma, S., Sui, Z., and Wei, F · 2023
Cited alongside, same era.
Faith and fate: Limits of transformers on compositionality, 2023
Dziri, N., Lu, X., Sclar, M., Li, X. L., Jiang, L., Lin, B. Y., West, P., Bhagavatula, C., Bras, R. L., Hwang, J. D., Sanyal, S., Welleck, S., Ren, X., Ettinger, A., Harchaoui, Z., and Choi, Y · 2023
Factscore: Fine-grained atomic evaluation of factual precision in long form text generation
Min, S., Krishna, K., Lyu, X., Lewis, M., Yih, W.-t., Koh, P. W., Iyyer, M., Zettlemoyer, L., and Hajishirzi, H · 2023
Later among the works it cites.
Progress measures for grokking via mechanistic interpretability, 2023
Nanda, N., Chan, L., Lieberum, T., Smith, J., and Steinhardt, J · 2023
Later among the works it cites.
Gpt-4 technical report, 2023
OpenAI · 2023
Later among the works it cites.
Is chatgpt a general-purpose natural language processing task solver?, 2023
Qin, C., Zhang, A., Zhang, Z., Chen, J., Yasunaga, M., and Yang, D · 2023
Later among the works it cites.
Positional description matters for transformers arithmetic, 2023
Shen, R., Bubeck, S., Eldan, R., Lee, Y. T., Li, Y., and Zhang, Y · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Towards revealing the mystery behind chain of thought: A theoretical perspective, 2023
Feng, G., Zhang, B., Gu, Y., Ye, H., He, D., and Wang, L · 2023
Cited alongside, same era.
What can transformers learn in-context? a case study of simple function classes, 2023
Garg, S., Tsipras, D., Liang, P., and Valiant, G · 2023
Cited alongside, same era.
Towards a mechanistic interpretation of multi-step reasoning capabilities of language models, 2023
Hou, Y., Li, J., Fei, Y., Stolfo, A., Zhou, W., Zeng, G., Bosselut, A., and Sachan, M · 2023
Cited alongside, same era.
The impact of positional encoding on length generalization in transformers, 2023
Kazemnejad, A., Padhi, I., Ramamurthy, K. N., Das, P., and Reddy, S · 2023
Cited alongside, same era.
Decomposed prompting: A modular approach for solving complex tasks, 2023
Khot, T., Trivedi, H., Finlayson, M., Fu, Y., Richardson, K., Clark, P., and Sabharwal, A · 2023
Cited alongside, same era.
Large language models are zero-shot reasoners, 2023
Kojima, T., Gu, S. S., Reid, M., Matsuo, Y., and Iwasawa, Y · 2023
Cited alongside, same era.
Humans in humans out: On gpt converging toward common sense in both success and failure, 2023
Koralus, P. and Wang-Maścianica, V · 2023
Cited alongside, same era.
von Oswald, J., Niklasson, E., Randazzo, E., Sacramento, J., Mordvintsev, A., Zhmoginov, A., and Vladymyrov, M · 2023
Later among the works it cites.
Chain-of-thought prompting elicits reasoning in large language models, 2023
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., Chi, E., Le, Q., and Zhou, D · 2023
Later among the works it cites.
Reasoning or reciting? exploring the capabilities and limitations of language models through counterfactual tasks, 2023
Wu, Z., Qiu, L., Ross, A., Akyürek, E., Chen, B., Wang, B., Kim, N., Andreas, J., and Kim, Y · 2023
Later among the works it cites.
It ain’t that bad: Understanding the mysterious performance drop in ood generalization for generative transformer models, 2023
Xu, X., Pan, Z., Zhang, H., and Yang, Y · 2023
Later among the works it cites.
Explaining the complex task reasoning of large language models with template-content structure, 2023
Yang, H., Meng, F., Lin, Z., and Zhang, M · 2023
Later among the works it cites.
Counterfactual memorization in neural language models, 2023
Zhang, C., Ippolito, D., Lee, K., Jagielski, M., Tramèr, F., and Carlini, N · 2023
Later among the works it cites.
The clock and the pizza: Two stories in mechanistic explanation of neural networks, 2023
Zhong, Z., Liu, Z., Tegmark, M., and Andreas, J · 2023
Later among the works it cites.
Large language models can learn rules, 2023
Zhu, Z., Xue, Y., Chen, X., Zhou, D., Tang, J., Schuurmans, D., and Dai, H · 2023
Later among the works it cites.
Transformers can achieve length generalization but not robustly, 2024
Zhou, Y., Alon, U., Chen, X., Wang, X., Agarwal, R., and Zhou, D · 2024
Closest in time.