Fetching the paper…
Reading the bibliography…
Large language models (LLMs) with memory are computationally universal.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. (2020) · 1901
Earlier work this paper cites.
Giving bert a calculator: Finding operations and arguments with reading comprehension
Andor, D., He, L., Lee, K., and Pitler, E. (2019) · 1909
Earlier work this paper cites.
Numnet: Machine reading comprehension with numerical reasoning
Ran, Q., Lin, Y., Li, P., Zhou, J., and Liu, Z. (2019) · 1910
Earlier work this paper cites.
Graves, A., Wayne, G., and Danihelka, I. (2014) · 2014
Earlier work this paper cites.
Learning graphical state transitions
Johnson, D. D. (2017) · 2017
Earlier work this paper cites.
Retrieval augmented language model pre-training
Guu, K., Lee, K., Tung, Z., Pasupat, P., and Chang, M. (2020) · 2020
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W.-t., Rocktäschel, T., et al. (2020) · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., et al. (2021) · 2021
Earlier work this paper cites.
The power of scale for parameter-efficient prompt tuning
Lester, B., Al-Rfou, R., and Constant, N. (2021) · 2021
Earlier work this paper cites.
Measuring and improving bert’s mathematical abilities by predicting the order of reasoning
Piękos, P., Michalewski, H., and Malinowski, M. (2021) · 2021
Earlier work this paper cites.
Finetuned language models are zero-shot learners
Wei, J., Bosma, M., Zhao, V. Y., Guu, K., Yu, A. W., Lester, B., Du, N., Dai, A. M., and Le, Q. V. (2021) · 2021
Earlier work this paper cites.
Recurrent memory transformer
Bulatov, A., Kuratov, Y., and Burtsev, M. (2022) · 2022
Earlier work this paper cites.
Chen, W., Ma, X., Wang, X., and Cohen, W. W. (2022) · 2022
Cited alongside, same era.
Binding language models in symbolic languages
Cheng, Z., Xie, T., Shi, P., Li, C., Nadkarni, R., Hu, Y., Xiong, C., Radev, D., Ostendorf, M., Zettlemoyer, L., et al. (2022) · 2022
Cited alongside, same era.
Glm: General language model pretraining with autoregressive blank infilling
Du, Z., Qian, Y., Liu, X., Ding, M., Qiu, J., Yang, Z., and Tang, J. (2022) · 2022
Cited alongside, same era.
Few-shot learning with retrieval augmented language models
Izacard, G., Lewis, P., Lomeli, M., Hosseini, L., Petroni, F., Schick, T., Dwivedi-Yu, J., Joulin, A., Riedel, S., and Grave, E. (2022) · 2022
Cited alongside, same era.
Demonstrate-search-predict: Composing retrieval and language models for knowledge-intensive nlp
Reasoning with language model is planning with world model
Hao, S., Gu, Y., Ma, H., Hong, J. J., Wang, Z., Wang, D. Z., and Hu, Z. (2023) · 2023
Closest in time.
Gpt-4 technical report
OpenAI (2023) · 2023
Closest in time.
Art: Automatic multi-step reasoning and tool-use for large language models
Paranjape, B., Lundberg, S., Singh, S., Hajishirzi, H., Zettlemoyer, L., and Ribeiro, M. T. (2023) · 2023
Closest in time.
Generative agents: Interactive simulacra of human behavior
Park, J. S., O’Brien, J. C., Cai, C. J., Morris, M. R., Liang, P., and Bernstein, M. S. (2023) · 2023
Closest in time.
Toolformer: Language models can teach themselves to use tools
Schick, T., Dwivedi-Yu, J., Dessì, R., Raileanu, R., Lomeli, M., Zettlemoyer, L., Cancedda, N., and Scialom, T. (2023) · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Khattab, O., Santhanam, K., Li, X. L., Hall, D., Liang, P., Potts, C., and Zaharia, M. (2022) · 2022
Cited alongside, same era.
Self-consistency improves chain of thought reasoning in language models
Wang, X., Wei, J., Schuurmans, D., Le, Q., Chi, E., and Zhou, D. (2022) · 2022
Cited alongside, same era.
Chain of thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Chi, E., Le, Q., and Zhou, D. (2022) · 2022
Cited alongside, same era.
Glm-130b: An open bilingual pre-trained model
Zeng, A., Liu, X., Du, Z., Wang, Z., Lai, H., Ding, M., Yang, Z., Xu, Y., Zheng, W., Xia, X., et al. (2022) · 2022
Cited alongside, same era.
Training language models with memory augmentation
Zhong, Z., Lei, T., and Chen, D. (2022) · 2022
Cited alongside, same era.
Anil, R., Dai, A. M., Firat, O., Johnson, M., Lepikhin, D., Passos, A., Shakeri, S., Taropa, E., Bailey, P., Chen, Z., et al. (2023) · 2023
Cited alongside, same era.
Scaling transformer to 1m tokens and beyond with rmt
Bulatov, A., Kuratov, Y., and Burtsev, M. S. (2023) · 2023
Cited alongside, same era.
Two failures of self-consistency in the multi-step reasoning of llms
Chen, A., Phang, J., Parrish, A., Padmakumar, V., Zhao, C., Bowman, S. R., and Cho, K. (2023) · 2023
Cited alongside, same era.
Closest in time.
Memory augmented large language models are computationally universal
Schuurmans, D. (2023) · 2023
Closest in time.
Hugginggpt: Solving ai tasks with chatgpt and its friends in huggingface
Shen, Y., Song, K., Tan, X., Li, D., Lu, W., and Zhuang, Y. (2023) · 2023
Closest in time.
Sql-palm: Improved large language modeladaptation for text-to-sql
Sun, R., Arik, S. O., Nakhost, H., Dai, H., Sinha, R., Yin, P., and Pfister, T. (2023) · 2023
Closest in time.
Vipergpt: Visual inference via python execution for reasoning
Surís, D., Menon, S., and Vondrick, C. (2023) · 2023
Closest in time.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al. (2023) · 2023
Closest in time.
Visionllm: Large language model is also an open-ended decoder for vision-centric tasks
Wang, W., Chen, Z., Chen, X., Wu, J., Zhu, X., Zeng, G., Luo, P., Lu, T., Zhou, J., Qiao, Y., et al. (2023) · 2023
Closest in time.