Fetching the paper…
Reading the bibliography…
This paper delves into the capabilities of large language models (LLMs), specifically focusing on advancing the theoretical comprehension of chain-of-thought prompting.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L. u., and Polosukhin, I · 2017
Earlier work this paper cites.
Language models are unsupervised multitask learners, 2019
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Are transformers universal approximators of sequence-to-sequence functions?
Yun, C., Bhojanapalli, S., Rawat, A. S., Reddi, S., and Kumar, S · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., et al · 2021
Earlier work this paper cites.
The power of scale for parameter-efficient prompt tuning
Lester, B., Al-Rfou, R., and Constant, N · 2021
Earlier work this paper cites.
Self-supervised learning is more robust to dataset imbalance
Liu, H., HaoChen, J. Z., Gaidon, A., and Ma, T · 2021
Earlier work this paper cites.
A new hybrid classification algorithm for predicting customer churn
Markapudi, B., Latha, K. J., and Chaduvula, K · 2021
Earlier work this paper cites.
Selfaugment: Automatic augmentation policies for self-supervised learning
Reed, C. J., Metzger, S., Srinivas, A., Darrell, T., and Keutzer, K · 2021
Earlier work this paper cites.
Self-supervised learning for large-scale item recommendations
Yao, T., Yi, X., Cheng, D. Z., Yu, F., Chen, T., Menon, A., Hong, L., Chi, E. H., Tjoa, S., Kang, J., et al · 2021
Earlier work this paper cites.
Do as i can, not as i say: Grounding language in robotic affordances
Ahn, M., Brohan, A., Brown, N., Chebotar, Y., Cortes, O., David, B., Finn, C., Fu, C., Gopalakrishnan, K., Hausman, K., et al · 2022
Earlier work this paper cites.
Attributed question answering: Evaluation and modeling for attributed large language models
Bohnet, B., Tran, V. Q., Verga, P., Aharoni, R., Andor, D., Soares, L. B., Eisenstein, J., Ganchev, K., Herzig, J., Hui, K., et al · 2022
Earlier work this paper cites.
Data distributional properties drive emergent in-context learning in transformers
Chan, S. C., Santoro, A., Lampinen, A. K., Wang, J. X., Singh, A. K., Richemond, P. H., McClelland, J., and Hill, F · 2022
Earlier work this paper cites.
Chen, W., Ma, X., Wang, X., and Cohen, W. W · 2022
Earlier work this paper cites.
Pangu-coder: Program synthesis with function-level language modeling, 2022
Christopoulou, F., Lampouras, G., Gritta, M., Zhang, G., Guo, Y., Li, Z., Zhang, Q., Xiao, M., Shen, B., Li, L., Yu, H., Yan, L., Zhou, P., Wang, X., Ma, Y., Iacobacci, I., Wang, Y., Liang, G., Wei, J., Jiang, X., Wang, Q., and Liu, Q · 2022
Earlier work this paper cites.
Large language models are zero-shot reasoners
Kojima, T., Gu, S. S., Reid, M., Matsuo, Y., and Iwasawa, Y · 2022
Earlier work this paper cites.
Rethinking the role of demonstrations: What makes in-context learning work?
Min, S., Lyu, X., Holtzman, A., Artetxe, M., Lewis, M., Hajishirzi, H., and Zettlemoyer, L · 2022
Earlier work this paper cites.
Impact of pretraining term frequencies on few-shot numerical reasoning
Razeghi, Y., IV, R. L. L., Gardner, M., and Singh, S · 2022
Earlier work this paper cites.
Leveraging large language models for multiple choice question answering
Robinson, J., Rytting, C. M., and Wingate, D · 2022
Earlier work this paper cites.
Multitask prompted training enables zero-shot task generalization
Sanh, V., Webson, A., Raffel, C., Bach, S., Sutawika, L., Alyafeai, Z., Chaffin, A., Stiegler, A., Raja, A., Dey, M., Bari, M. S., Xu, C., Thakker, U., Sharma, S. S., Szczechla, E., Kim, T., Chhablani, G., Nayak, N., Datta, D., Chang, J., Jiang, M. T.-J., Wang, H., Manica, M., Shen, S., Yong, Z. X., Pandey, H., Bawden, R., Wang, T., Neeraj, T., Rozen, J., Sharma, A., Santilli, A., Fevry, T., Fries, J. A., Teehan, R., Scao, T. L., Biderman, S., Gao, L., Wolf, T., and Rush, A. M · 2022
Earlier work this paper cites.
Emergent structures and training dynamics in large language models
Teehan, R., Clinciu, M., Serikov, O., Szczechla, E., Seelam, N., Mirkin, S., and Gokaslan, A · 2022
Earlier work this paper cites.
Code4struct: Code generation for few-shot structured prediction from natural language
Wang, X., Li, S., and Ji, H · 2022
Cited alongside, same era.
An explanation of in-context learning as implicit bayesian inference
Xie, S. M., Raghunathan, A., Liang, P., and Ma, T · 2022
Cited alongside, same era.
Graph of thoughts: Solving elaborate problems with large language models
Besta, M., Blach, N., Kubicek, A., Gerstenberger, R., Gianinazzi, L., Gajda, J., Lehmann, T., Podstawski, M., Niewiadomski, H., Nyczyk, P., et al · 2023
Cited alongside, same era.
Emergent autonomous scientific research capabilities of large language models
Boiko, D. A., MacKnight, R., and Gomes, G · 2023
Cited alongside, same era.
Capabilities of gpt-4 on medical challenge problems
Nori, H., King, N., McKinney, S. M., Carignan, D., and Horvitz, E · 2023
Closest in time.
OpenAI · 2023
Closest in time.
Detecting llm-generated text in computing education: A comparative study for chatgpt cases
Orenstrakh, M. S., Karnalim, O., Suarez, C. A., and Liut, M · 2023
Closest in time.
Ai assistant for document management using lang chain and pinecone
Pesaru, A., Gill, T. S., and Tangella, A. R · 2023
Closest in time.
Lm-nav: Robotic navigation with large pre-trained models of language, vision, and action
Shah, D., Osiński, B., Levine, S., et al · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chang, J. D., Brantley, K., Ramamurthy, R., Misra, D., and Sun, W · 2023
Cited alongside, same era.
Evoprompting: Language models for code-level neural architecture search
Chen, A., Dohan, D. M., and So, D. R · 2023
Cited alongside, same era.
Towards revealing the mystery behind chain of thought: A theoretical perspective, 2023
Feng, G., Zhang, B., Gu, Y., Ye, H., He, D., and Wang, L · 2023
Cited alongside, same era.
Promptbreeder: Self-referential self-improvement via prompt evolution
Fernando, C., Banarse, D., Michalewski, H., Osindero, S., and Rocktäschel, T · 2023
Cited alongside, same era.
Openagi: When llm meets domain experts
Ge, Y., Hua, W., Ji, J., Tan, J., Xu, S., and Zhang, Y · 2023
Cited alongside, same era.
Large language model ai chatbots require approval as medical devices
Gilbert, S., Harvey, H., Melvin, T., Vollebregt, E., and Wicks, P · 2023
Cited alongside, same era.
Dimensions for designing llm-based writing support
Gmeiner, F. and Yildirim, N · 2023
Cited alongside, same era.
A theory of emergent in-context learning as implicit structure induction
Hahn, M. and Goyal, N · 2023
Cited alongside, same era.
Pangu-coder2: Boosting large language models for code with ranking feedback
Shen, B., Zhang, J., Chen, T., Zan, D., Geng, B., Fu, A., Zeng, M., Yu, A., Ji, J., Zhao, J., et al · 2023
Closest in time.
Progprompt: Generating situated robot task plans using large language models
Singh, I., Blukis, V., Mousavian, A., Goyal, A., Xu, D., Tremblay, J., Fox, D., Thomason, J., and Garg, A · 2023
Closest in time.
Towards expert-level medical question answering with large language models
Singhal, K., Tu, T., Gottweis, J., Sayres, R., Wulczyn, E., Hou, L., Clark, K., Pfohl, S., Cole-Lewis, H., Neal, D., et al · 2023
Closest in time.
Challenging big-bench tasks and whether chain-of-thought can solve them
Suzgun, M., Scales, N., Schärli, N., Gehrmann, S., Tay, Y., Chung, H. W., Chowdhery, A., Le, Q. V., Chi, E., Zhou, D., and Wei, J · 2023
Closest in time.
Evaluation of chatgpt as a question answering system for answering complex questions
Tan, Y., Min, D., Li, Y., Li, W., Hu, N., Chen, Y., and Qi, G · 2023
Closest in time.
Large language models in medicine
Thirunavukarasu, A. J., Ting, D. S. J., Elangovan, K., Gutierrez, L., Tan, T. F., and Ting, D. S. W · 2023
Closest in time.
Llama: Open and efficient foundation language models, 2023
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., Rodriguez, A., Joulin, A., Grave, E., and Lample, G · 2023
Closest in time.
Chatgpt for robotics: Design principles and model abilities
Vemprala, S., Bonatti, R., Bucker, A., and Kapoor, A · 2023
Closest in time.
Freshllms: Refreshing large language models with search engine augmentation
Vu, T., Iyyer, M., Wang, X., Constant, N., Wei, J., Wei, J., Tar, C., Sung, Y.-H., Zhou, D., Le, Q., et al · 2023
Closest in time.
Emergent analogical reasoning in large language models
Webb, T., Holyoak, K. J., and Lu, H · 2023
Closest in time.
The learnability of in-context learning
Wies, N., Levine, Y., and Shashua, A · 2023
Closest in time.
Large language models as optimizers
Yang, C., Wang, X., Lu, Y., Liu, H., Le, Q. V., Zhou, D., and Chen, X · 2023
Closest in time.
Tree of thoughts: Deliberate problem solving with large language models
Yao, S., Yu, D., Zhao, J., Shafran, I., Griffiths, T. L., Cao, Y., and Narasimhan, K · 2023
Closest in time.
Sentiment analysis in the era of large language models: A reality check
Zhang, W., Deng, Y., Liu, B., Pan, S. J., and Bing, L · 2023
Closest in time.
Codegeex: A pre-trained model for code generation with multilingual evaluations on humaneval-x
Zheng, Q., Xia, X., Zou, X., Dong, Y., Wang, S., Xue, Y., Wang, Z., Shen, L., Wang, A., Li, Y., et al · 2023
Closest in time.