Fetching the paper…
Reading the bibliography…
During both pretraining and fine-tuning, Large Language Models (\textbf{LLMs}) are trained on trillions of tokens of text of widely varying quality.
Learning from massive noisy labeled data for image classification
Xiao, T., Xia, T., Yang, Y., Huang, C., and Wang, X · 2015
Earlier work this paper cites.
Synthetic and natural noise both break neural machine translation
Belinkov, Y. and Bisk, Y · 2017
Earlier work this paper cites.
Toward robust neural machine translation for noisy input sequences
Sperber, M., Niehues, J., and Waibel, A · 2017
Earlier work this paper cites.
Neural machine translation of text from non-native speakers
Anastasopoulos, A., Lui, A., Nguyen, T. Q., and Chiang, D · 2019
Earlier work this paper cites.
Realistic noisy text generation for fortified nlp
Pruthi, D., Kumar, S., Field, A., and Vogler, N · 2019
Earlier work this paper cites.
Beyond class-conditional assumption: A primary attempt to combat instance-dependent label noise
Chen, P., Ye, J., Chen, G., Zhao, J., and Heng, P.-A · 2020
Earlier work this paper cites.
Learning from noisy labels with deep neural networks: A survey
Song, H., Kim, M., Park, D., Shin, Y., and Lee, J.-G · 2020
Earlier work this paper cites.
Towards a better understanding of noise in natural language processing
Al Sharou, K., Li, Z., and Specia, L · 2021
Earlier work this paper cites.
Analysing the noise model error for realistic noisy label data
Hedderich, M. A., Zhu, D., and Klakow, D · 2021
Earlier work this paper cites.
Investigating the limitations of the transformers with simple arithmetic tasks
Nogueira, R., Jiang, Z., and Li, J. J · 2021
Earlier work this paper cites.
Show your work: Scratchpads for intermediate computation with language models
Nye, M., Andreassen, A., Gur-Ari, G., Michalewski, H., Austin, J., Bieber, D., Dohan, D., Lewkowycz, A., Bosma, M., Luan, D., Sutton, C., and Odena, A · 2021
Earlier work this paper cites.
Chen, W., Ma, X., Wang, X., and Cohen, W. W · 2022
Earlier work this paper cites.
Scaling instruction-finetuned language models
Chung, H. W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., Li, E., Wang, X., Dehghani, M., Brahma, S., Webson, A., Gu, S. S., Dai, Z., Suzgun, M., Chen, X., Chowdhery, A., Valter, D., Narang, S., Mishra, G., Yu, A. W., Zhao, V., Huang, Y., Dai, A. M., Yu, H., Petrov, S., hsin Chi, E. H., Dean, J., Devlin, J., Roberts, A., Zhou, D., Le, Q. V., and Wei, J · 2022
Earlier work this paper cites.
Dohan, D., Xu, W., Lewkowycz, A., Austin, J., Bieber, D., Lopes, R. G., Wu, Y., Michalewski, H., Saurous, R. A., Sohl-Dickstein, J. N., Murphy, K., and Sutton, C · 2022
Earlier work this paper cites.
Pal: Program-aided language models
Gao, L., Madaan, A., Zhou, S., Alon, U., Liu, P., Yang, Y., Callan, J., and Neubig, G · 2022
Cited alongside, same era.
Large language models are reasoning teachers
Ho, N., Schmid, L., and Yun, S.-Y · 2022
Cited alongside, same era.
Large language models can self-improve
Huang, J., Gu, S. S., Hou, L., Wu, Y., Wang, X., Yu, H., and Han, J · 2022
Cited alongside, same era.
Text and patterns: For effective chain of thought, it takes two to tango
Madaan, A. and Yazdanbakhsh, A · 2022
Cited alongside, same era.
Rethinking the role of demonstrations: What makes in-context learning work?
Min, S., Lyu, X., Holtzman, A., Artetxe, M., Lewis, M., Hajishirzi, H., and Zettlemoyer, L · 2022
Distilling step-by-step! outperforming larger language models with less training data and smaller model sizes
Hsieh, C.-Y., Li, C.-L., Yeh, C.-K., Nakhost, H., Fujii, Y., Ratner, A., Krishna, R., Lee, C.-Y., and Pfister, T · 2023
Later among the works it cites.
Recursion of thought: Divide and conquer reasoning with language models, 2023
Lee, S. and Kim, G · 2023
Later among the works it cites.
Self-alignment with instruction backtranslation
Li, X., Yu, P., Zhou, C., Schick, T., Zettlemoyer, L., Levy, O., Weston, J., and Lewis, M · 2023
Later among the works it cites.
Goat: Fine-tuned llama outperforms gpt-4 on arithmetic tasks
Liu, T. and Low, B. K. H · 2023
Later among the works it cites.
Faithful chain-of-thought reasoning
LYU, Q., Havaldar, S., Stein, A., Zhang, L., Rao, D., Wong, E., Apidianaki, M., and Callison-Burch, C · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Reasoning like program executors
Pi, X., Liu, Q., Chen, B., Ziyadi, M., Lin, Z., Gao, Y., Fu, Q., Lou, J.-G., and Chen, W · 2022
Cited alongside, same era.
Limitations of language models in arithmetic and symbolic induction
Qian, J., Wang, H., Li, Z., LI, S., and Yan, X · 2022
Cited alongside, same era.
Transformers learn in-context by gradient descent
von Oswald, J., Niklasson, E., Randazzo, E., Sacramento, J., Mordvintsev, A., Zhmoginov, A., and Vladymyrov, M · 2022
Cited alongside, same era.
Chain of thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., hsin Chi, E. H., Xia, F., Le, Q., and Zhou, D · 2022
Cited alongside, same era.
Star: Bootstrapping reasoning with reasoning
Zelikman, E., Wu, Y., and Goodman, N. D · 2022
Cited alongside, same era.
Teaching algorithmic reasoning via in-context learning
Zhou, H., Nova, A., Larochelle, H., Courville, A. C., Neyshabur, B., and Sedghi, H · 2022
Cited alongside, same era.
Pythia: A suite for analyzing large language models across training and scaling
Biderman, S. R., Schoelkopf, H., Anthony, Q. G., Bradley, H., O’Brien, K., Hallahan, E., Khan, M. A., Purohit, S., Prashanth, U. S., Raff, E., Skowron, A., Sutawika, L., and van der Wal, O · 2023
Cited alongside, same era.
Evaluating transformer language models on arithmetic operations using number decomposition
Muffo, M., Cocco, A., and Bertino, E · 2023
Later among the works it cites.
Orca: Progressive learning from complex explanation traces of gpt-4
Mukherjee, S., Mitra, A., Jawahar, G., Agarwal, S., Palangi, H., and Awadallah, A. H · 2023
Later among the works it cites.
Toolformer: Language models can teach themselves to use tools
Schick, T., Dwivedi-Yu, J., Dessì, R., Raileanu, R., Lomeli, M., Zettlemoyer, L., Cancedda, N., and Scialom, T · 2023
Later among the works it cites.
Large language models can be easily distracted by irrelevant context
Shi, F., Chen, X., Misra, K., Scales, N., Dohan, D., hsin Chi, E. H., Scharli, N., and Zhou, D · 2023
Later among the works it cites.
Turpin, M., Michael, J., Perez, E., and Bowman, S · 2023
Later among the works it cites.
Noisywikihow: A benchmark for learning with real-world noisy labels in natural language processing
Wu, T., Ding, X., Tang, M., Zhang, H., Qin, B., and Liu, T · 2023
Later among the works it cites.
Wizardlm: Empowering large language models to follow complex instructions
Xu, C., Sun, Q., Zheng, K., Geng, X., Zhao, P., Feng, J., Tao, C., and Jiang, D · 2023
Later among the works it cites.
Lima: Less is more for alignment
Zhou, C., Liu, P., Xu, P., Iyer, S., Sun, J., Mao, Y., Ma, X., Efrat, A., Yu, P., Yu, L., Zhang, S., Ghosh, G., Lewis, M., Zettlemoyer, L., and Levy, O · 2023
Later among the works it cites.