Fetching the paper…
Reading the bibliography…
Chain-of-thought (CoT) prompting has been shown to empirically improve the accuracy of large language models (LLMs) on various question answering tasks.
Recursive deep models for semantic compositionality over a sentiment treebank
Socher, R., Perelygin, A., Wu, J., Chuang, J., Manning, C. D., Ng, A., and Potts, C · 2013
Earlier work this paper cites.
Visualizing and understanding neural models in NLP
Li, J., Chen, X., Hovy, E., and Jurafsky, D · 2016
Earlier work this paper cites.
Axiomatic attribution for deep networks, 2017
Sundararajan, M., Taly, A., and Yan, Q · 2017
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Earlier work this paper cites.
CommonsenseQA: A question answering challenge targeting commonsense knowledge
Talmor, A., Herzig, J., Lourie, N., and Berant, J · 2019
Earlier work this paper cites.
Allennlp interpret: A framework for explaining predictions of nlp models
Wallace, E., Tuyls, J., Wang, J., Subramanian, S., Gardner, M., and Singh, S · 2019
Earlier work this paper cites.
GPT-Neo: Large Scale Autoregressive Language Modeling with Mesh-Tensorflow, March 2021
Black, S., Leo, G., Wang, P., Leahy, C., and Biderman, S · 2021
Earlier work this paper cites.
An interpretability illusion for bert, 2021
Bolukbasi, T., Pearce, A., Yuan, A., Coenen, A., Reif, E., Viégas, F., and Wattenberg, M · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., Hesse, C., and Schulman, J · 2021
Cited alongside, same era.
Contrastive explanations for model interpretability
Jacovi, A., Swayamdipta, S., Ravfogel, S., Elazar, Y., Choi, Y., and Goldberg, Y · 2021
Cited alongside, same era.
GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model
Wang, B. and Komatsuzaki, A · 2021
Cited alongside, same era.
Roscoe: A suite of metrics for scoring step-by-step reasoning, 2022
Golovneva, O., Chen, M., Poff, S., Corredor, M., Zettlemoyer, L., Fazel-Zarandi, M., and Celikyilmaz, A · 2022
Cited alongside, same era.
Large language models are zero-shot reasoners, 2022
Kojima, T., Gu, S. S., Reid, M., Matsuo, Y., and Iwasawa, Y · 2022
Cited alongside, same era.
Post-hoc interpretability for neural NLP: A survey
Madsen, A., Reddy, S., and Chandar, S · 2022
Later among the works it cites.
Towards understanding chain-of-thought prompting: An empirical study of what matters
Wang, B., Min, S., Deng, X., Shen, J., Wu, Y., Zettlemoyer, L., and Sun, H · 2022
Later among the works it cites.
Chain of thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., Chi, E. H., Le, Q. V., and Zhou, D · 2022
Later among the works it cites.
Interpreting language models with contrastive explanations
Yin, K. and Neubig, G · 2022
Later among the works it cites.
Automatic chain of thought prompting in large language models
Zhang, Z., Zhang, A., Li, M., and Smola, A · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The disagreement problem in explainable machine learning: A practitioner’s perspective, 2022
Krishna, S., Han, T., Gu, A., Pombra, J., Jabbari, S., Wu, S., and Lakkaraju, H · 2022
Cited alongside, same era.
Can language models learn from explanations in context?
Lampinen, A. K., Dasgupta, I., Chan, S. C., Matthewson, K., Tessler, M. H., Creswell, A., McClelland, J. L., Wang, J. X., and Hill, F · 2022
Cited alongside, same era.
Text and patterns: For effective chain of thought, it takes two to tango
Madaan, A. and Yazdanbakhsh, A · 2022
Cited alongside, same era.
Star: Bootstrapping reasoning with reasoning, 2022a
Zelikman, E., Wu, Y., Mu, J., and Goodman, N. D
Cited in the paper.
Star: Bootstrapping reasoning with reasoning
Zelikman, E., Wu, Y., Mu, J., and Goodman, N. D
Cited in the paper.
Language models don’t always say what they think: Unfaithful explanations in chain-of-thought prompting, 2023
Turpin, M., Michael, J., Perez, E., and Bowman, S. R · 2023
Closest in time.
Self-consistency improves chain of thought reasoning in language models, 2023
Wang, X., Wei, J., Schuurmans, D., Le, Q., Chi, E., Narang, S., Chowdhery, A., and Zhou, D · 2023
Closest in time.