Fetching the paper…
Reading the bibliography…
As large language models (LLMs) perform more difficult tasks, it becomes harder to verify the correctness and safety of their behavior.
Building watson: An overview of the deepqa project
Ferrucci, D. A., Brown, E. W., Chu-Carroll, J., Fan, J., Gondek, D., Kalyanpur, A., Lally, A., Murdock, J. W., Nyberg, E., Prager, J. M., Schlaefer, N., and Welty, C · 2010
Earlier work this paper cites.
Towards a rigorous science of interpretable machine learning
Doshi-Velez, F. and Kim, B · 2017
Earlier work this paper cites.
Supervising strong learners by amplifying weak experts
Christiano, P., Shlegeris, B., and Amodei, D · 2018
Earlier work this paper cites.
Can a suit of armor conduct electricity? a new dataset for open book question answering
Mihaylov, T., Clark, P., Khot, T., and Sabharwal, A · 2018
Earlier work this paper cites.
Factored cognition
Stuhlmüeller, A · 2018
Earlier work this paper cites.
HotpotQA: A dataset for diverse, explainable multi-hop question answering
Yang, Z., Qi, P., Zhang, S., Bengio, Y., Cohen, W. W., Salakhutdinov, R., and Manning, C. D · 2018
Earlier work this paper cites.
Multi-hop reading comprehension through question decomposition and rescoring
Min, S., Zhong, V., Zettlemoyer, L., and Hajishirzi, H · 2019
Earlier work this paper cites.
Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead
Rudin, C · 2019
Earlier work this paper cites.
Leakage-adjusted simulatability: Can models generate non-trivial explanations of their behavior in natural language?
Hase, P., Zhang, S., Xie, H., and Bansal, M · 2020
Earlier work this paper cites.
The curious case of neural text degeneration
Holtzman, A., Buys, J., Du, L., Forbes, M., and Choi, Y · 2020
Earlier work this paper cites.
Towards faithfully interpretable NLP systems: How should we define and evaluate faithfulness?
Jacovi, A. and Goldberg, Y · 2020
Earlier work this paper cites.
Unsupervised question decomposition for question answering
Perez, E., Lewis, P., Yih, W.-t., Cho, K., and Kiela, D · 2020
Earlier work this paper cites.
Evaluating large language models trained on code, 2021
Chen, M., Tworek, J., Jun, H., Yuan, Q., de Oliveira Pinto, H. P., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., Ray, A., Puri, R., Krueger, G., Petrov, M., Khlaaf, H., Sastry, G., Mishkin, P., Chan, B., Gray, S., Ryder, N., Pavlov, M., Power, A., Kaiser, L., Bavarian, M., Winter, C., Tillet, P., Such, F. P., Cummings, D., Plappert, M., Chantzis, F., Barnes, E., Herbert-Voss, A., Guss, W. H., Nichol, A., Paino, A., Tezak, N., Tang, J., Babuschkin, I., Balaji, S., Jain, S., Saunders, W., Hesse, C., Carr, A. N., Leike, J., Achiam, J., Misra, V., Morikawa, E., Radford, A., Knight, M., Brundage, M., Murati, M., Mayer, K., Welinder, P., McGrew, B., Amodei, D., McCandlish, S., Sutskever, I., and Zaremba, W · 2021
Earlier work this paper cites.
Decomposing complex questions makes multi-hop QA easier and more interpretable
Fu, R., Wang, H., Zhang, X., Zhou, J., and Yan, Y · 2021
Earlier work this paper cites.
Did Aristotle Use a Laptop? A Question Answering Benchmark with Implicit Reasoning Strategies
Geva, M., Khashabi, D., Segal, E., Khot, T., Roth, D., and Berant, J · 2021
Earlier work this paper cites.
Webgpt: Browser-assisted question-answering with human feedback
Nakano, R., Hilton, J., Balaji, S., Wu, J., Ouyang, L., Kim, C., Hesse, C., Jain, S., Kosaraju, V., Saunders, W., Jiang, X., Cobbe, K., Eloundou, T., Krueger, G., Button, K., Knight, M., Chess, B., and Schulman, J · 2021
Earlier work this paper cites.
Show your work: Scratchpads for intermediate computation with language models
Nye, M., Johan Andreassen, A., Gur-Ari, G., Michalewski, H., Austin, J., Bieber, D., Dohan, D., Lewkowycz, A., Bosma, M., Luan, D., Sutton, C., and Odena, A · 2021
Cited alongside, same era.
Prompt programming for large language models: Beyond the few-shot paradigm
Reynolds, L. and McDonell, K · 2021
Cited alongside, same era.
Measuring association between labels and free-text rationales
Wiegreffe, S., Marasović, A., and Smith, N. A · 2021
Cited alongside, same era.
ToKen: Task decomposition and knowledge infusion for few-shot hate speech detection
AlKhamissi, B., Ladhak, F., Iyer, S., Stoyanov, V., Kozareva, Z., Li, X., Fung, P., Mathias, L., Celikyilmaz, A., and Diab, M · 2022
Cited alongside, same era.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Galactica: A large language model for science, 2022
Taylor, R., Kardas, M., Cucurull, G., Scialom, T., Hartshorn, A., Saravia, E., Poulton, A., Kerkez, V., and Stojnic, R · 2022
Later among the works it cites.
Solving math word problems with process- and outcome-based feedback
Uesato, J., Kushman, N., Kumar, R., Song, F., Siegel, N., Wang, L., Creswell, A., Irving, G., and Higgins, I · 2022
Later among the works it cites.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., ichter, b., Xia, F., Chi, E., Le, Q. V., and Zhou, D · 2022
Later among the works it cites.
The unreliability of explanations in few-shot prompting for textual reasoning
Ye, X. and Durrett, G · 2022
Later among the works it cites.
Introducing claude, 2023
Anthropic · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bai, Y., Jones, A., Ndousse, K., Askell, A., Chen, A., DasSarma, N., Drain, D., Fort, S., Ganguli, D., Henighan, T., Joseph, N., Kadavath, S., Kernion, J., Conerly, T., El-Showk, S., Elhage, N., Hatfield-Dodds, Z., Hernandez, D., Hume, T., Johnston, S., Kravec, S., Lovitt, L., Nanda, N., Olsson, C., Amodei, D., Brown, T., Clark, J., McCandlish, S., Olah, C., Mann, B., and Kaplan, J · 2022
Cited alongside, same era.
Successive prompting for decomposing complex questions
Dua, D., Gupta, S., Singh, S., and Gardner, M · 2022
Cited alongside, same era.
Complex reading comprehension through question decomposition
Guo, X.-Y., Li, Y.-F., and Haffari, G · 2022
Cited alongside, same era.
Maieutic prompting: Logically consistent reasoning with recursive explanations
Jung, J., Qin, L., Welleck, S., Brahman, F., Bhagavatula, C., Le Bras, R., and Choi, Y · 2022
Cited alongside, same era.
Large language models are zero-shot reasoners
Kojima, T., Gu, S. S., Reid, M., Matsuo, Y., and Iwasawa, Y · 2022
Cited alongside, same era.
Explanations from large language models make small reasoners better
Li, S., Chen, J., Shen, Y., Chen, Z., Zhang, X., Li, Z., Wang, H., Qian, J., Peng, B., Mao, Y., Chen, W., and Yan, X · 2022
Cited alongside, same era.
TruthfulQA: Measuring how models mimic human falsehoods
Lin, S., Hilton, J., and Evans, O · 2022
Cited alongside, same era.
Text and patterns: For effective chain of thought, it takes two to tango, 2022
Madaan, A. and Yazdanbakhsh, A · 2022
Cited alongside, same era.
Creswell, A., Shanahan, M., and Higgins, I · 2023
Closest in time.
Shapley value attribution in chain of thought
Gao, L · 2023
Closest in time.
ROSCOE: A suite of metrics for scoring step-by-step reasoning
Golovneva, O., Chen, M. P., Poff, S., Corredor, M., Zettlemoyer, L., Fazel-Zarandi, M., and Celikyilmaz, A · 2023
Closest in time.
Measuring faithfulness in chain-of-thought reasoning
Lanham, T., Chen, A., Radhakrishnan, A., Steiner, B., Denison, C., Hernandez, D., Li, D., Durmus, E., Hubinger, E., Kernion, J., Lukosuite, K., Nguyen, K., Cheng, N., Joseph, N., Schiefer, N., Rausch, O., Larson, R., McCandlish, S., Kundu, S., Kadavath, S., Yang, S., Henighan, T., Maxwell, T., Telleen-Lawton, T., Hume, T., Hatfield-Dodds, Z., Kaplan, J., Brauner, J., Bowman, S. R., and Perez, E · 2023
Closest in time.
Lightman, H., Kosaraju, V., Burda, Y., Edwards, H., Baker, B., Lee, T., Leike, J., Schulman, J., Sutskever, I., and Cobbe, K · 2023
Closest in time.
Faithful chain-of-thought reasoning
Lyu, Q., Havaldar, S., Stein, A., Zhang, L., Rao, D., Wong, E., Apidianaki, M., and Callison-Burch, C · 2023
Closest in time.
Interpretability Dreams, 2023
Olah, C · 2023
Closest in time.
Iterated decomposition: Improving science Q&A by supervising reasoning processes
Reppert, J., Rachbach, B., George, C., Stebbing, L., Byun, J., Appleton, M., and Stuhlmüeller, A · 2023
Closest in time.
Turpin, M., Michael, J., Perez, E., and Bowman, S. R · 2023
Closest in time.
Plan-and-solve prompting: Improving zero-shot chain-of-thought reasoning by large language models
Wang, L., Xu, W., Lan, Y., Hu, Z., Lan, Y., Lee, R. K.-W., and Lee, E.-P · 2023
Closest in time.
Least-to-most prompting enables complex reasoning in large language models
Zhou, D., Schärli, N., Hou, L., Wei, J., Scales, N., Wang, X., Schuurmans, D., Cui, C., Bousquet, O., Le, Q., and Chi, E · 2023
Closest in time.