Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) are increasingly used as powerful tools for several high-stakes natural language processing (NLP) applications.
“Why should I trust you?" Explaining the predictions of any classifier
Ribeiro, M. T., Singh, S., and Guestrin, C. (2016) · 2016
Earlier work this paper cites.
The (un)reliability of saliency methods
Kindermans, P.-J., Hooker, S., Adebayo, J., Alber, M., Schütt, K. T., Dähne, S., Erhan, D., and Kim, B. (2017) · 2017
Earlier work this paper cites.
Smoothgrad: Removing noise by adding noise
Smilkov, D., Thorat, N., Kim, B., Viégas, F., and Wattenberg, M. (2017) · 2017
Earlier work this paper cites.
Axiomatic attribution for deep networks
Sundararajan, M., Taly, A., and Yan, Q. (2017) · 2017
Earlier work this paper cites.
Sentence-bert: Sentence embeddings using siamese bert-networks
Reimers, N. and Gurevych, I. (2019) · 2019
Earlier work this paper cites.
Visualizing attention in transformer-based language representation models
Vig, J. (2019) · 2019
Earlier work this paper cites.
Attention flows: Analyzing and comparing attention mechanisms in language models
DeRose, J. F., Wang, J., and Berger, M. (2020) · 2020
Earlier work this paper cites.
Is bert really robust? a strong baseline for natural language attack on text classification and entailment
Jin, D., Jin, Z., Zhou, J. T., and Szolovits, P. (2020) · 2020
Earlier work this paper cites.
Perturbed masking: Parameter-free probing for analyzing and interpreting BERT
Wu, Z., Chen, Y., Kao, B., and Liu, Q. (2020) · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., et al. (2021) · 2021
Earlier work this paper cites.
Did aristotle use a laptop? a question answering benchmark with implicit reasoning strategies
Geva, M., Khashabi, D., Segal, E., Khot, T., Roth, D., and Berant, J. (2021) · 2021
Earlier work this paper cites.
How can we know when language models know? on the calibration of language models for question answering
Jiang, Z., Araki, J., Ding, H., and Neubig, G. (2021) · 2021
Cited alongside, same era.
A diverse corpus for evaluating and developing english math word problem solvers
Miao, S.-Y., Liang, C.-C., and Su, K.-Y. (2021) · 2021
Cited alongside, same era.
Are nlp models really able to solve simple math word problems?
Patel, A., Bhattamishra, S., and Goyal, N. (2021) · 2021
Cited alongside, same era.
Ethical and social risks of harm from language models
Weidinger, L., Mellor, J., Rauh, M., Griffin, C., Uesato, J., Huang, P.-S., Cheng, M., Glaese, M., Balle, B., Kasirzadeh, A., Kenton, Z., Brown, S., Hawkins, W., Stepleton, T., Biles, C., Birhane, A., Haas, J., Rimell, L., Hendricks, L. A., Isaac, W., Legassick, S., Irving, G., and Gabriel, I. (2021) · 2021
Cited alongside, same era.
Polyjuice: Generating counterfactuals for explaining, evaluating, and improving models
Wu, T., Ribeiro, M. T., Heer, J., and Weld, D. (2021) · 2021
Cited alongside, same era.
Challenges and applications of large language models
Kaddour, J., Harris, J., Mozes, M., Bradley, H., Raileanu, R., and McHardy, R. (2023) · 2023
Closest in time.
Measuring faithfulness in chain-of-thought reasoning
Lanham, T., Chen, A., Radhakrishnan, A., Steiner, B., Denison, C., Hernandez, D., Li, D., Durmus, E., Hubinger, E., Kernion, J., Lukošiūtė, K., Nguyen, K., Cheng, N., Joseph, N., Schiefer, N., Rausch, O., Larson, R., McCandlish, S., Kundu, S., Kadavath, S., Yang, S., Henighan, T., Maxwell, T., Telleen-Lawton, T., Hume, T., Hatfield-Dodds, Z., Kaplan, J., Brauner, J., Bowman, S. R., and Perez, E. (2023) · 2023
Closest in time.
Faithful chain-of-thought reasoning
Lyu, Q., Havaldar, S., Stein, A., Zhang, L., Rao, D., Wong, E., Apidianaki, M., and Callison-Burch, C. (2023) · 2023
Closest in time.
An overview of bard: an early experiment with generative ai
Manyika, J. (2023) · 2023
Closest in time.
Gpt-4 technical report
OpenAI, R. (2023) · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Teaching models to express their uncertainty in words
Lin, S., Hilton, J., and Evans, O. (2022) · 2022
Cited alongside, same era.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Srivastava, A., Rastogi, A., Rao, A., Shoeb, A. A. M., Abid, A., Fisch, A., Brown, A. R., Santoro, A., Gupta, A., Garriga-Alonso, A., et al. (2022) · 2022
Cited alongside, same era.
Semattack: Natural textual attacks via different semantic spaces
Wang, B., Xu, C., Liu, X., Cheng, Y., and Li, B. (2022) · 2022
Cited alongside, same era.
Model-card-claude-2.pdf
Anthropic · 2023
Cited alongside, same era.
Faithfulness tests for natural language explanations
Atanasova, P., Camburu, O.-M., Lioma, C., Lukasiewicz, T., Simonsen, J. G., and Augenstein, I. (2023) · 2023
Cited alongside, same era.
Survey of hallucination in natural language generation
Ji, Z., Lee, N., Frieske, R., Yu, T., Su, D., Xu, Y., Ishii, E., Bang, Y. J., Madotto, A., and Fung, P. (2023) · 2023
Cited alongside, same era.
Visualizing and understanding neural models in nlp
Li, J., Chen, X., Hovy, E., and Jurafsky, D. (2016a)
Cited in the paper.
Touvron, H. (2023) · 2023
Closest in time.
CREST: A joint framework for rationalization and counterfactual text generation
Treviso, M., Ross, A., Guerreiro, N. M., and Martins, A. (2023) · 2023
Closest in time.
Language models don’t always say what they think: Unfaithful explanations in chain-of-thought prompting
Turpin, M., Michael, J., Perez, E., and Bowman, S. R. (2023) · 2023
Closest in time.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., Chi, E., Le, Q., and Zhou, D. (2023) · 2023
Closest in time.
Can llms express their uncertainty? an empirical evaluation of confidence elicitation in llms
Xiong, M., Hu, Z., Lu, X., Li, Y., Fu, J., He, J., and Hooi, B. (2023) · 2023
Closest in time.
The potential and pitfalls of using a large language model such as chatgpt or gpt-4 as a clinical assistant
Zhang, J., Sun, K., Jagadeesh, A., Ghahfarokhi, M., Gupta, D., Gupta, A., Gupta, V., and Guo, Y. (2023) · 2023
Closest in time.