Fetching the paper…
Reading the bibliography…
Large language models (LLMs) frequently generate hallucinations-content that deviates from factual accuracy or provided context-posing challenges for diagnosis due to the complex interplay of underlying causes.
Word association norms, mutual information, and lexicography
Church, K. and Hanks, P · 1990
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Bahdanau, D · 2014
Earlier work this paper cites.
Scheduled sampling for sequence prediction with recurrent neural networks
Bengio, S., Vinyals, O., Jaitly, N., and Shazeer, N · 2015
Earlier work this paper cites.
Rationalizing neural predictions
Lei, T., Barzilay, R., and Jaakkola, T · 2016
Earlier work this paper cites.
Visualizing and understanding neural models in nlp
Li, J., Chen, X., Hovy, E., and Jurafsky, D · 2016
Earlier work this paper cites.
"why should I trust you?": Explaining the predictions of any classifier
Ribeiro, M. T., Singh, S., and Guestrin, C · 2016
Earlier work this paper cites.
A unified approach to interpreting model predictions
Lundberg, S · 2017
Earlier work this paper cites.
A unified approach to interpreting model predictions
Scott, M., Su-In, L., et al · 2017
Earlier work this paper cites.
Axiomatic attribution for deep networks
Sundararajan, M., Taly, A., and Yan, Q · 2017
Earlier work this paper cites.
Sanity checks for saliency maps
Adebayo, J., Gilmer, J., Muelly, M., Goodfellow, I., Hardt, M., and Kim, B · 2018
Earlier work this paper cites.
Neural network attributions: A causal perspective
Chattopadhyay, A., Manupriya, P., Sarkar, A., and Balasubramanian, V. N · 2019
Earlier work this paper cites.
Probing neural network comprehension of natural language arguments
Niven, T. and Kao, H.-Y · 2019
Earlier work this paper cites.
Eraser: A benchmark to evaluate rationalized nlp models
DeYoung, J., Jain, S., Rajani, N. F., Lehman, E., Xiong, C., Socher, R., and Wallace, B. C · 2020
Earlier work this paper cites.
Shortcut learning in deep neural networks
Geirhos, R., Jacobsen, J.-H., Michaelis, C., Zemel, R., Brendel, W., Bethge, M., and Wichmann, F. A · 2020
Earlier work this paper cites.
Learning to faithfully rationalize by construction
Jain, S., Wiegreffe, S., Pinter, Y., and Wallace, B. C · 2020
Earlier work this paper cites.
Look at the first sentence: Position bias in question answering
Ko, M., Lee, J., Kim, H., Kim, G., and Kang, J · 2020
Earlier work this paper cites.
Beyond accuracy: Behavioral testing of nlp models with checklist
Ribeiro, M. T., Wu, T., Guestrin, C., and Singh, S · 2020
Earlier work this paper cites.
When explanations lie: Why many modified bp attributions fail
Sixt, L., Granz, M., and Landgraf, T · 2020
Earlier work this paper cites.
An empirical study on robustness to spurious correlations using pre-trained language models
Tu, L., Lalwani, G., Gella, S., and He, H · 2020
Earlier work this paper cites.
Explaining by removing: A unified framework for model explanation
Covert, I., Lundberg, S., and Lee, S.-I · 2021
Earlier work this paper cites.
Towards interpreting and mitigating shortcut learning behavior of nlu models
Du, M., Manjunatha, V., Jain, R., Deshpande, R., Dernoncourt, F., Gu, J., Sun, T., and Hu, X · 2021
Earlier work this paper cites.
Mind the style of text! adversarial and backdoor attacks based on text style transfer
Qi, F., Chen, Y., Zhang, X., Li, M., Liu, Z., and Sun, M · 2021
Earlier work this paper cites.
Discretized integrated gradients for explaining language models
Sanyal, S. and Ren, X · 2021
Earlier work this paper cites.
Rationales for sequential predictions
Vafa, K., Deng, Y., Blei, D., and Rush, A. M · 2021
Cited alongside, same era.
A comparative study of faithfulness metrics for model interpretability methods
Chan, C. S., Kong, H., and Guanqing, L · 2022
Cited alongside, same era.
On the origin of hallucinations in conversational models: Is it the datasets or the models?
Dziri, N., Milton, S., Yu, M., Zaiane, O., and Reddy, S · 2022
Cited alongside, same era.
Logic traps in evaluating attribution scores
Ju, Y., Zhang, Y., Yang, Z., Jiang, Z., Liu, K., and Zhao, J · 2022
Cited alongside, same era.
The disagreement problem in explainable machine learning: A practitioner’s perspective
Krishna, S., Han, T., Gu, A., Wu, S., Jabbari, S., and Lakkaraju, H · 2022
Cited alongside, same era.
Incorporating attribution importance for improving faithfulness metrics
Zhao, Z. and Aletras, N · 2023
Later among the works it cites.
Webarena: A realistic web environment for building autonomous agents
Zhou, S., Xu, F. F., Zhu, H., Zhou, X., Lo, R., Sridhar, A., Cheng, X., Ou, T., Bisk, Y., Fried, D., et al · 2023
Later among the works it cites.
Syntaxshap: Syntax-aware explainability method for text generation
Amara, K., Sevastjanova, R., and El-Assady, M · 2024
Later among the works it cites.
Llms will always hallucinate, and we need to live with this
Banerjee, S., Agarwal, A., and Singla, S · 2024
Later among the works it cites.
The reversal curse: Llms trained on “a is b” fail to learn “b is a”
Berglund, L., Tong, M., Kaufmann, M., Balesni, M., Stickland, A. C., Korbak, T., and Evans, O · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Factuality enhanced language models for open-ended text generation
Lee, N., Ping, W., Xu, P., Patwary, M., Fung, P. N., Shoeybi, M., and Catanzaro, B · 2022
Cited alongside, same era.
How pre-trained language models capture factual knowledge? a causal-inspired analysis
Li, S., Li, X., Shang, L., Dong, Z., Sun, C., Liu, B., Ji, Z., Jiang, X., and Liu, Q · 2022
Cited alongside, same era.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2023
Cited alongside, same era.
Transformers learn to implement preconditioned gradient descent for in-context learning
Ahn, K., Cheng, X., Daneshmand, H., and Sra, S · 2023
Cited alongside, same era.
Sequential integrated gradients: a simple but effective method for explaining language models
Enguehard, J · 2023
Cited alongside, same era.
How do transformers learn topic structure: Towards a mechanistic understanding
Li, Y., Li, Y., and Risteski, A · 2023
Cited alongside, same era.
Mahankali, A., Hashimoto, T. B., and Ma, T · 2023
Cited alongside, same era.
Impossibility theorems for feature attribution
Bilodeau, B., Jaques, N., Koh, P. W., and Kim, B · 2024
Later among the works it cites.
Unveiling induction heads: Provable training dynamics and feature learning in transformers
Chen, S., Sheen, H., Wang, T., and Yang, Z · 2024
Later among the works it cites.
The browsergym ecosystem for web agent research
Chezelles, D., Le Sellier, T., Gasse, M., Lacoste, A., Drouin, A., Caccia, M., Boisvert, L., Thakkar, M., Marty, T., Assouel, R., et al · 2024
Later among the works it cites.
OLMo: Accelerating the science of language models
Groeneveld, D., Beltagy, I., Walsh, E., Bhagia, A., Kinney, R., Tafjord, O., Jha, A., Ivison, H., Magnusson, I., Wang, Y., Arora, S., Atkinson, D., Authur, R., Chandu, K., Cohan, A., Dumas, J., Elazar, Y., Gu, Y., Hessel, J., Khot, T., Merrill, W., Morrison, J., Muennighoff, N., Naik, A., Nam, C., Peters, M., Pyatkin, V., Ravichander, A., Schwenk, D., Shah, S., Smith, W., Strubell, E., Subramani, N., Wortsman, M., Dasigi, P., Lambert, N., Richardson, K., Zettlemoyer, L., Dodge, J., Lo, K., Soldaini, L., Smith, N., and Hajishirzi, H · 2024
Later among the works it cites.
Hijacking context in large multi-modal models
Jeong, J · 2024
Later among the works it cites.
Calibrated language models must hallucinate
Kalai, A. T. and Vempala, S. S · 2024
Later among the works it cites.
Omniact: A dataset and benchmark for enabling multimodal generalist autonomous agents for desktop and web, 2024
Kapoor, R., Butala, Y. P., Russak, M., Koh, J. Y., Kamble, K., Alshikh, W., and Salakhutdinov, R · 2024
Later among the works it cites.
Screenagent: A vision language model-driven computer control agent
Niu, R., Li, J., Wang, S., Fu, Y., Hu, X., Leng, X., Kong, H., Chang, Y., and Wang, Q · 2024
Later among the works it cites.
Androidinthewild: A large-scale dataset for android device control
Rawles, C., Li, A., Rodriguez, D., Riva, O., and Lillicrap, T · 2024
Later among the works it cites.
Dolma: An Open Corpus of Three Trillion Tokens for Language Model Pretraining Research
Soldaini, L., Kinney, R., Bhagia, A., Schwenk, D., Atkinson, D., Authur, R., Bogin, B., Chandu, K., Dumas, J., Elazar, Y., Hofmann, V., Jha, A. H., Kumar, S., Lucy, L., Lyu, X., Lambert, N., Magnusson, I., Morrison, J., Muennighoff, N., Naik, A., Nam, C., Peters, M. E., Ravichander, A., Richardson, K., Shen, Z., Strubell, E., Subramani, N., Tafjord, O., Walsh, P., Zettlemoyer, L., Smith, N. A., Hajishirzi, H., Beltagy, I., Groeneveld, D., Dodge, J., and Lo, K · 2024
Later among the works it cites.
Localizing paragraph memorization in language models
Stoehr, N., Gordon, M., Zhang, C., and Lewis, O · 2024
Later among the works it cites.
Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments, 2024
Xie, T., Zhang, D., Chen, J., Li, X., Zhao, S., Cao, R., Hua, T. J., Cheng, Z., Shin, D., Lei, F., Liu, Y., Xu, Y., Zhou, S., Savarese, S., Xiong, C., Zhong, V., and Yu, T · 2024
Later among the works it cites.
Plausible extractive rationalization through semi-supervised entailment signal
Yeo, W. J., Satapathy, R., and Cambria, E · 2024
Later among the works it cites.
Do llms overcome shortcut learning? an evaluation of shortcut challenges in large language models
Yuan, Y., Zhao, L., Zhang, K., Zheng, G., and Liu, Q · 2024
Later among the works it cites.
Towards faithful explanations: Boosting rationalization with shortcuts discovery
Yue, L., Liu, Q., Du, Y., Wang, L., Gao, W., and An, Y · 2024
Later among the works it cites.
Knowledge overshadowing causes amalgamated hallucination in large language models
Zhang, Y., Li, S., Liu, J., Yu, P., Fung, Y. R., Li, J., Li, M., and Ji, H · 2024
Later among the works it cites.
Reagent: Towards a model-agnostic feature attribution method for generative language models
Zhao, Z. and Shan, B · 2024
Later among the works it cites.
Halogen: Fantastic llm hallucinations and where to find them, 2025
Ravichander, A., Ghela, S., Wadden, D., and Choi, Y · 2025
Closest in time.