Fetching the paper…
Reading the bibliography…
We propose semantic entropy probes (SEPs), a cheap and reliable method for uncertainty quantification in Large Language Models (LLMs).
Classification and regression trees
Loh, W.-Y · 2011
Earlier work this paper cites.
Scikit-learn: Machine learning in Python
Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., et al · 2011
Earlier work this paper cites.
An overview of the BIOASQ large-scale biomedical semantic indexing and question answering competition
Tsatsaronis, G., Balikas, G., Malakasiotis, P., Partalas, I., Zschunke, M., Alvers, M. R., Weissenborn, D., Krithara, A., Petridis, S., Polychronopoulos, D., Almirantis, Y., Pavlopoulos, J., Baskiotis, N., Gallinari, P., Artiéres, T., Ngomo, A.-C. N., Heino, N., Gaussier, E., Barrio-Alvers, L., Schroeder, M., Androutsopoulos, I., and Paliouras, G · 2015
Earlier work this paper cites.
Understanding intermediate layers using linear classifier probes
Alain, G. and Bengio, Y · 2017
Earlier work this paper cites.
TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension
Joshi, M., Choi, E., Weld, D. S., and Zettlemoyer, L · 2017
Earlier work this paper cites.
Correcting length bias in neural machine translation
Murray, K. and Chiang, D · 2018
Earlier work this paper cites.
Know what you don’t know: Unanswerable questions for squad
Rajpurkar, P., Jia, R., and Liang, P · 2018
Earlier work this paper cites.
Natural questions: a benchmark for question answering research
Kwiatkowski, T., Palomaki, J., Redfield, O., Collins, M., Parikh, A., Alberti, C., Epstein, D., Polosukhin, I., Kelcey, M., Devlin, J., Lee, K., Toutanova, K. N., Jones, L., Chang, M.-W., Dai, A., Uszkoreit, J., Le, Q., and Petrov, S · 2019
Earlier work this paper cites.
Language models as knowledge bases?
Petroni, F., Rocktäschel, T., Lewis, P., Bakhtin, A., Wu, Y., Miller, A. H., and Riedel, S · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Feqa: A question answering evaluation framework for faithfulness assessment in abstractive summarization
Durmus, E., He, H., and Diab, M · 2020
Earlier work this paper cites.
Controlled hallucinations: Learning to generate faithfully from noisy data
Filippova, K · 2020
Earlier work this paper cites.
On faithfulness and factuality in abstractive summarization
Maynez, J., Narayan, S., Bohnet, B., and McDonald, R · 2020
Earlier work this paper cites.
How much knowledge can you pack into the parameters of a language model?
Roberts, A., Raffel, C., and Shazeer, N · 2020
Earlier work this paper cites.
Asking and answering questions to evaluate the factual consistency of summaries
Wang, A., Cho, K., and Lewis, M · 2020
Earlier work this paper cites.
Probing classifiers: Promises, shortcomings, and advances
Belinkov, Y · 2021
Earlier work this paper cites.
Towards question-answering as an automatic metric for evaluating the content quality of a summary
Deutsch, D., Bedrax-Weiss, T., and Roth, D · 2021
Earlier work this paper cites.
Neural path hunter: Reducing hallucination in dialogue systems via path grounding
Dziri, N., Madotto, A., Zaïane, O., and Bose, A. J · 2021
Earlier work this paper cites.
Deberta: Decoding-enhanced bert with disentangled attention
He, P., Liu, X., Gao, J., and Chen, W · 2021
Earlier work this paper cites.
Uncertainty estimation in autoregressive structured prediction
Malinin, A. and Gales, M · 2021
Earlier work this paper cites.
Improving factual consistency of abstractive summarization via question answering
Nan, F., Santos, C. N. d., Zhu, H., Ng, P., McKeown, K., Nallapati, R., Zhang, D., Wang, Z., Arnold, A. O., and Xiang, B · 2021
Earlier work this paper cites.
Hallucinated but factual! inspecting the factuality of hallucinations in abstractive summarization
Cao, M., Dong, Y., and Cheung, J. C. K · 2022
Earlier work this paper cites.
Rarr: Researching and revising what language models say, using language models
Gao, L., Dai, Z., Pasupat, P., Chen, A., Chaganty, A. T., Fan, Y., Zhao, V. Y., Lao, N., Lee, H., Juan, D.-C., et al · 2022
Earlier work this paper cites.
Language models (mostly) know what they know
Kadavath, S., Conerly, T., Askell, A., Henighan, T., Drain, D., Perez, E., Schiefer, N., Dodds, Z. H., DasSarma, N., Tran-Johnson, E., et al · 2022
Earlier work this paper cites.
Factuality enhanced language models for open-ended text generation
Lee, N., Ping, W., Xu, P., Patwary, M., Fung, P. N., Shoeybi, M., and Catanzaro, B · 2022
Cited alongside, same era.
Reducing conversational agents’ overconfidence through linguistic calibration
Mielke, S. J., Szlam, A., Boureau, Y.-L., and Dinan, E · 2022
Cited alongside, same era.
Read before generate! faithful long form question answering with machine reading
Su, D., Li, X., Zhang, J., Shang, L., Jiang, X., Liu, Q., and Fung, P · 2022
Cited alongside, same era.
Extracting latent steering vectors from pretrained language models
Subramani, N., Suresh, N., and Peters, M. E · 2022
Cited alongside, same era.
The internal state of an llm knows when it’s lying
Azaria, A. and Mitchell, T · 2023
Cited alongside, same era.
Eliciting latent predictions from transformers with the tuned lens
Belrose, N., Furman, Z., Smith, L., Halawi, D., Ostrovsky, I., McKinney, L., Biderman, S., and Steinhardt, J · 2023
Self-contradictory hallucinations of large language models: Evaluation, detection and mitigation
Mündler, N., He, J., Jenko, S., and Vechev, M · 2023
Later among the works it cites.
Fact finding: Attempting to reverse-engineer factual recall on the neuron level, Dec 2023
Nanda, N., Rajamanoharan, S., Kramar, J., and Shah, R · 2023
Later among the works it cites.
Trustworthy journalism through AI
Opdahl, A. L., Tessem, B., Dang-Nguyen, D.-T., Motta, E., Setty, V., Throndsen, E., Tverberg, A., and Trattner, C · 2023
Later among the works it cites.
GPT-4 technical report, 2023
OpenAI · 2023
Later among the works it cites.
Future lens: Anticipating subsequent tokens from a single hidden state
Pal, K., Sun, J., Yuan, A., Wallace, B., and Bau, D · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Discovering latent knowledge in language models without supervision
Burns, C., Ye, H., Klein, D., and Steinhardt, J · 2023
Cited alongside, same era.
Quantifying uncertainty in answers from any language model and enhancing their trustworthiness
Chen, J. and Mueller, J · 2023
Cited alongside, same era.
Selectively answering ambiguous questions
Cole, J. R., Zhang, M. J., Gillick, D., Eisenschlos, J. M., Dhingra, B., and Eisenstein, J · 2023
Cited alongside, same era.
Chain-of-verification reduces hallucination in large language models
Dhuliawala, S., Komeili, M., Xu, J., Raileanu, R., Li, X., Celikyilmaz, A., and Weston, J · 2023
Cited alongside, same era.
Shifting attention to relevance: Towards the uncertainty estimation of large language models
Duan, J., Cheng, H., Wang, S., Wang, C., Zavalny, A., Xu, R., Kailkhura, B., and Xu, K · 2023
Cited alongside, same era.
Halo: Estimation and reduction of hallucinations in open-source weak large language models
Elaraby, M., Lu, M., Dunn, J., Zhang, X., Wang, Y., and Liu, S · 2023
Cited alongside, same era.
Peng, B., Galley, M., He, P., Cheng, H., Xie, Y., Hu, Y., Huang, Q., Liden, L., Yu, Z., Chen, W., et al · 2023
Later among the works it cites.
A survey of hallucination in large foundation models
Rawte, V., Sheth, A., and Das, A · 2023
Later among the works it cites.
Steering llama 2 via contrastive activation addition
Rimsky, N., Gabrieli, N., Schulz, J., Tong, M., Hubinger, E., and Turner, A. M · 2023
Later among the works it cites.
ChatGPT and other large language models are double-edged swords
Shen, Y., Heacock, L., Elias, J., Hentel, K. D., Reig, B., Shih, G., and Moy, L · 2023
Later among the works it cites.
Trusting your evidence: Hallucinate less with context-aware decoding
Shi, W., Han, X., Lewis, M., Tsvetkov, Y., Zettlemoyer, L., and Yih, S. W.-t · 2023
Later among the works it cites.
Large language models encode clinical knowledge
Singhal, K., Azizi, S., Tu, T., Mahdavi, S. S., Wei, J., Chung, H. W., Scales, N., Tanwani, A., Cole-Lewis, H., Pfohl, S., Payne, P., Seneviratne, M., Gamble, P., Kelly, C., Scharli, N., Chowdhery, A., Mansfield, P., y Arcas, B. A., Webster, D., Corrado, G. S., Matias, Y., Chou, K., Gottweis, J., Tomasev, N., Liu, Y., Rajkomar, A., Barral, J., Semturs, C., Karthikesalingam, A., and Natarajan, V · 2023
Later among the works it cites.
Gemini: a family of highly capable multimodal models
Team, T. G · 2023
Later among the works it cites.
Language models don’t always say what they think: unfaithful explanations in chain-of-thought prompting
Turpin, M., Michael, J., Perez, E., and Bowman, S · 2023
Later among the works it cites.
A stitch in time saves nine: Detecting and mitigating hallucinations of llms by actively validating low-confidence generation
Varshney, N., Yao, W., Zhang, H., Chen, J., and Yu, D · 2023
Later among the works it cites.
Lawyer who used ChatGPT faces penalty for made up citations
Weiser, B · 2023
Later among the works it cites.
Representation engineering: A top-down approach to ai transparency
Zou, A., Phan, L., Chen, S., Campbell, J., Guo, P., Ren, R., Pan, A., Yin, X., Mazeika, M., Dombrowski, A.-K., et al · 2023
Later among the works it cites.
Do language models know when they’re hallucinating references?
Agrawal, A., Mackey, L., and Kalai, A. T · 2024
Closest in time.
Linguistic calibration of language models
Band, N., Li, X., Ma, T., and Hashimoto, T · 2024
Closest in time.
Dola: Decoding by contrasting layers improves factuality in large language models
Chuang, Y.-S., Xie, Y., Luo, H., Kim, Y., Glass, J., and He, P · 2024
Closest in time.
Detecting Hallucinations in Large Language Models Using Semantic Entropy
Farquhar, S., Kossen, J., Kuhn, L., and Gal, Y · 2024
Closest in time.
Inference-time intervention: Eliciting truthful answers from a language model
Li, K., Patel, O., Viégas, F., Pfister, H., and Wattenberg, M · 2024
Closest in time.
Simple probes can catch sleeper agents, 2024
MacDiarmid, M., Maxwell, T., Schiefer, N., Mu, J., Kaplan, J., Duvenaud, D., Bowman, S., Tamkin, A., Perez, E., Sharma, M., Denison, C., and Hubinger, E · 2024
Closest in time.
Introducing meta llama 3: The most capable openly available llm to date, 2024
Meta · 2024
Closest in time.
Fine-tuning language models for factuality
Tian, K., Mitchell, E., Yao, H., Manning, C. D., and Finn, C · 2024
Closest in time.