Fetching the paper…
Reading the bibliography…
We study how to characterize and predict the truthfulness of texts generated from large language models (LLMs), which serves as a crucial step in building trust between humans and LLMs.
Crafting papers on machine learning
Langley, P · 2000
Earlier work this paper cites.
Maximum likelihood estimation of intrinsic dimension
Levina, E. and Bickel, P · 2004
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Lin, C.-Y · 2004
Earlier work this paper cites.
Estimating local intrinsic dimension with k-nearest neighbor graphs
Costa, J. A., Girotra, A., and Hero, A. O · 2005
Earlier work this paper cites.
Visualizing data using t-sne
Van der Maaten, L. and Hinton, G · 2008
Earlier work this paper cites.
Estimating local intrinsic dimensionality
Amsaleg, L., Chelly, O., Furon, T., Girard, S., Houle, M. E., Kawarabayashi, K.-i., and Nett, M · 2015
Earlier work this paper cites.
Intrinsic dimension estimation: Relevant techniques and a benchmark framework
Campadelli, P., Casiraghi, E., Ceruti, C., Rozza, A., et al · 2015
Earlier work this paper cites.
Dropout as a bayesian approximation: Representing model uncertainty in deep learning
Gal, Y. and Ghahramani, Z · 2016
Earlier work this paper cites.
Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension
Joshi, M., Choi, E., Weld, D. S., and Zettlemoyer, L · 2017
Earlier work this paper cites.
Characterizing adversarial subspaces using local intrinsic dimensionality
Ma, X., Li, B., Wang, Y., Erfani, S. M., Wijewickrema, S., Schoenebeck, G., Song, D., Houle, M. E., and Bailey, J · 2018
Earlier work this paper cites.
Hotpotqa: A dataset for diverse, explainable multi-hop question answering
Yang, Z., Qi, P., Zhang, S., Bengio, Y., Cohen, W., Salakhutdinov, R., and Manning, C. D · 2018
Earlier work this paper cites.
Intrinsic dimension of data representations in deep neural networks
Ansuini, A., Laio, A., Macke, J. H., and Zoccolan, D · 2019
Earlier work this paper cites.
Geometry-aware maximum likelihood estimation of intrinsic dimension
Gomtsyan, M., Mokrov, N., Panov, M., and Yanovich, Y · 2019
Earlier work this paper cites.
Coqa: A conversational question answering challenge
Reddy, S., Chen, D., and Manning, C. D · 2019
Earlier work this paper cites.
Tydi qa: A benchmark for information-seeking question answering in ty pologically di verse languages
Clark, J. H., Choi, E., Collins, M., Garrette, D., Kwiatkowski, T., Nikolaev, V., and Palomaki, J · 2020
Cited alongside, same era.
Selective question answering under domain shift
Kamath, A., Jia, R., and Liang, P · 2020
Cited alongside, same era.
Uncertainty estimation in autoregressive structured prediction
Malinin, A. and Gales, M · 2020
Cited alongside, same era.
The intrinsic dimension of images and its impact on learning
Pope, P., Zhu, C., Abdelkader, A., Goldblum, M., and Goldstein, T · 2020
Cited alongside, same era.
Intrinsic dimension, persistent homology and generalization in neural networks
Birdal, T., Lou, A., Guibas, L. J., and Simsekli, U · 2021
Cited alongside, same era.
The internal state of an llm knows when its lying
Azaria, A. and Mitchell, T · 2023
Later among the works it cites.
Palm: Scaling language modeling with pathways
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al · 2023
Later among the works it cites.
Shifting attention to relevance: Towards the uncertainty estimation of large language models
Duan, J., Cheng, H., Wang, S., Wang, C., Zavalny, A., Xu, R., Kailkhura, B., and Xu, K · 2023
Later among the works it cites.
Survey of hallucination in natural language generation
Ji, Z., Lee, N., Frieske, R., Yu, T., Su, D., Xu, Y., Ishii, E., Bang, Y. J., Madotto, A., and Fung, P · 2023
Later among the works it cites.
Inference-time intervention: Eliciting truthful answers from a language model
Li, K., Patel, O., Viégas, F., Pfister, H., and Wattenberg, M · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M. S., Bohg, J., Bosselut, A., Brunskill, E., et al · 2021
Cited alongside, same era.
Unsolved problems in ml safety
Hendrycks, D., Carlini, N., Schulman, J., and Steinhardt, J · 2021
Cited alongside, same era.
Discovering latent knowledge in language models without supervision
Burns, C., Ye, H., Klein, D., and Steinhardt, J · 2022
Cited alongside, same era.
TRUE: Re-evaluating factual consistency evaluation
Honovich, O., Aharoni, R., Herzig, J., Taitelbaum, H., Kukliansy, D., Cohen, V., Scialom, T., Szpektor, I., Hassidim, A., and Matias, Y · 2022
Cited alongside, same era.
Language models (mostly) know what they know, 2022
Kadavath, S., Conerly, T., Askell, A., Henighan, T., Drain, D., Perez, E., Schiefer, N., Hatfield-Dodds, Z., DasSarma, N., Tran-Johnson, E., Johnston, S., El-Showk, S., Jones, A., Elhage, N., Hume, T., Chen, A., Bai, Y., Bowman, S., Fort, S., Ganguli, D., Hernandez, D., Jacobson, J., Kernion, J., Kravec, S., Lovitt, L., Ndousse, K., Olsson, C., Ringer, S., Amodei, D., Brown, T., Clark, J., Joseph, N., Mann, B., McCandlish, S., Olah, C., and Kaplan, J · 2022
Cited alongside, same era.
Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation
Kuhn, L., Gal, Y., and Farquhar, S · 2022
Cited alongside, same era.
Out-of-distribution detection and selective generation for conditional language models
Ren, J., Luo, J., Zhao, Y., Krishna, K., Saleh, M., Lakshminarayanan, B., and Liu, P. J · 2022
Cited alongside, same era.
Later among the works it cites.
Generating with confidence: Uncertainty quantification for black-box large language models
Lin, Z., Trivedi, S., and Sun, J · 2023
Later among the works it cites.
Marks, S. and Tegmark, M · 2023
Later among the works it cites.
OpenAI · 2023
Later among the works it cites.
Stanford alpaca: An instruction-following llama model
Taori, R., Gulrajani, I., Zhang, T., Dubois, Y., Li, X., Guestrin, C., Liang, P., and Hashimoto, T. B · 2023
Later among the works it cites.
Tian, K., Mitchell, E., Zhou, A., Sharma, A., Rafailov, R., Yao, H., Finn, C., and Manning, C. D · 2023
Later among the works it cites.
Intrinsic dimension estimation for robust detection of ai-generated texts
Tulchinskii, E., Kuznetsov, K., Kushnareva, L., Cherniavskii, D., Barannikov, S., Piontkovskaya, I., Nikolenko, S., and Burnaev, E · 2023
Later among the works it cites.
Can llms express their uncertainty? an empirical evaluation of confidence elicitation in llms
Xiong, M., Hu, Z., Lu, X., Li, Y., Fu, J., He, J., and Hooi, B · 2023
Later among the works it cites.
Navigating the grey area: Expressions of overconfidence and uncertainty in language models
Zhou, K., Jurafsky, D., and Hashimoto, T · 2023
Later among the works it cites.
Representation engineering: A top-down approach to ai transparency
Zou, A., Phan, L., Chen, S., Campbell, J., Guo, P., Ren, R., Pan, A., Yin, X., Mazeika, M., Dombrowski, A.-K., et al · 2023
Later among the works it cites.