Fetching the paper…
Reading the bibliography…
Guaranteeing the correctness and factuality of language model (LM) outputs is a major open problem.
Conditional validity of inductive conformal predictors
Vovk, V. (2012) · 2012
Earlier work this paper cites.
Conformal prediction for reliable machine learning: theory, adaptations and applications
Balasubramanian, V., Ho, S.-S., and Vovk, V. (2014) · 2014
Earlier work this paper cites.
Natural logic and natural language inference
MacCartney, B. and Manning, C. D. (2015) · 2015
Earlier work this paper cites.
Multi-step retriever-reader interaction for scalable open-domain question answering
Das, R., Dhuliawala, S., Zaheer, M., and McCallum, A. (2019) · 2019
Earlier work this paper cites.
Natural questions: A benchmark for question answering research
Kwiatkowski, T., Palomaki, J., Redfield, O., Collins, M., Parikh, A., Alberti, C., Epstein, D., Polosukhin, I., Devlin, J., Lee, K., Toutanova, K., Jones, L., Kelcey, M., Chang, M.-W., Dai, A. M., Uszkoreit, J., Le, Q., and Petrov, S. (2019) · 2019
Earlier work this paper cites.
Improved natural language generation via loss truncation
Kang, D. and Hashimoto, T. B. (2020) · 2020
Earlier work this paper cites.
Dense passage retrieval for open-domain question answering
Karpukhin, V., Oğuz, B., Min, S., Lewis, P., Wu, L., Edunov, S., Chen, D., and tau Yih, W. (2020) · 2020
Earlier work this paper cites.
On faithfulness and factuality in abstractive summarization
Maynez, J., Narayan, S., Bohnet, B., and McDonald, R. (2020) · 2020
Earlier work this paper cites.
Conformal prediction under covariate shift
Tibshirani, R. J., Barber, R. F., Candes, E. J., and Ramdas, A. (2020) · 2020
Earlier work this paper cites.
Measuring mathematical problem solving with the math dataset
Hendrycks, D., Burns, C., Kadavath, S., Arora, A., Basart, S., Tang, E., Song, D., and Steinhardt, J. (2021) · 2021
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks
Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., tau Yih, W., Rocktäschel, T., Riedel, S., and Kiela, D. (2021) · 2021
Earlier work this paper cites.
A gentle introduction to conformal prediction and distribution-free uncertainty quantification
Angelopoulos, A. N. and Bates, S. (2022) · 2022
Earlier work this paper cites.
Training uncertainty-aware classifiers with conformalized deep learning
Einbinder, B.-S., Romano, Y., Sesia, M., and Zhou, Y. (2022) · 2022
Earlier work this paper cites.
Nested conformal prediction and quantile out-of-bag ensemble methods
Gupta, C., Kuchibhotla, A. K., and Ramdas, A. (2022) · 2022
Earlier work this paper cites.
Rethinking with retrieval: Faithful large language model inference
He, H., Zhang, H., and Roth, D. (2022) · 2022
Earlier work this paper cites.
Emergent abilities of large language models
Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S., Yogatama, D., Bosma, M., Zhou, D., Metzler, D., Chi, E. H., Hashimoto, T., Vinyals, O., Liang, P., Dean, J., and Fedus, W. (2022) · 2022
Earlier work this paper cites.
Conformal risk control
Angelopoulos, A. N., Bates, S., Fisch, A., Lei, L., and Schuster, T. (2023) · 2023
Cited alongside, same era.
Conformal prediction beyond exchangeability
Barber, R. F., Candes, E. J., Ramdas, A., and Tibshirani, R. J. (2023) · 2023
Cited alongside, same era.
Sparks of artificial general intelligence: Early experiments with gpt-4
Bubeck, S., Chandrasekaran, V., Eldan, R., Gehrke, J., Horvitz, E., Kamar, E., Lee, P., Lee, Y. T., Li, Y., Lundberg, S., Nori, H., Palangi, H., Ribeiro, M. T., and Zhang, Y. (2023) · 2023
Cited alongside, same era.
Hallucination is the last thing you need
Curran, S., Lansley, S., and Bethell, O. (2023) · 2023
Cited alongside, same era.
Class-conditional conformal prediction with many classes
Ding, T., Angelopoulos, A. N., Bates, S., Jordan, M. I., and Tibshirani, R. J. (2023) · 2023
Cited alongside, same era.
Improving factuality and reasoning in language models through multiagent debate
Du, Y., Li, S., Torralba, A., Tenenbaum, J. B., and Mordatch, I. (2023) · 2023
FActScore: Fine-grained atomic evaluation of factual precision in long form text generation
Min, S., Krishna, K., Lyu, X., Lewis, M., Yih, W.-t., Koh, P. W., Iyyer, M., Zettlemoyer, L., and Hajishirzi, H. (2023) · 2023
Later among the works it cites.
Learning to reject with a fixed predictor: Application to decontextualization
Mohri, C., Andor, D., Choi, E., Collins, M., Mao, A., and Zhong, Y. (2023) · 2023
Later among the works it cites.
Gpt-4 technical report
OpenAI (2023) · 2023
Later among the works it cites.
Conformal language modeling
Quach, V., Fisch, A., Schuster, T., Yala, A., Sohn, J. H., Jaakkola, T. S., and Barzilay, R. (2023) · 2023
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J. (2023) · 2023
Later among the works it cites.
Conformal nucleus sampling
Ravfogel, S., Goldberg, Y., and Goldberger, J. (2023) · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Conformal prediction with conditional guarantees
Gibbs, I., Cherian, J. J., and Candès, E. J. (2023) · 2023
Cited alongside, same era.
Language models hallucinate, but may excel at fact verification
Guan, J., Dodge, J., Wadden, D., Huang, M., and Peng, H. (2023) · 2023
Cited alongside, same era.
Survey of hallucination in natural language generation
Ji, Z., Lee, N., Frieske, R., Yu, T., Su, D., Xu, Y., Ishii, E., Bang, Y. J., Madotto, A., and Fung, P. (2023) · 2023
Cited alongside, same era.
Kuhn, L., Gal, Y., and Farquhar, S. (2023) · 2023
Cited alongside, same era.
Conformal prediction with large language models for multi-choice question answering
Kumar, B., Lu, C., Gupta, G., Palepu, A., Bellamy, D., Raskar, R., and Beam, A. (2023) · 2023
Cited alongside, same era.
Factuality enhanced language models for open-ended text generation
Lee, N., Ping, W., Xu, P., Patwary, M., Fung, P., Shoeybi, M., and Catanzaro, B. (2023) · 2023
Cited alongside, same era.
Later among the works it cites.
Robots that ask for help: Uncertainty alignment for large language model planners
Ren, A. Z., Dixit, A., Bodrova, A., Singh, S., Tu, S., Brown, N., Xu, P., Takayama, L., Xia, F., Varley, J., Xu, Z., Sadigh, D., Zeng, A., and Majumdar, A. (2023) · 2023
Later among the works it cites.
WikiChat: Stopping the hallucination of large language model chatbots by few-shot grounding on Wikipedia
Semnani, S., Yao, V., Zhang, H., and Lam, M. (2023) · 2023
Later among the works it cites.
Trusting your evidence: Hallucinate less with context-aware decoding
Shi, W., Han, X., Lewis, M., Tsvetkov, Y., Zettlemoyer, L., and Yih, S. (2023) · 2023
Later among the works it cites.
Aligning factual consistency for clinical studies summarization through reinforcement learning
Tang, X., Cohan, A., and Gerstein, M. (2023) · 2023
Later among the works it cites.
Large language models in medicine
Thirunavukarasu, A. J., Ting, D. S. J., Elangovan, K., Gutierrez, L., Tan, T. F., and Ting, D. S. W. (2023) · 2023
Later among the works it cites.
Just ask for calibration: Strategies for eliciting calibrated confidence scores from language models fine-tuned with human feedback
Tian, K., Mitchell, E., Zhou, A., Sharma, A., Rafailov, R., Yao, H., Finn, C., and Manning, C. D. (2023) · 2023
Later among the works it cites.
Large language models for robotics: A survey
Zeng, F., Gan, W., Wang, Y., Liu, N., and Yu, P. S. (2023) · 2023
Later among the works it cites.
Can ai assistants know what they don’t know?
Cheng, Q., Sun, T., Liu, X., Zhang, W., Yin, Z., Li, S., Li, L., Chen, K., and Qiu, X. (2024) · 2024
Closest in time.
Non-exchangeable conformal language generation with nearest neighbors
Ulmer, D., Zerva, C., and Martins, A. F. T. (2024) · 2024
Closest in time.