Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have demonstrated impressive performance on several tasks and are increasingly deployed in real-world applications.
Verification of forecasts expressed in terms of probability
Brier, G. W · 1950
Earlier work this paper cites.
Hierarchical mixtures of experts and the em algorithm
Jordan, M. I. and Jacobs, R. A · 1994
Earlier work this paper cites.
Solving general arithmetic word problems
Roy, S. and Roth, D · 2015
Earlier work this paper cites.
Learning with rejection
Cortes, C., DeSalvo, G., and Mohri, M · 2016
Earlier work this paper cites.
On calibration of modern neural networks
Guo, C., Pleiss, G., Sun, Y., and Weinberger, K. Q · 2017
Earlier work this paper cites.
Can a suit of armor conduct electricity? a new dataset for open book question answering
Mihaylov, T., Clark, P., Khot, T., and Sabharwal, A · 2018
Earlier work this paper cites.
Selective question answering under domain shift
Kamath, A., Jia, R., and Liang, P · 2020
Earlier work this paper cites.
Consistent estimators for learning to defer to an expert
Mozannar, H. and Sontag, D · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., Hesse, C., and Schulman, J · 2021
Earlier work this paper cites.
What disease does this patient have? a large-scale open domain question answering dataset from medical exams
Jin, D., Pan, E., Oufattole, N., Weng, W.-H., Fang, H., and Szolovits, P · 2021
Earlier work this paper cites.
Generating soap notes from doctor-patient conversations using modular summarization techniques
Krishna, K., Khosla, S., Bigham, J. P., and Lipton, Z. C · 2021
Earlier work this paper cites.
Knowing more about questions can help: Improving calibration in question answering
Zhang, S., Gong, C., and Choi, E · 2021
Earlier work this paper cites.
Langchain-ai/langchain: build context-aware reasoning applications, Oct 2022
Harrison, C · 2022
Earlier work this paper cites.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2023
Earlier work this paper cites.
A unifying theory of distance from calibration
Błasiok, J., Gopalan, P., Hu, L., and Nakkiran, P · 2023
Earlier work this paper cites.
Look before you leap: An exploratory study of uncertainty measurement for large language models
Huang, Y., Song, J., Wang, Z., Zhao, S., Chen, H., Juefei-Xu, F., and Ma, L · 2023
Earlier work this paper cites.
Mistral 7b
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., de las Casas, D., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., Lavaud, L. R., Lachaux, M.-A., Stock, P., Scao, T. L., Lavril, T., Wang, T., Lacroix, T., and Sayed, W. E · 2023
Earlier work this paper cites.
Optimizing natural language processing, large language models (llms) for efficient customer service, and hyper-personalization to enable sustainable growth and revenue
Kolasani, S · 2023
Earlier work this paper cites.
Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation
Kuhn, L., Gal, Y., and Farquhar, S · 2023
Earlier work this paper cites.
Fast inference from transformers via speculative decoding
Leviathan, Y., Kalman, M., and Matias, Y · 2023
Earlier work this paper cites.
Large language models with controllable working memory
Li, D., Rawat, A. S., Zaheer, M., Wang, X., Lukasik, M., Veit, A., Felix, X. Y., and Kumar, S · 2023
Cited alongside, same era.
Generating with confidence: Uncertainty quantification for black-box large language models
Lin, Z., Trivedi, S., and Sun, J · 2023
Cited alongside, same era.
Selfcheckgpt: Zero-resource black-box hallucination detection for generative large language models
Manakul, P., Liusie, A., and Gales, M · 2023
Cited alongside, same era.
Structured prediction with stronger consistency guarantees
Mao, A., Mohri, M., and Zhong, Y · 2023
Cited alongside, same era.
Disentqa: Disentangling parametric and contextual knowledge with counterfactual question answering
Neeman, E., Aharoni, R., Honovich, O., Choshen, L., Szpektor, I., and Abend, O · 2023
Cited alongside, same era.
Ruffle&riley: Towards the automated induction of conversational tutoring systems
Shifting attention to relevance: Towards the predictive uncertainty quantification of free-form large language models
Duan, J., Cheng, H., Wang, S., Zavalny, A., Wang, C., Xu, R., Kailkhura, B., and Xu, K · 2024
Closest in time.
Lora+: Efficient low rank adaptation of large models
Hayou, S., Ghosh, N., and Yu, B · 2024
Closest in time.
Decomposing uncertainty for large language models through input clarification ensembling
Hou, B., Liu, Y., Qian, K., Andreas, J., Chang, S., and Zhang, Y · 2024
Closest in time.
Routerbench: A benchmark for multi-llm routing system
Hu, Q. J., Bieker, J., Li, X., Jiang, N., Keigwin, B., Ranganath, G., Keutzer, K., and Upadhyay, S. K · 2024
Closest in time.
Context-aware assistant selection for improved inference acceleration with large language models
Huang, J., Parthasarathi, P., Rezagholizadeh, M., and Chandar, S · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Schmucker, R., Xia, M., Azaria, A., and Mitchell, T · 2023
Cited alongside, same era.
Just ask for calibration: Strategies for eliciting calibrated confidence scores from language models fine-tuned with human feedback
Tian, K., Mitchell, E., Zhou, A., Sharma, A., Rafailov, R., Yao, H., Finn, C., and Manning, C. D · 2023
Cited alongside, same era.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al · 2023
Cited alongside, same era.
Strength in numbers: Estimating confidence of large language models by prompt agreement
Wightman, G. P., Delucia, A., and Dredze, M · 2023
Cited alongside, same era.
Autogen: Enabling next-gen llm applications via multi-agent conversation
Wu, Q., Bansal, G., Zhang, J., Wu, Y., Li, B., Zhu, E., Jiang, L., Zhang, X., Zhang, S., Liu, J., et al · 2023
Cited alongside, same era.
Evaluating reading comprehension exercises generated by llms: A showcase of chatgpt in education applications
Xiao, C., Xu, S. X., Zhang, K., Wang, Y., and Xia, L · 2023
Cited alongside, same era.
Large language model cascades with mixture of thought representations for cost-efficient reasoning
Yue, M., Zhao, J., Zhang, M., Du, L., and Yao, Z · 2023
Cited alongside, same era.
Jeong, H., Jabbour, S., Yang, Y., Thapta, R., Mozannar, H., Han, W. J., Mehandru, N., Wornow, M., Lialin, V., Liu, X., et al · 2024
Closest in time.
Jiang, A. Q., Sablayrolles, A., Roux, A., Mensch, A., Savary, B., Bamford, C., Chaplot, D. S., Casas, D. d. l., Hanna, E. B., Bressand, F., et al · 2024
Closest in time.
When no-rejection learning is consistent for regression with rejection
Li, X., Liu, S., Sun, C., and Wang, H · 2024
Closest in time.
Uncertainty estimation and quantification for llms: A simple supervised approach
Liu, L., Pan, Y., Li, X., and Chen, G · 2024
Closest in time.
Factual confidence of llms: on reliability and robustness of current estimators
Mahaut, M., Aina, L., Czarnowska, P., Hardalov, M., Mueller, T., and Màrquez, L · 2024
Closest in time.
Two-stage learning to defer with multiple experts
Mao, A., Mohri, C., Mohri, M., and Zhong, Y · 2024
Closest in time.
Learning to reject with a fixed predictor: Application to decontextualization
Mohri, C., Andor, D., Choi, E., Collins, M., Mao, A., and Zhong, Y · 2024
Closest in time.
Routellm: Learning to route llms from preference data
Ong, I., Almahairi, A., Wu, V., Chiang, W.-L., Wu, T., Gonzalez, J. E., Kadous, M. W., and Stoica, I · 2024
Closest in time.
Polyrouter: A multi-llm querying system
Stripelis, D., Hu, Z., Zhang, J., Xu, Z., Shah, A. D., Jin, H., Yao, Y., Avestimehr, S., and He, C · 2024
Closest in time.
Language models don’t always say what they think: unfaithful explanations in chain-of-thought prompting
Turpin, M., Michael, J., Perez, E., and Bowman, S · 2024
Closest in time.
Chain-of-thought reasoning without prompting
Wang, X. and Zhou, D · 2024
Closest in time.
Harnessing the power of llms in practice: A survey on chatgpt and beyond
Yang, J., Jin, H., Tang, R., Han, X., Feng, Q., Jiang, H., Zhong, S., Yin, B., and Hu, X · 2024
Closest in time.
Can large language models faithfully express their intrinsic uncertainty in words?
Yona, G., Aharoni, R., and Geva, M · 2024
Closest in time.
R-tuning: Instructing large language models to say ‘I don’t know’
Zhang, H., Diao, S., Lin, Y., Fung, Y., Lian, Q., Wang, X., Chen, Y., Ji, H., and Zhang, T · 2024
Closest in time.
Chuang, Y.-N., Yu, L., Wang, G., Zhang, L., Liu, Z., Cai, X., Sui, Y., Braverman, V., and Hu, X · 2025
Closest in time.