Fetching the paper…
Reading the bibliography…
Understanding the reliability of large language models (LLMs) has recently garnered significant attention.
When quantization affects confidence of large language models?
Proskurina, I., Brun, L., Metzler, G., and Velcin, J. (2024) · 1928
Earlier work this paper cites.
Principles of Mathematical Analysis
Rudin, W. (1976) · 1976
Earlier work this paper cites.
Maximum likelihood from incomplete data via the EM algorithm
Dempster, A. P., Laird, N. M., and Rubin, D. B. (1977) · 1977
Earlier work this paper cites.
On the limited memory BFGS method for large scale optimization
Liu, D. C. and Nocedal, J. (1989) · 1989
Earlier work this paper cites.
Probabilistic outputs for support vector machines and comparisons to regularized likelihood methods
Platt, J. (1999) · 1999
Earlier work this paper cites.
The skyline operator
Börzsönyi, S., Kossmann, D., and Stocker, K. (2001) · 2001
Earlier work this paper cites.
Statistical Inference
Casella, G. and Berger, R. (2002) · 2002
Earlier work this paper cites.
Transforming classifier scores into accurate multiclass probability estimates
Zadrozny, B. and Elkan, C. (2002) · 2002
Earlier work this paper cites.
An Introduction to Copulas
Nelsen, R. B. (2006) · 2006
Earlier work this paper cites.
Goodness-of-fit tests for copulas: A review and a power study
Genest, C., Rémillard, B., and Beaudoin, D. (2009) · 2009
Earlier work this paper cites.
Obtaining well calibrated probabilities using bayesian binning
Naeini, M. P., Cooper, G. F., and Hauskrecht, M. (2015) · 2015
Earlier work this paper cites.
Taking the human out of the loop: A review of bayesian optimization
Shahriari, B., Swersky, K., Wang, Z., Adams, R. P., and de Freitas, N. (2016) · 2016
Earlier work this paper cites.
On calibration of modern neural networks
Guo, C., Pleiss, G., Sun, Y., and Weinberger, K. Q. (2017) · 2017
Earlier work this paper cites.
TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension
Joshi, M., Choi, E., Weld, D., and Zettlemoyer, L. (2017) · 2017
Earlier work this paper cites.
A baseline for detecting misclassified and out-of-distribution examples in neural networks
Hendrycks, D. and Gimpel, K. (2018) · 2018
Earlier work this paper cites.
Don’t give me the details, just the summary! Topic-aware convolutional neural networks for extreme summarization
Narayan, S., Cohen, S. B., and Lapata, M. (2018) · 2018
Earlier work this paper cites.
Neural text summarization: A critical evaluation
Kryściński, W., Keskar, N. S., McCann, B., Xiong, C., and Socher, R. (2019) · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D. (2020) · 2020
Cited alongside, same era.
Training verifiers to solve math word problems
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., Hesse, C., and Schulman, J. (2021) · 2021
Cited alongside, same era.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J. (2021) · 2021
Cited alongside, same era.
How can we know when language models know? on the calibration of language models for question answering
Jiang, Z., Araki, J., Ding, H., and Neubig, G. (2021) · 2021
Cited alongside, same era.
Hebo: Pushing the limits of sample-efficient hyperparameter optimisation
AutoMix: Automatically mixing language models
Aggarwal, P., Madaan, A., Anand, A., Potharaju, S. P., Mishra, S., Zhou, P., Gupta, A., Rajagopal, D., Kappaganthu, K., Yang, Y., Upadhyay, S., Faruqui, M., and Mausam (2024) · 2024
Later among the works it cites.
Discovering latent knowledge in language models without supervision
Burns, C., Ye, H., Klein, D., and Steinhardt, J. (2024) · 2024
Later among the works it cites.
LLM.int8(): 8-bit matrix multiplication for transformers at scale
Dettmers, T., Lewis, M., Belkada, Y., and Zettlemoyer, L. (2024) · 2024
Later among the works it cites.
Hybrid LLM: Cost-efficient and quality-aware query routing
Ding, D., Mallick, A., Wang, C., Sim, R., Mukherjee, S., Ruhle, V., Lakshmanan, L. V. S., and Awadallah, A. H. (2024) · 2024
Later among the works it cites.
Detecting hallucinations in large language models using semantic entropy
Farquhar, S., Kossen, J., Kuhn, L., and Gal, Y. (2024) · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cowen-Rivers, A., Lyu, W., Tutunov, R., Wang, Z., Grosnit, A., Griffiths, R.-R., Maravel, A., Hao, J., Wang, J., Peters, J., and Bou Ammar, H. (2022) · 2022
Cited alongside, same era.
Language models (mostly) know what they know
Kadavath, S., Conerly, T., Askell, A., Henighan, T., Drain, D., Perez, E., Schiefer, N., Hatfield-Dodds, Z., DasSarma, N., Tran-Johnson, E., Johnston, S., El-Showk, S., Jones, A., Elhage, N., Hume, T., Chen, A., Bai, Y., Bowman, S., Fort, S., Ganguli, D., Hernandez, D., Jacobson, J., Kernion, J., Kravec, S., Lovitt, L., Ndousse, K., Olsson, C., Ringer, S., Amodei, D., Brown, T., Clark, J., Joseph, N., Mann, B., McCandlish, S., Olah, C., and Kaplan, J. (2022) · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P., Leike, J., and Lowe, R. (2022) · 2022
Cited alongside, same era.
MedMCQA: A large-scale multi-subject multi-choice dataset for medical domain question answering
Pal, A., Umapathi, L. K., and Sankarasubbu, M. (2022) · 2022
Cited alongside, same era.
The internal state of an LLM knows when it’s lying
Azaria, A. and Mitchell, T. (2023) · 2023
Cited alongside, same era.
Frugalgpt: How to use large language models while reducing cost and improving performance
Chen, L., Zaharia, M., and Zou, J. (2023) · 2023
Cited alongside, same era.
Tryage: Real-time, intelligent routing of user prompts to large language models
Hari, S. N. and Thomson, M. (2023) · 2023
Cited alongside, same era.
LLM-blender: Ensembling large language models with pairwise ranking and generative fusion
Jiang, D., Ren, X., and Lin, B. Y. (2023) · 2023
Cited alongside, same era.
Gupta, N., Narasimhan, H., Jitkrittum, W., Rawat, A. S., Menon, A. K., and Kumar, S. (2024) · 2024
Later among the works it cites.
When does confidence-based cascade deferral suffice?
Jitkrittum, W., Gupta, N., Menon, A. K., Narasimhan, H., Rawat, A. S., and Kumar, S. (2024) · 2024
Later among the works it cites.
Semantic entropy probes: Robust and cheap hallucination detection in LLMs
Kossen, J., Han, J., Razzak, M., Schut, L., Malik, S., and Gal, Y. (2024) · 2024
Later among the works it cites.
Generating with confidence: Uncertainty quantification for black-box large language models
Lin, Z., Trivedi, S., and Sun, J. (2024) · 2024
Later among the works it cites.
GPT-4 Technical Report
OpenAI (2024) · 2024
Later among the works it cites.
Softmax probabilities (mostly) predict large language model correctness on multiple-choice q&a
Plaut, B., Nguyen, K., and Trinh, T. (2024) · 2024
Later among the works it cites.
Fly-swat or cannon? cost-effective language model choice via meta-modeling
Sakota, M., Peyrard, M., and West, R. (2024) · 2024
Later among the works it cites.
Cascade-aware training of language models
Wang, C., Augenstein, S., Rush, K., Jitkrittum, W., Narasimhan, H., Rawat, A. S., Menon, A. K., and Go, A. (2024) · 2024
Later among the works it cites.
Can LLMs express their uncertainty? an empirical evaluation of confidence elicitation in LLMs
Xiong, M., Hu, Z., Lu, X., Li, Y., Fu, J., He, J., and Hooi, B. (2024) · 2024
Later among the works it cites.
Large language model cascades with mixture of thoughts representations for cost-efficient reasoning
Yue, M., Zhao, J., Zhang, M., Du, L., and Yao, Z. (2024) · 2024
Later among the works it cites.
Efficiently deploying LLMs with controlled risk
Zellinger, M. J. and Thomson, M. (2024) · 2024
Later among the works it cites.
The shift from models to compound AI systems
Zaharia, M., Khattab, O., Chen, L., Davis, J. Q., Miller, H., Potts, C., Zou, J., Carbin, M., Frankle, J., Rao, N., and Ghodsi, A. (2024) · 2025
Closest in time.