Fetching the paper…
Reading the bibliography…
Deploying large language models in production requires simultaneous attention to efficiency and risk control.
Probabilistic outputs for support vector machines and comparisons to regularized likelihood methods
Platt, J. (1999) · 1999
Earlier work this paper cites.
The skyline operator
Börzsönyi, S., Kossmann, D., and Stocker, K. (2001) · 2001
Earlier work this paper cites.
Investigating selective prediction approaches across several tasks in IID, OOD, and adversarial settings
Varshney, N., Mishra, S., and Baral, C. (2022) · 2002
Earlier work this paper cites.
Transforming classifier scores into accurate multiclass probability estimates
Zadrozny, B. and Elkan, C. (2002) · 2002
Earlier work this paper cites.
On the foundations of noise-free selective classification
El-Yaniv, R. and Wiener, Y. (2010) · 2010
Earlier work this paper cites.
Distilling the knowledge in a neural network
Hinton, G., Vinyals, O., and Dean, J. (2015) · 2015
Earlier work this paper cites.
Obtaining well calibrated probabilities using bayesian binning
Naeini, M. P., Cooper, G. F., and Hauskrecht, M. (2015) · 2015
Earlier work this paper cites.
Selective classification for deep neural networks
Geifman, Y. and El-Yaniv, R. (2017) · 2017
Earlier work this paper cites.
On calibration of modern neural networks
Guo, C., Pleiss, G., Sun, Y., and Weinberger, K. Q. (2017) · 2017
Earlier work this paper cites.
A baseline for detecting misclassified and out-of-distribution examples in neural networks
Hendrycks, D. and Gimpel, K. (2018) · 2018
Earlier work this paper cites.
Training verifiers to solve math word problems
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., Hesse, C., and Schulman, J. (2021) · 2021
Earlier work this paper cites.
How can we know when language models know? on the calibration of language models for question answering
Jiang, Z., Araki, J., Ding, H., and Neubig, G. (2021) · 2021
Earlier work this paper cites.
The art of abstention: Selective prediction and error regularization for natural language processing
Xin, J., Tang, R., Yu, Y., and Lin, J. (2021) · 2021
Earlier work this paper cites.
Flashattention: Fast and memory-efficient exact attention with io-awareness
Dao, T., Fu, D. Y., Ermon, S., Rudra, A., and Ré, C. (2022) · 2022
Earlier work this paper cites.
Language models (mostly) know what they know
Kadavath, S., Conerly, T., Askell, A., Henighan, T., Drain, D., Perez, E., Schiefer, N., Hatfield-Dodds, Z., DasSarma, N., Tran-Johnson, E., Johnston, S., El-Showk, S., Jones, A., Elhage, N., Hume, T., Chen, A., Bai, Y., Bowman, S., Fort, S., Ganguli, D., Hernandez, D., Jacobson, J., Kernion, J., Kravec, S., Lovitt, L., Ndousse, K., Olsson, C., Ringer, S., Amodei, D., Brown, T., Clark, J., Joseph, N., Mann, B., McCandlish, S., Olah, C., and Kaplan, J. (2022) · 2022
Earlier work this paper cites.
Teaching models to express their uncertainty in words
Lin, S., Hilton, J., and Evans, O. (2022) · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P., Leike, J., and Lowe, R. (2022) · 2022
Cited alongside, same era.
The internal state of an llm knows when it’s lying
Azaria, A. and Mitchell, T. (2023) · 2023
Cited alongside, same era.
GPTCache: An open-source semantic cache for LLM applications enabling faster answers and cost savings
Bang, F. (2023) · 2023
Cited alongside, same era.
Tryage: Real-time, intelligent routing of user prompts to large language models
Hari, S. N. and Thomson, M. (2023) · 2023
Cited alongside, same era.
Llm-blender: Ensembling large language models with pairwise ranking and generative fusion
Jiang, D., Ren, X., and Lin, B. Y. (2023) · 2023
Cited alongside, same era.
Inside: Llms’ internal states retain the power of hallucination detection
Chen, C., Liu, K., Chen, Z., Gu, Y., Wu, Y., Tao, M., Fu, Z., and Ye, J. (2024) · 2024
Closest in time.
Hybrid llm: Cost-efficient and quality-aware query routing
Ding, D., Mallick, A., Wang, C., Sim, R., Mukherjee, S., Ruhle, V., Lakshmanan, L. V. S., and Awadallah, A. H. (2024) · 2024
Closest in time.
Language model cascades: Token-level uncertainty and beyond
Gupta, N., Narasimhan, H., Jitkrittum, W., Rawat, A. S., Menon, A. K., and Kumar, S. (2024) · 2024
Closest in time.
Routerbench: A benchmark for multi-llm routing system
Hu, Q. J., Bieker, J., Li, X., Jiang, N., Keigwin, B., Ranganath, G., Keutzer, K., and Upadhyay, S. K. (2024) · 2024
Closest in time.
Semantic entropy probes: Robust and cheap hallucination detection in llms
Kossen, J., Han, J., Razzak, M., Schut, L., Malik, S., and Gal, Y. (2024) · 2024
Closest in time.
Orchestrallm: Efficient orchestration of language models for dialogue state tracking
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kag, A., Fedorov, I., Gangrade, A., Whatmough, P., and Saligrama, V. (2023) · 2023
Cited alongside, same era.
Semantic uncertainty: Linguistic invariances for uncertainty estimation in natural language generation
Kuhn, L., Gal, Y., and Farquhar, S. (2023) · 2023
Cited alongside, same era.
Fast inference from transformers via speculative decoding
Leviathan, Y., Kalman, M., and Matias, Y. (2023) · 2023
Cited alongside, same era.
Routing to the expert: Efficient reward-guided ensemble of large language models
Lu, K., Yuan, H., Lin, R., Lin, J., Yuan, Z., Zhou, C., and Zhou, J. (2023) · 2023
Cited alongside, same era.
Selfcheckgpt: Zero-resource black-box hallucination detection for generative large language models
Manakul, P., Liusie, A., and Gales, M. J. F. (2023) · 2023
Cited alongside, same era.
Out-of-distribution detection and selective generation for conditional language models
Ren, J., Luo, J., Zhao, Y., Krishna, K., Saleh, M., Lakshminarayanan, B., and Liu, P. J. (2023) · 2023
Cited alongside, same era.
Post-abstention: Towards reliably re-attempting the abstained instances in qa
Varshney, N. and Baral, C. (2023) · 2023
Cited alongside, same era.
Lee, C.-H., Cheng, H., and Ostendorf, M. (2024) · 2024
Closest in time.
Generating with confidence: Uncertainty quantification for black-box large language models
Lin, Z., Trivedi, S., and Sun, J. (2024) · 2024
Closest in time.
Kernel language entropy: Fine-grained uncertainty quantification for llms from semantic similarities
Nikitin, A., Kossen, J., Gal, Y., and Marttinen, P. (2024) · 2024
Closest in time.
Gpt-4 technical report
OpenAI, Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., Avila, R., Babuschkin, I., Balaji, S., Balcom, V., Baltescu, P., Bao, H., Bavarian, M., Belgum, J., Bello, I., Berdine, J., Bernadett-Shapiro, G., Berner, C., Bogdonoff, L., Boiko, O., Boyd, M., Brakman, A.-L., Brockman, G., Brooks, T., Brundage, M., Button, K., Cai, T., Campbell, R., Cann, A., Carey, B., Carlson, C., Carmichael, R., Chan, B., Chang, C., Chantzis, F., Chen, D., Chen, S., Chen, R., Chen, J., Chen, M., Chess, B., Cho, C., Chu, C., Chung, H. W., Cummings, D., Currier, J., Dai, Y., Decareaux, C., Degry, T., Deutsch, N., Deville, D., Dhar, A., Dohan, D., Dowling, S., Dunning, S., Ecoffet, A., Eleti, A., Eloundou, T., Farhi, D., Fedus, L., Felix, N., Fishman, S. P., Forte, J., Fulford, I., Gao, L., Georges, E., Gibson, C., Goel, V., Gogineni, T., Goh, G., Gontijo-Lopes, R., Gordon, J., Grafstein, M., Gray, S., Greene, R., Gross, J., Gu, S. S., Guo, Y., Hallacy, C., Han, J., Harris, J., He, Y., Heaton, M., Heidecke, J., Hesse, C., Hickey, A., Hickey, W., Hoeschele, P., Houghton, B., Hsu, K., Hu, S., Hu, X., Huizinga, J., Jain, S., Jain, S., Jang, J., Jiang, A., Jiang, R., Jin, H., Jin, D., Jomoto, S., Jonn, B., Jun, H., Kaftan, T., Łukasz Kaiser, Kamali, A., Kanitscheider, I., Keskar, N. S., Khan, T., Kilpatrick, L., Kim, J. W., Kim, C., Kim, Y., Kirchner, J. H., Kiros, J., Knight, M., Kokotajlo, D., Łukasz Kondraciuk, Kondrich, A., Konstantinidis, A., Kosic, K., Krueger, G., Kuo, V., Lampe, M., Lan, I., Lee, T., Leike, J., Leung, J., Levy, D., Li, C. M., Lim, R., Lin, M., Lin, S., Litwin, M., Lopez, T., Lowe, R., Lue, P., Makanju, A., Malfacini, K., Manning, S., Markov, T., Markovski, Y., Martin, B., Mayer, K., Mayne, A., McGrew, B., McKinney, S. M., McLeavey, C., McMillan, P., McNeil, J., Medina, D., Mehta, A., Menick, J., Metz, L., Mishchenko, A., Mishkin, P., Monaco, V., Morikawa, E., Mossing, D., Mu, T., Murati, M., Murk, O., Mély, D., Nair, A., Nakano, R., Nayak, R., Neelakantan, A., Ngo, R., Noh, H., Ouyang, L., O’Keefe, C., Pachocki, J., Paino, A., Palermo, J., Pantuliano, A., Parascandolo, G., Parish, J., Parparita, E., Passos, A., Pavlov, M., Peng, A., Perelman, A., de Avila Belbute Peres, F., Petrov, M., de Oliveira Pinto, H. P., Michael, Pokorny, Pokrass, M., Pong, V. H., Powell, T., Power, A., Power, B., Proehl, E., Puri, R., Radford, A., Rae, J., Ramesh, A., Raymond, C., Real, F., Rimbach, K., Ross, C., Rotsted, B., Roussez, H., Ryder, N., Saltarelli, M., Sanders, T., Santurkar, S., Sastry, G., Schmidt, H., Schnurr, D., Schulman, J., Selsam, D., Sheppard, K., Sherbakov, T., Shieh, J., Shoker, S., Shyam, P., Sidor, S., Sigler, E., Simens, M., Sitkin, J., Slama, K., Sohl, I., Sokolowsky, B., Song, Y., Staudacher, N., Such, F. P., Summers, N., Sutskever, I., Tang, J., Tezak, N., Thompson, M. B., Tillet, P., Tootoonchian, A., Tseng, E., Tuggle, P., Turley, N., Tworek, J., Uribe, J. F. C., Vallone, A., Vijayvergiya, A., Voss, C., Wainwright, C., Wang, J. J., Wang, A., Wang, B., Ward, J., Wei, J., Weinmann, C., Welihinda, A., Welinder, P., Weng, J., Weng, L., Wiethoff, M., Willner, D., Winter, C., Wolrich, S., Wong, H., Workman, L., Wu, S., Wu, J., Wu, M., Xiao, K., Xu, T., Yoo, S., Yu, K., Yuan, Q., Zaremba, W., Zellers, R., Zhang, C., Zhang, M., Zhao, S., Zheng, T., Zhuang, J., Zhuk, W., and Zoph, B. (2024) · 2024
Closest in time.
Softmax probabilities (mostly) predict large language model correctness on multiple-choice q&a
Plaut, B., Nguyen, K., and Trinh, T. (2024) · 2024
Closest in time.
Fly-swat or cannon? cost-effective language model choice via meta-modeling
Sakota, M., Peyrard, M., and West, R. (2024) · 2024
Closest in time.
Fusing models with complementary expertise
Wang, H., Polo, F. M., Sun, Y., Kundu, S., Xing, E., and Yurochkin, M. (2024) · 2024
Closest in time.
Can llms express their uncertainty? an empirical evaluation of confidence elicitation in llms
Xiong, M., Hu, Z., Lu, X., Li, Y., Fu, J., He, J., and Hooi, B. (2024) · 2024
Closest in time.
A survey on knowledge distillation of large language models
Xu, X., Li, M., Tao, C., Shen, T., Cheng, R., Li, J., Xu, C., Tao, D., and Zhou, T. (2024) · 2024
Closest in time.
Large language model cascades with mixture of thoughts representations for cost-efficient reasoning
Yue, M., Zhao, J., Zhang, M., Du, L., and Yao, Z. (2024) · 2024
Closest in time.
Selective-LAMA: Selective prediction for confidence-aware evaluation of language models
Yoshikawa, H. and Okazaki, N. (2023) · 2028
Closest in time.