Fetching the paper…
Reading the bibliography…
One of the key technologies for the success of Large Language Models (LLMs) is preference alignment.
On the convergence properties of the em algorithm
Wu, C. J · 1983
Earlier work this paper cites.
Transforming classifier scores into accurate multiclass probability estimates
Zadrozny, B. and Elkan, C · 2002
Earlier work this paper cites.
Rethinking the inception architecture for computer vision
Szegedy, C., Vanhoucke, V., Ioffe, S., Shlens, J., and Wojna, Z · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D · 2017
Earlier work this paper cites.
On calibration of modern neural networks
Guo, C., Pleiss, G., Sun, Y., and Weinberger, K. Q · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Earlier work this paper cites.
mixup: Beyond empirical risk minimization
Zhang, H., Cisse, M., Dauphin, Y. N., and Lopez-Paz, D · 2017
Earlier work this paper cites.
Trainable calibration measures for neural networks from kernel mean embeddings
Kumar, A., Sarawagi, S., and Jain, U · 2018
Earlier work this paper cites.
Beyond temperature scaling: Obtaining well-calibrated multi-class probabilities with dirichlet calibration
Kull, M., Perello Nieto, M., Kängsepp, M., Silva Filho, T., Song, H., and Flach, P · 2019
Earlier work this paper cites.
Verified uncertainty calibration
Kumar, A., Liang, P. S., and Ma, T · 2019
Earlier work this paper cites.
When does label smoothing help?
Müller, R., Kornblith, S., and Hinton, G. E · 2019
Earlier work this paper cites.
Distribution-free binary classification: prediction sets, confidence intervals and calibration
Gupta, C., Podkopaev, A., and Ramdas, A · 2020
Earlier work this paper cites.
Improving model calibration with accuracy versus uncertainty optimization
Krishnan, R. and Tickoo, O · 2020
Earlier work this paper cites.
Calibrating deep neural networks using focal loss
Mukhoti, J., Kulharia, V., Sanyal, A., Golodetz, S., Torr, P., and Dokania, P · 2020
Earlier work this paper cites.
Distribution-free calibration guarantees for histogram binning without sample splitting
Gupta, C. and Ramdas, A · 2021
Earlier work this paper cites.
How can we know when language models know? on the calibration of language models for question answering
Jiang, Z., Araki, J., Ding, H., and Neubig, G · 2021
Earlier work this paper cites.
Soft calibration objectives for neural networks
Karandikar, A., Cain, N., Tran, D., Lakshminarayanan, B., Shlens, J., Mozer, M. C., and Roelofs, B · 2021
Earlier work this paper cites.
Knowing more about questions can help: Improving calibration in question answering
Zhang, S., Gong, C., and Choi, E · 2021
Earlier work this paper cites.
Calibrate before use: Improving few-shot performance of language models
Zhao, Z., Wallace, E., Feng, S., Klein, D., and Singh, S · 2021
Earlier work this paper cites.
A close look into the calibration of pre-trained language models
Chen, Y., Yuan, L., Cui, G., Liu, Z., and Ji, H · 2022
Earlier work this paper cites.
Better uncertainty calibration via proper scores for classification and beyond
Gruber, S. and Buettner, F · 2022
Earlier work this paper cites.
Language models (mostly) know what they know
Kadavath, S., Conerly, T., Askell, A., Henighan, T., Drain, D., Perez, E., Schiefer, N., Hatfield-Dodds, Z., DasSarma, N., Tran-Johnson, E., et al · 2022
Earlier work this paper cites.
Teaching models to express their uncertainty in words
Lin, S., Hilton, J., and Evans, O · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback, 2022
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P., Leike, J., and Lowe, R · 2022
Cited alongside, same era.
A consistent and differentiable lp canonical calibration error estimator
Popordanoska, T., Sayer, R., and Blaschko, M · 2022
Cited alongside, same era.
Uncertainty quantification with pre-trained language models: A large-scale empirical analysis
Xiao, Y., Liang, P. P., Bhatt, U., Neiswanger, W., Salakhutdinov, R., and Morency, L.-P · 2022
Cited alongside, same era.
Robust calibration with multi-domain temperature scaling
Yu, Y., Bates, S., Ma, Y., and Jordan, M · 2022
Cited alongside, same era.
Sparks of artificial general intelligence: Early experiments with gpt-4
Bubeck, S., Chandrasekaran, V., Eldan, R., Gehrke, J., Horvitz, E., Kamar, E., Lee, P., Lee, Y. T., Li, Y., Lundberg, S., et al · 2023
Maxmin-rlhf: Towards equitable alignment of large language models with diverse human preferences
Chakraborty, S., Qiu, J., Yuan, H., Koppel, A., Huang, F., Manocha, D., Bedi, A. S., and Wang, M · 2024
Later among the works it cites.
Dataset reset policy optimization for rlhf
Chang, J. D., Shan, W., Oertell, O., Brantley, K., Misra, D., Lee, J. D., and Sun, W · 2024
Later among the works it cites.
Qlora: Efficient finetuning of quantized llms
Dettmers, T., Pagnoni, A., Holtzman, A., and Zettlemoyer, L · 2024
Later among the works it cites.
Pac-bayes analysis for recalibration in classification
Fujisawa, M. and Futami, F · 2024
Later among the works it cites.
Information-theoretic generalization analysis for expected calibration error
Futami, F. and Fujisawa, M · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Chiang, W.-L., Li, Z., Lin, Z., Sheng, Y., Wu, Z., Zhang, H., Zheng, L., Zhuang, S., Zhuang, Y., Gonzalez, J. E., et al · 2023
Cited alongside, same era.
Palm: Scaling language modeling with pathways
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al · 2023
Cited alongside, same era.
Prototypical calibration for few-shot learning of language models
Han, Z., Hao, Y., Dong, L., Sun, Y., and Wei, F · 2023
Cited alongside, same era.
Generative calibration for in-context learning
Jiang, Z., Zhang, Y., Liu, C., Zhao, J., and Liu, K · 2023
Cited alongside, same era.
Sample-dependent adaptive temperature scaling for improved calibration
Joy, T., Pinto, F., Lim, S.-N., Torr, P. H., and Dokania, P. K · 2023
Cited alongside, same era.
Statistical rejection sampling improves preference optimization
Liu, T., Zhao, Y., Joshi, R., Khalman, M., Saleh, M., Liu, P. J., and Liu, J · 2023
Cited alongside, same era.
Nash learning from human feedback
Munos, R., Valko, M., Calandriello, D., Azar, M. G., Rowland, M., Guo, Z. D., Tang, Y., Geist, M., Mesnard, T., Michi, A., et al · 2023
Cited alongside, same era.
Later among the works it cites.
Learn your reference model for real good alignment
Gorbatovski, A., Shaposhnikov, B., Malakhov, A., Surnachev, N., Aksenov, Y., Maksimov, I., Balagansky, N., and Gavrilov, D · 2024
Later among the works it cites.
Enhancing confidence expression in large language models through learning from past experience
Han, H., Li, T., Chen, S., Shi, J., Du, C., Xiao, Y., Liang, J., and Lin, X · 2024
Later among the works it cites.
T \ \backslash " ulu 3: Pushing frontiers in open language model post-training
Lambert, N., Morrison, J., Pyatkin, V., Huang, S., Ivison, H., Brahman, F., Miranda, L. J. V., Liu, A., Dziri, N., Lyu, S., et al · 2024
Later among the works it cites.
Taming overconfidence in llms: Reward calibration in rlhf
Leng, J., Huang, C., Zhu, B., and Huang, J · 2024
Later among the works it cites.
OLMo, T., Walsh, P., Soldaini, L., Groeneveld, D., Lo, K., Arora, S., Bhagia, A., Gu, Y., Huang, S., Jordan, M., et al · 2024
Later among the works it cites.
Thermometer: Towards universal calibration for large language models
Shen, M., Das, S., Greenewald, K., Sattigeri, P., Wornell, G., and Ghosh, S · 2024
Later among the works it cites.
Preference fine-tuning of llms should leverage suboptimal, on-policy data
Tajwar, F., Singh, A., Sharma, A., Rafailov, R., Schneider, J., Xie, T., Ermon, S., Finn, C., and Kumar, A · 2024
Later among the works it cites.
Understanding the performance gap between online and offline alignment algorithms
Tang, Y., Guo, D. Z., Zheng, Z., Calandriello, D., Cao, Y., Tarassov, E., Munos, R., Pires, B. Á., Valko, M., Cheng, Y., et al · 2024
Later among the works it cites.
When to trust llms: Aligning confidence with response quality
Tao, S., Yao, L., Ding, H., Xie, Y., Cao, Q., Sun, F., Gao, J., Shen, H., and Ding, B · 2024
Later among the works it cites.
Sayself: Teaching llms to express confidence with self-reflective rationales
Xu, T., Wu, S., Diao, S., Liu, X., Wang, X., Chen, Y., and Gao, J · 2024
Later among the works it cites.
Asymptotics of language model alignment
Yang, J. Q., Salamatian, S., Sun, Z., Suresh, A. T., and Beirami, A · 2024
Later among the works it cites.
A theoretical analysis of nash learning from human feedback under general kl-regularized preference
Ye, C., Xiong, W., Zhang, Y., Jiang, N., and Zhang, T · 2024
Later among the works it cites.
Fine-tuning attention modules only: Enhancing weight disentanglement in task arithmetic
Jin, R., Hou, B., Xiao, J., Su, W. J., and Shen, L · 2025
Closest in time.
Preserving diversity in supervised fine-tuning of large language models
Li, Z., Chen, C., Xu, T., Qin, Z., Xiao, J., Luo, Z.-Q., and Sun, R · 2025
Closest in time.
Liu, K., Long, Q., Shi, Z., Su, W. J., and Xiao, J · 2025
Closest in time.
Large language model uncertainty proxies: discrimination and calibration for medical diagnosis and treatment
Savage, T., Wang, J., Gallo, R., Boukil, A., Patel, V., Safavi-Naini, S. A. A., Soroush, A., and Chen, J. H · 2025
Closest in time.
Fundamental limits of game-theoretic llm alignment: Smith consistency and preference matching
Shi, Z., Liu, K., Long, Q., Su, W. J., and Xiao, J · 2025
Closest in time.
Magnetic preference optimization: Achieving last-iterate convergence for language model alignment
Wang, M., Ma, C., Chen, Q., Meng, L., Han, Y., Xiao, J., Zhang, Z., Huo, J., Su, W. J., and Yang, Y · 2025
Closest in time.