Fetching the paper…
Reading the bibliography…
Current LLM alignment techniques use pairwise human preferences at a sample level, and as such, they do not imply an alignment on the distributional level.
Computational optimal transport
G. Peyré and M. Cuturi · 1935
Earlier work this paper cites.
Dual stochastic dominance and related mean-risk models
W. Ogryczak and A. Ruszczynski · 2002
Earlier work this paper cites.
Are loss functions all the same?
L. Rosasco, E. De Vito, A. Caponnetto, M. Piana, and A. Verri · 2004
Earlier work this paper cites.
Convexity, classification, and risk bounds
P. L. Bartlett, M. I. Jordan, and J. D. McAuliffe · 2006
Earlier work this paper cites.
Optimal Transport: old and new , volume 338
C. Villani · 2009
Earlier work this paper cites.
Optimal Transport for Applied Mathematicians: Calculus of Variations, PDEs, and Modeling
F. Santambrogio · 2015
Earlier work this paper cites.
Deep reinforcement learning from human preferences
P. F. Christiano, J. Leike, T. Brown, M. Martic, S. Legg, and D. Amodei · 2017
Earlier work this paper cites.
An optimal transportation approach for assessing almost stochastic order
E. Del Barrio, J. A. Cuesta-Albertos, and C. Matrán · 2018
Earlier work this paper cites.
Sorting out lipschitz function approximation
C. Anil, J. Lucas, and R. Grosse · 2019
Earlier work this paper cites.
Differentiable ranking and sorting using optimal transport
M. Cuturi, O. Teboul, and J.-P. Vert · 2019
Earlier work this paper cites.
Fast differentiable sorting and ranking
M. Blondel, O. Teboul, Q. Berthet, and J. Djolonga · 2020
Earlier work this paper cites.
Learning to summarize with human feedback
N. Stiennon, L. Ouyang, J. Wu, D. Ziegler, R. Lowe, C. Voss, A. Radford, D. Amodei, and P. F. Christiano · 2020
Earlier work this paper cites.
Trl: Transformer reinforcement learning
L. von Werra, Y. Belkada, L. Tunstall, E. Beeching, T. Thrush, N. Lambert, and S. Huang · 2020
Cited alongside, same era.
LoRa: Low-rank adaptation of large language models
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen · 2021
Cited alongside, same era.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Y. Bai, A. Jones, K. Ndousse, A. Askell, A. Chen, N. DasSarma, D. Drain, S. Fort, D. Ganguli, T. Henighan, et al · 2022
Cited alongside, same era.
Learning with stochastic orders
C. Domingo-Enrich, Y. Schiff, and Y. Mroueh · 2022
Cited alongside, same era.
Empirical optimal transport between different measures adapts to lower complexity
S. Hundrieser, T. Staudt, and A. Munk · 2022
Cited alongside, same era.
Calibrating sequence likelihood improves conditional language generation
Y. Zhao, M. Khalman, R. Joshi, S. Narayan, M. Saleh, and P. J. Liu · 2023
Later among the works it cites.
Starling-7b: Improving llm helpfulness and harmlessness with rlaif, November 2023
B. Zhu, E. Frick, T. Wu, H. Zhu, and J. Jiao · 2023
Later among the works it cites.
Llama 3 model card
AI@Meta · 2024
Closest in time.
A general theoretical paradigm to understand learning from human preferences
M. G. Azar, Z. D. Guo, B. Piot, R. Munos, M. Rowland, M. Valko, and D. Calandriello · 2024
Closest in time.
Length-controlled alpacaeval: A simple way to debias automatic evaluators
Y. Dubois, B. Galambosi, P. Liang, and T. B. Hashimoto · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al · 2022
Cited alongside, same era.
Open LLM Leaderboard
E. Beeching, C. Fourrier, N. Habib, S. Han, N. Lambert, N. Rajani, O. Sanseviero, L. Tunstall, and T. Wolf · 2023
Cited alongside, same era.
Human-centered loss functions (halos)
K. Ethayarajh, W. Xu, D. Jurafsky, and D. Kiela · 2023
Cited alongside, same era.
Lower complexity adaptation for empirical entropic optimal transport
M. Groppe and S. Hundrieser · 2023
Cited alongside, same era.
Beavertails: Towards improved safety alignment of llm via a human-preference dataset
J. Ji, M. Liu, J. Dai, X. Pan, C. Zhang, C. Bian, C. Zhang, R. Sun, Y. Wang, and Y. Yang · 2023
Cited alongside, same era.
A. Jiang, A. Sablayrolles, A. Mensch, C. Bamford, D. Chaplot, D. de las Casas, F. Bressand, G. Lengyel, G. Lample, L. Saulnier, et al · 2023
Cited alongside, same era.
Helpsteer: Multi-attribute helpfulness dataset for steerlm
Z. Wang, Y. Dong, J. Zeng, V. Adams, M. N. Sreedhar, D. Egert, O. Delalleau, J. P. Scowcroft, N. Kant, A. Swope, et al · 2023
Cited alongside, same era.
K. Ethayarajh, W. Xu, N. Muennighoff, D. Jurafsky, and D. Kiela · 2024
Closest in time.
tinyBenchmarks: evaluating LLMs with fewer examples
F. Maia Polo, L. Weber, L. Choshen, Y. Sun, G. Xu, and M. Yurochkin · 2024
Closest in time.
Sharp convergence rates for empirical optimal transport with smooth costs
T. Manole and J. Niles-Weed · 2024
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
R. Rafailov, A. Sharma, E. Mitchell, C. D. Manning, S. Ermon, and C. Finn · 2024
Closest in time.
Lab: Large-scale alignment for chatbots
S. Sudalairaj, A. Bhandwaldar, A. Pareja, K. Xu, D. D. Cox, and A. Srivastava · 2024
Closest in time.
Openhermes-2.5-mistral-7b
Teknium · 2024
Closest in time.
Self-rewarding language models, 2024
W. Yuan, R. Y. Pang, K. Cho, X. Li, S. Sukhbaatar, J. Xu, and J. Weston · 2024
Closest in time.