Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have demonstrated remarkable performances in various tasks.
MM Algorithms for Generalized Bradley-Terry Models
D. R. Hunter · 2004
Earlier work this paper cites.
The k-armed dueling bandits problem
Y. Yue, J. Broder, R. Kleinberg, and T. Joachims · 2012
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
A. Jacot, F. Gabriel, and C. Hongler · 2018
Earlier work this paper cites.
Language models are unsupervised multitask learners
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever, et al · 2019
Earlier work this paper cites.
Fine-tuning language models from human preferences
D. M. Ziegler, N. Stiennon, J. Wu, T. B. Brown, A. Radford, D. Amodei, P. Christiano, and G. Irving · 2019
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al · 2020
Earlier work this paper cites.
AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts
T. Shin, Y. Razeghi, R. L. Logan IV, E. Wallace, and S. Singh · 2020
Earlier work this paper cites.
Denoising diffusion implicit models
J. Song, C. Meng, and S. Ermon · 2020
Earlier work this paper cites.
Mpnet: Masked and permuted pre-training for language understanding
K. Song, X. Tan, T. Qin, J. Lu, and T.-Y. Liu · 2020
Earlier work this paper cites.
Neural contextual bandits with UCB-based exploration
D. Zhou, L. Li, and Q. Gu · 2020
Earlier work this paper cites.
The Power of Scale for Parameter-Efficient Prompt Tuning
B. Lester, R. Al-Rfou, and N. Constant · 2021
Earlier work this paper cites.
Prefix-Tuning: Optimizing Continuous Prompts for Generation
X. L. Li and P. Liang · 2021
Earlier work this paper cites.
Value-at-risk optimization with Gaussian processes
Q. P. Nguyen, Z. Dai, B. K. H. Low, and P. Jaillet · 2021
Earlier work this paper cites.
Learning transferable visual models from natural language supervision
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al · 2021
Earlier work this paper cites.
Optimal algorithms for stochastic contextual preference bandits
A. Saha · 2021
Earlier work this paper cites.
Neural Thompson sampling
W. Zhang, D. Zhou, L. Li, and Q. Gu · 2021
Earlier work this paper cites.
Factual probing is [MASK]: Learning vs. learning to recall
Z. Zhong, D. Friedman, and D. Chen · 2021
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Y. Bai, A. Jones, K. Ndousse, A. Askell, A. Chen, N. DasSarma, D. Drain, S. Fort, D. Ganguli, T. Henighan, et al · 2022
Earlier work this paper cites.
Stochastic Contextual Dueling Bandits under Linear Stochastic Transitivity Models
V. Bengs, A. Saha, and E. Hüllermeier · 2022
Earlier work this paper cites.
Clip-tuning: Towards derivative-free prompt learning with a mixture of rewards
Y. Chai, S. Wang, Y. Sun, H. Tian, H. Wu, and H. Wang · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al · 2022
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer · 2022
Cited alongside, same era.
BBTv2: Pure black-box optimization can be comparable to gradient descent for few-shot learning
T. Sun, Z. He, H. Qian, X. Huang, and X. Qiu · 2022
Cited alongside, same era.
Black-box tuning for language-model-as-a-service
T. Sun, Y. Shao, H. Qian, X. Huang, and X. Qiu · 2022
Cited alongside, same era.
Improving image generation with better captions
Direct preference optimization with an offset
A. Amini, T. Vieira, and R. Cotterell · 2024
Closest in time.
A general theoretical paradigm to understand learning from human preferences
M. G. Azar, Z. D. Guo, B. Piot, R. Munos, M. Rowland, M. Valko, and D. Calandriello · 2024
Closest in time.
RLHF Deciphered: A Critical Analysis of Reinforcement Learning from Human Feedback for LLMs
S. Chaudhari, P. Aggarwal, V. Murahari, T. Rajpurohit, A. Kalyan, K. Narasimhan, A. Deshpande, and B. C. da Silva · 2024
Closest in time.
Evaluating text-to-image generative models: An empirical study on human image synthesis
M. Chen, Y. Liu, J. Yi, C. Xu, Q. Lai, H. Wang, T.-Y. Ho, and Q. Xu · 2024
Closest in time.
Online personalizing white-box llms generation with neural bandits
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Betker, G. Goh, L. Jing, T. Brooks, J. Wang, L. Li, L. Ouyang, J. Zhuang, J. Lee, Y. Guo, et al · 2023
Cited alongside, same era.
Open problems and fundamental limitations of reinforcement learning from human feedback
S. Casper, X. Davies, C. Shi, T. K. Gilbert, J. Scheurer, J. Rando, R. Freedman, T. Korbak, D. Lindner, P. Freire, et al · 2023
Cited alongside, same era.
InstructZero: Efficient Instruction Optimization for Black-Box Large Language Models
L. Chen, J. Chen, T. Goldstein, H. Huang, and T. Zhou · 2023
Cited alongside, same era.
Black-box prompt learning for pre-trained language models
S. Diao, Z. Huang, R. Xu, X. Li, L. Yong, X. Zhou, and T. Zhang · 2023
Cited alongside, same era.
Promptbreeder: Self-referential self-improvement via prompt evolution
C. Fernando, D. Banarse, H. Michalewski, S. Osindero, and T. Rocktäschel · 2023
Cited alongside, same era.
Google · 2023
Cited alongside, same era.
OpenAI · 2023
Cited alongside, same era.
Z. Chen, W. Daniel, P.-y. Chen, and F. Buet-Golfouse · 2024
Closest in time.
Alpacafarm: A simulation framework for methods that learn from human feedback
Y. Dubois, C. X. Li, R. Taori, T. Zhang, I. Gulrajani, J. Ba, C. Guestrin, P. S. Liang, and T. B. Hashimoto · 2024
Closest in time.
Efficient exploration for llms
V. Dwaracherla, S. M. Asghari, B. Hao, and B. Van Roy · 2024
Closest in time.
Mixed preference optimization: Reinforcement learning with data selection and better reference model
Q. Gou and C.-T. Nguyen · 2024
Closest in time.
Connecting large language models with evolutionary algorithms yields powerful prompt optimizers
Q. Guo, R. Wang, J. Guo, B. Li, K. Song, X. Tan, G. Liu, J. Bian, and Y. Yang · 2024
Closest in time.
Localized zeroth-order prompt optimization
W. Hu, Y. Shu, Z. Yu, Z. Wu, X. Lin, Z. Dai, S.-K. Ng, and B. K. H. Low · 2024
Closest in time.
PRewrite: Prompt Rewriting with Reinforcement Learning
W. Kong, S. A. Hombaiah, M. Zhang, Q. Mei, and M. Bendersky · 2024
Closest in time.
Use Your INSTINCT: INSTruction optimization usIng Neural bandits Coupled with Transformers
X. Lin, Z. Wu, Z. Dai, W. Hu, Y. Shu, S.-K. Ng, P. Jaillet, and B. K. H. Low · 2024
Closest in time.
LiPO: Listwise Preference Optimization through Learning-to-Rank
T. Liu, Z. Qin, J. Wu, J. Shen, M. Khalman, R. Joshi, Y. Zhao, M. Saleh, S. Baumgartner, J. Liu, et al · 2024
Closest in time.
Improving text-to-image consistency via automatic prompt optimization
O. Mañas, P. Astolfi, M. Hall, C. Ross, J. Urbanek, A. Williams, A. Agrawal, A. Romero-Soriano, and M. Drozdzal · 2024
Closest in time.
Filtered direct preference optimization
T. Morimura, M. Sakamoto, Y. Jinnai, K. Abe, and K. Air · 2024
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
R. Rafailov, A. Sharma, E. Mitchell, C. D. Manning, S. Ermon, and C. Finn · 2024
Closest in time.
Best arm identification for prompt learning under a limited budget
C. Shi, K. Yang, J. Yang, and C. Shen · 2024
Closest in time.
Generalized preference optimization: A unified approach to offline alignment
Y. Tang, Z. D. Guo, Z. Zheng, D. Calandriello, R. Munos, M. Rowland, P. H. Richemond, M. Valko, B. Á. Pires, and B. Piot · 2024
Closest in time.
Large language models as optimizers
C. Yang, X. Wang, Y. Lu, H. Liu, Q. V. Le, D. Zhou, and X. Chen · 2024
Closest in time.