Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) exhibit remarkably powerful capabilities.
“The Curious Case of Neural Text Degeneration”, 2020
Ari Holtzman et al · 1904
Earlier work this paper cites.
“Rank analysis of incomplete block designs: I. The method of paired comparisons”
Ralph Bradley and Milton Terry · 1952
Earlier work this paper cites.
“Non-null ranking models. I”
Colin Mallows · 1957
Earlier work this paper cites.
“Simple statistical gradient-following algorithms for connectionist reinforcement learning”
Ronald Williams · 1992
Earlier work this paper cites.
“ROUGE: A Package for Automatic Evaluation of Summaries”
Chin-Yew Lin · 2004
Earlier work this paper cites.
“Bandit Based Monte-Carlo Planning”
Levente Kocsis and Csaba Szepesvári · 2006
Earlier work this paper cites.
“Learning to summarize from human feedback”, 2022
Nisan Stiennon et al · 2009
Earlier work this paper cites.
“Sequence Transduction with Recurrent Neural Networks”, 2012
Alex Graves · 2012
Earlier work this paper cites.
“SQuAD: 100,000+ Questions for Machine Comprehension of Text”
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev and Percy Liang · 2016
Earlier work this paper cites.
“Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension”
Mandar Joshi, Eunsol Choi, Daniel Weld and Luke Zettlemoyer · 2017
Earlier work this paper cites.
“Proximal policy optimization algorithms”
John Schulman et al · 2017
Earlier work this paper cites.
“TL;DR: Mining Reddit to Learn Automatic Summarization”
Michael Völske, Martin Potthast, Shahbaz Syed and Benno Stein · 2017
Earlier work this paper cites.
“Complex sequential question answering: Towards learning to converse over linked question answer pairs with a knowledge graph”
Amrita Saha et al · 2018
Earlier work this paper cites.
“Natural questions: a benchmark for question answering research”
Tom Kwiatkowski et al · 2019
Earlier work this paper cites.
“Measuring massive multitask language understanding”
Dan Hendrycks et al · 2020
Earlier work this paper cites.
“Program synthesis with large language models”
Jacob Austin et al · 2021
Earlier work this paper cites.
“Evaluating large language models trained on code”
Mark Chen et al · 2021
Earlier work this paper cites.
“Training verifiers to solve math word problems”
Karl Cobbe et al · 2021
Earlier work this paper cites.
“Did aristotle use a laptop? a question answering benchmark with implicit reasoning strategies”
Mor Geva et al · 2021
Earlier work this paper cites.
“Alignment of Language Agents”, 2021
Zachary Kenton et al · 2021
Earlier work this paper cites.
“The Lean 4 Theorem Prover and Programming Language”
Leonardo Moura and Sebastian Ullrich · 2021
Earlier work this paper cites.
“WebGPT: Browser-assisted question-answering with human feedback”
Reiichiro Nakano et al · 2021
Earlier work this paper cites.
“FUDGE: Controlled text generation with future discriminators”
Kevin Yang and Dan Klein · 2021
Earlier work this paper cites.
“Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback”, 2022
Yuntao Bai et al · 2022
Earlier work this paper cites.
“Understanding Dataset Difficulty with 𝒱 \mathcal{V} -Usable Information”
Kawin Ethayarajh, Yejin Choi and Swabha Swayamdipta · 2022
Earlier work this paper cites.
“Training language models to follow instructions with human feedback”, 2022
Long Ouyang et al · 2022
Earlier work this paper cites.
“Challenging big-bench tasks and whether chain-of-thought can solve them”
Mirac Suzgun et al · 2022
Earlier work this paper cites.
“Solving math word problems with process- and outcome-based feedback”, 2022
Jonathan Uesato et al · 2022
Earlier work this paper cites.
“Star: Self-taught reasoner bootstrapping reasoning with reasoning”
Eric Zelikman, Yuhuai Wu, Jesse Mu and Noah Goodman · 2022
Earlier work this paper cites.
“Calibrating sequence likelihood improves conditional language generation”
Yao Zhao et al · 2022
Earlier work this paper cites.
Josh Achiam et al · 2023
Earlier work this paper cites.
“A general theoretical paradigm to understand learning from human preferences”
Mohammad Azar et al · 2023
Earlier work this paper cites.
“Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling”, 2023
Stella Biderman et al · 2023
Earlier work this paper cites.
“Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision”, 2023
Collin Burns et al · 2023
Earlier work this paper cites.
“Black-box prompt optimization: Aligning large language models without model training”
Jiale Cheng et al · 2023
Earlier work this paper cites.
“Can large language models be an alternative to human evaluations?”
Cheng-Han Chiang and Hung-yi Lee · 2023
Earlier work this paper cites.
“Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90%* ChatGPT Quality”, 2023
Wei-Lin Chiang et al · 2023
Earlier work this paper cites.
“UltraFeedback: Boosting Language Models with High-quality Feedback”, 2023
Ganqu Cui et al · 2023
Earlier work this paper cites.
“Reward-augmented decoding: Efficient controlled text generation with a unidirectional reward model”
Haikang Deng and Colin Raffel · 2023
Earlier work this paper cites.
“Enhancing Chat Language Models by Scaling High-quality Instructional Conversations”, 2023
Ning Ding et al · 2023
Earlier work this paper cites.
“Raft: Reward ranked finetuning for generative foundation model alignment”
Hanze Dong et al · 2023
Earlier work this paper cites.
“AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback”, 2023
Yann Dubois et al · 2023
Earlier work this paper cites.
“The false promise of imitating proprietary llms”
Arnav Gudibande et al · 2023
Earlier work this paper cites.
“Contrastive prefence learning: Learning from human feedback without rl”
Joey Hejna et al · 2023
Earlier work this paper cites.
Jixiang Hong et al · 2023
Earlier work this paper cites.
“Multi-dimensional evaluation of text summarization with in-context learning”
Sameer Jain et al · 2023
Earlier work this paper cites.
“LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative Fusion”, 2023
Dongfu Jiang, Xiang Ren and Bill Lin · 2023
Earlier work this paper cites.
“Swe-bench: Can language models resolve real-world github issues?”
Carlos Jimenez et al · 2023
Earlier work this paper cites.
“Prometheus: Inducing fine-grained evaluation capability in language models”
Seungone Kim et al · 2023
Earlier work this paper cites.
“Large language models are state-of-the-art evaluators of translation quality”
Tom Kocmi and Christian Federmann · 2023
Earlier work this paper cites.
“Certifying llm safety against adversarial prompting”
Aounon Kumar et al · 2023
Earlier work this paper cites.
“RLAIF: Scaling Reinforcement Learning from Human Feedback with AI Feedback”, 2023
Harrison Lee et al · 2023
Earlier work this paper cites.
“Cmmlu: Measuring massive multitask language understanding in chinese”
Haonan Li et al · 2023
Earlier work this paper cites.
“Generative judge for evaluating alignment”
Junlong Li et al · 2023
Earlier work this paper cites.
“Prd: Peer rank and discussion improve large language model based evaluations”
Ruosen Li, Teerth Patel and Xinya Du · 2023
Earlier work this paper cites.
“AlpacaEval: An Automatic Evaluator of Instruction-following Models”
Xuechen Li et al · 2023
Earlier work this paper cites.
“Rain: Your language models can align themselves without finetuning”
Yuhui Li et al · 2023
Earlier work this paper cites.
“Remax: A simple, effective, and efficient reinforcement learning method for aligning large language models”
Ziniu Li et al · 2023
Earlier work this paper cites.
“Let’s Verify Step by Step”, 2023
Hunter Lightman et al · 2023
Earlier work this paper cites.
“The unlocking spell on base llms: Rethinking alignment via in-context learning”
Bill Lin et al · 2023
Cited alongside, same era.
“RLTF: Reinforcement Learning from Unit Test Feedback”, 2023
Jiate Liu et al · 2023
Cited alongside, same era.
“Statistical rejection sampling improves preference optimization”
Tianqi Liu et al · 2023
Cited alongside, same era.
“G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment (2023)”
Yang Liu et al · 2023
Cited alongside, same era.
“Ml-bench: Large language models leverage open-source libraries for machine learning tasks”
“Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs”
Chris Liu et al · 2024
Closest in time.
“Chain of Hindsight aligns Language Models with Feedback”
Hao Liu, Carmelo Sferrazza and Pieter Abbeel · 2024
Closest in time.
“Decoding-time Realignment of Language Models”
Tianlin Liu et al · 2024
Closest in time.
“LiPO: Listwise Preference Optimization through Learning-to-Rank”
Tianqi Liu et al · 2024
Closest in time.
“Improve Mathematical Reasoning in Language Models by Automated Process Supervision”, 2024
Liangchen Luo et al · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yuliang Liu et al · 2023
Cited alongside, same era.
“Controlled decoding from language models”
Sidharth Mudgal et al · 2023
Cited alongside, same era.
“Practices for governing agentic AI systems”
Yonadav Shavit et al · 2023
Cited alongside, same era.
“PanGu-Coder2: Boosting Large Language Models for Code with Ranking Feedback”, 2023
Bo Shen et al · 2023
Cited alongside, same era.
“Large Language Model Alignment: A Survey”, 2023
Tianhao Shen et al · 2023
Cited alongside, same era.
“Llama 2: Open foundation and fine-tuned chat models”
Hugo Touvron et al · 2023
Cited alongside, same era.
“Llama: Open and efficient foundation language models”
Hugo Touvron et al · 2023
Cited alongside, same era.
“Beyond reverse kl: Generalizing direct preference optimization with diverse divergence constraints”
Chaoqi Wang et al · 2023
Cited alongside, same era.
“KnowTuning: Knowledge-aware Fine-tuning for Large Language Models”
Yougang Lyu et al · 2024
Closest in time.
“MACPO: Weak-to-Strong Alignment via Multi-Agent Contrastive Preference Optimization”
Yougang Lyu et al · 2024
Closest in time.
“Don’t Forget Your Reward Values: Language Model Alignment via Value-based Calibration”
Xin Mao et al · 2024
Closest in time.
“LLM Critics Help Catch LLM Bugs”, 2024
Nat McAleese et al · 2024
Closest in time.
“SimPO: Simple Preference Optimization with a Reference-Free Reward”
Yu Meng, Mengzhou Xia and Danqi Chen · 2024
Closest in time.
“Aligning CodeLLMs with Direct Preference Optimization”
Yibo Miao et al · 2024
Closest in time.
“Filtered Direct Preference Optimization”, 2024
Tetsuro Morimura et al · 2024
Closest in time.
“GPT-4 Technical Report”, 2024
OpenAI et al · 2024
Closest in time.
“West-of-N: Synthetic Preference Generation for Improved Reward Modeling”, 2024
Alizée Pace et al · 2024
Closest in time.
“Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive”
Arka Pal et al · 2024
Closest in time.
“Mutual Reasoning Makes Smaller LLMs Stronger Problem-Solvers”, 2024
Zhenting Qi et al · 2024
Closest in time.
“DMoERM: Recipes of Mixture-of-Experts for Effective Reward Modeling”
Shanghaoran Quan · 2024
Closest in time.
“Direct preference optimization: Your language model is secretly a reward model”
Rafael Rafailov et al · 2024
Closest in time.
“WARM: On the Benefits of Weight Averaged Reward Models”, 2024
Alexandre Ramé et al · 2024
Closest in time.
“Group Robust Preference Optimization in Reward-free RLHF”
Shyam Ramesh et al · 2024
Closest in time.
“Insights into Alignment: Evaluating DPO and its Variants Across Multiple Tasks”, 2024
Amir Saeidi, Shivanshu Verma and Chitta Baral · 2024
Closest in time.
“DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models”, 2024
Zhihong Shao et al · 2024
Closest in time.
“MAPO: Advancing Multilingual Reasoning through Multilingual Alignment-as-Preference Optimization”
Shuaijie She et al · 2024
Closest in time.
Feifan Song et al · 2024
Closest in time.
“Preference ranking optimization for human alignment”
Feifan Song et al · 2024
Closest in time.
“Scaling Data Diversity for Fine-Tuning Language Models in Human Alignment”
Feifan Song et al · 2024
Closest in time.
“Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context”, 2024
Gemini Team et al · 2024
Closest in time.
“Conditioned language policy: A general framework for steerable multi-objective finetuning”
Kaiwen Wang et al · 2024
Closest in time.
“Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations”, 2024
Peiyi Wang et al · 2024
Closest in time.
“On the Essence and Prospect: An Investigation of Alignment Approaches for Big Models”, 2024
Xinpeng Wang et al · 2024
Closest in time.
“Jailbroken: How does llm safety training fail?”
Alexander Wei, Nika Haghtalab and Jacob Steinhardt · 2024
Closest in time.
“Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge”
Tianhao Wu et al · 2024
Closest in time.
“Self-play preference optimization for language model alignment”
Yue Wu et al · 2024
Closest in time.
“Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning”, 2024
Yuxi Xie et al · 2024
Closest in time.
Huajian Xin et al · 2024
Closest in time.
“DeepSeek-Prover: Advancing Theorem Proving in LLMs through Large-Scale Synthetic Data”, 2024
Huajian Xin et al · 2024
Closest in time.
“Watch Every Step! LLM Agent Learning via Iterative Step-Level Process Refinement”, 2024
Weimin Xiong et al · 2024
Closest in time.
Haoran Xu et al · 2024
Closest in time.
“The Perfect Blend: Redefining RLHF with Mixture of Judges”
Tengyu Xu et al · 2024
Closest in time.
“3D-Properties: Identifying Challenges in DPO and Charting a Path Forward”
Yuzi Yan et al · 2024
Closest in time.
“Qwen2 Technical Report”, 2024
An Yang et al · 2024
Closest in time.
Kailai Yang et al · 2024
Closest in time.
“Regularizing Hidden States Enables Learning Generalizable Reward Model for LLMs”
Rui Yang et al · 2024
Closest in time.
Rui Yang et al · 2024
Closest in time.
“Beyond Scalar Reward Model: Learning Generative Judge from Preference Data”
Ziyi Ye et al · 2024
Closest in time.
“OVM, Outcome-supervised Value Models for Planning in Mathematical Reasoning”, 2024
Fei Yu, Anningzhe Gao and Benyou Wang · 2024
Closest in time.
“Direct Alignment of Language Models via Quality-Aware Self-Refinement”
Runsheng Yu et al · 2024
Closest in time.
“RRHF: Rank responses to align language models with human feedback”
Hongyi Yuan et al · 2024
Closest in time.
“Self-rewarding language models”
Weizhe Yuan et al · 2024
Closest in time.
Chen Zhang et al · 2024
Closest in time.
“ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search”, 2024
Dan Zhang et al · 2024
Closest in time.
Di Zhang et al · 2024
Closest in time.
“Generative Verifiers: Reward Modeling as Next-Token Prediction”, 2024
Lunjun Zhang et al · 2024
Closest in time.
“Improving Reinforcement Learning from Human Feedback with Efficient Reward Model Ensemble”, 2024
Shun Zhang et al · 2024
Closest in time.
“Judging llm-as-a-judge with mt-bench and chatbot arena”
Lianmin Zheng et al · 2024
Closest in time.
“Prior Constraints-based Reward Model Training for Aligning Large Language Models”, 2024
Hang Zhou et al · 2024
Closest in time.
“Recovering Mental Representations from Large Language Models with Markov Chain Monte Carlo”, 2024
Jian-Qiao Zhu, Haijiang Yan and Thomas. Griffiths · 2024
Closest in time.
“LIRE: listwise reward enhancement for preference alignment”
Mingye Zhu et al · 2024
Closest in time.