Fetching the paper…
Reading the bibliography…
Although preference optimization methods have improved reasoning performance in Large Language Models (LLMs), they often lack transparency regarding why one reasoning outcome is preferred over another.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Ralph Allan Bradley and Milton E. Terry. 1952 · 1952
Earlier work this paper cites.
Automatic essay grading using text categorization techniques
Leah S. Larkey. 1998 · 1998
Earlier work this paper cites.
The hewlett foundation: Automated essay scoring
Ben Hamner, Jaison Morgan, Mark Shermis Lynnvandev, and Tom Vander Ark. 2012 · 2012
Earlier work this paper cites.
Automatic text scoring using neural networks
Dimitrios Alikaniotis, Helen Yannakoudakis, and Marek Rei. 2016 · 2016
Earlier work this paper cites.
Automatic features for essay scoring – an empirical study
Fei Dong and Yue Zhang. 2016 · 2016
Earlier work this paper cites.
A neural approach to automated essay scoring
Kaveh Taghipour and Hwee Tou Ng. 2016 · 2016
Earlier work this paper cites.
Attention-based recurrent convolutional neural network for automatic essay scoring
Fei Dong, Yue Zhang, and Jie Yang. 2017 · 2017
Earlier work this paper cites.
Should you fine-tune BERT for automated essay scoring?
Elijah Mayfield and Alan W Black. 2020 · 2020
Earlier work this paper cites.
Enhancing automated essay scoring performance via fine-tuning pre-trained language models with combination of regression and ranking
Ruosong Yang, Jiannong Cao, Zhiyuan Wen, Youzheng Wu, and Xiaodong He. 2020 · 2020
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F Christiano, Jan Leike, and Ryan Lowe. 2022 · 2022
Earlier work this paper cites.
RL4F: Generating natural language feedback with reinforcement learning for repairing model outputs
Afra Feyza Akyurek, Ekin Akyurek, Ashwin Kalyan, Peter Clark, Derry Tanti Wijaya, and Niket Tandon. 2023 · 2023
Earlier work this paper cites.
DeBERTav3: Improving deBERTa using ELECTRA-style pre-training with gradient-disentangled embedding sharing
Pengcheng He, Jianfeng Gao, and Weizhu Chen. 2023 · 2023
Earlier work this paper cites.
Active retrieval augmented generation
Zhengbao Jiang, Frank Xu, Luyu Gao, Zhiqing Sun, Qian Liu, Jane Dwivedi-Yu, Yiming Yang, Jamie Callan, and Graham Neubig. 2023 · 2023
Earlier work this paper cites.
Language models can solve computer tasks
Geunwoo Kim, Pierre Baldi, and Stephen Marcus McAleer. 2023 · 2023
Earlier work this paper cites.
Efficient memory management for large language model serving with pagedattention
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, and Ion Stoica. 2023 · 2023
Earlier work this paper cites.
Distilling ChatGPT for explainable automated student answer assessment
Jiazheng Li, Lin Gui, Yuxiang Zhou, David West, Cesare Aloisi, and Yulan He. 2023a · 2023
Earlier work this paper cites.
Self-refine: Iterative refinement with self-feedback
Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, Shashank Gupta, Bodhisattwa Prasad Majumder, Katherine Hermann, Sean Welleck, Amir Yazdanbakhsh, and Peter Clark. 2023 · 2023
Earlier work this paper cites.
A note on dpo with noisy preferences & relationship to ipo
Eric Mitchell. 2023 · 2023
Cited alongside, same era.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn. 2023 · 2023
Cited alongside, same era.
Reflexion: language agents with verbal reinforcement learning
Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik R Narasimhan, and Shunyu Yao. 2023 · 2023
Cited alongside, same era.
Autograder: A feature-based quantitative essay grading system using bert
Roopchand Reddy Vanga, C. Sindhu, M. S. Bharath, T. Charandeep Reddy, and Meghana Kanneganti. 2023 · 2023
Cited alongside, same era.
Generating sequences by learning to self-correct
Sean Welleck, Ximing Lu, Peter West, Faeze Brahman, Tianxiao Shen, Daniel Khashabi, and Yejin Choi. 2023 · 2023
Cited alongside, same era.
Llama 3 model card
AI@Meta. 2024 · 2024
Iterative reasoning preference optimization
Richard Yuanzhe Pang, Weizhe Yuan, He He, Kyunghyun Cho, Sainbayar Sukhbaatar, and Jason E Weston. 2024 · 2024
Later among the works it cites.
REFINER: Reasoning feedback on intermediate representations
Debjit Paul, Mete Ismayilzada, Maxime Peyrard, Beatriz Borges, Antoine Bosselut, Robert West, and Boi Faltings. 2024 · 2024
Later among the works it cites.
Qwen2.5 technical report
Qwen, An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, Huan Lin, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxin Yang, Jingren Zhou, Junyang Lin, Kai Dang, Keming Lu, Keqin Bao, Kexin Yang, Le Yu, Mei Li, Mingfeng Xue, Pei Zhang, Qin Zhu, Rui Men, Runji Lin, Tianhao Li, Tingyu Xia, Xingzhang Ren, Xuancheng Ren, Yang Fan, Yang Su, Yi-Chao Zhang, Yunyang Wan, Yuqi Liu, Zeyu Cui, Zhenru Zhang, Zihan Qiu, and Shanghaoran Quan. 2024 · 2024
Later among the works it cites.
Qwen2.5: A party of foundation models
QwenTeam. 2024 · 2024
Later among the works it cites.
Scaling laws for reward model overoptimization in direct alignment algorithms
Rafael Rafailov, Yaswanth Chittepu, Ryan Park, Harshit Sikchi, Joey Hejna, W. Bradley Knox, Chelsea Finn, and Scott Niekum. 2024 · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Can LLM-generated misinformation be detected?
Canyu Chen and Kai Shu. 2024 · 2024
Cited alongside, same era.
Step-level value preference optimization for mathematical reasoning
Guoxin Chen, Minpeng Liao, Chengxi Li, and Kai Fan. 2024a · 2024
Cited alongside, same era.
Provably robust dpo: aligning language models with noisy feedback
Sayak Ray Chowdhury, Anush Kini, and Nagarajan Natarajan. 2024 · 2024
Cited alongside, same era.
PACE: Improving prompt with actor-critic editing for large language model
Yihong Dong, Kangcheng Luo, Xue Jiang, Zhi Jin, and Ge Li. 2024 · 2024
Cited alongside, same era.
Large language models cannot self-correct reasoning yet
Jie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng, Adams Wei Yu, Xinying Song, and Denny Zhou. 2024 · 2024
Cited alongside, same era.
Step-dpo: Step-wise preference optimization for long-chain reasoning of llms
Xin Lai, Zhuotao Tian, Yukang Chen, Senqiao Yang, Xiangru Peng, and Jiaya Jia. 2024 · 2024
Cited alongside, same era.
Later among the works it cites.
Can LLMs learn from previous mistakes? investigating LLMs’ errors to boost for reasoning
Yongqi Tong, Dawei Li, Sizhe Wang, Yujia Wang, Fei Teng, and Jingbo Shang. 2024 · 2024
Later among the works it cites.
LLMs cannot find reasoning errors, but can correct them given the error location
Gladys Tyen, Hassan Mansoor, Victor Carbune, Peter Chen, and Tony Mak. 2024 · 2024
Later among the works it cites.
Math-shepherd: Verify and reinforce LLMs step-by-step without human annotations
Peiyi Wang, Lei Li, Zhihong Shao, Runxin Xu, Damai Dai, Yifei Li, Deli Chen, Yu Wu, and Zhifang Sui. 2024 · 2024
Later among the works it cites.
How interpretable are reasoning explanations from prompting large language models?
Yeo Wei Jie, Ranjan Satapathy, Rick Goh, and Erik Cambria. 2024 · 2024
Later among the works it cites.
Yueqin Yin, Zhendong Wang, Yi Gu, Hai Huang, Weizhu Chen, and Mingyuan Zhou. 2024 · 2024
Later among the works it cites.
LlamaFactory: Unified efficient fine-tuning of 100+ language models
Yaowei Zheng, Richong Zhang, Junhao Zhang, Yanhan Ye, and Zheyan Luo. 2024 · 2024
Later among the works it cites.
The mystery of in-context learning: A comprehensive survey on interpretation and analysis
Yuxiang Zhou, Jiazheng Li, Yanzheng Xiang, Hanqi Yan, Lin Gui, and Yulan He. 2024 · 2024
Later among the works it cites.
Bridging and modeling correlations in pairwise data for direct preference optimization
Yuxin Jiang, Bo Huang, Yufei Wang, Xingshan Zeng, Liangyou Li, Yasheng Wang, Xin Jiang, Lifeng Shang, Ruiming Tang, and Wei Wang. 2025 · 2025
Closest in time.
Training language models to self-correct via reinforcement learning
Aviral Kumar, Vincent Zhuang, Rishabh Agarwal, Yi Su, John D. Co-Reyes, Avi Singh, Kate Baumli, Shariq Iqbal, Colton Bishop, Rebecca Roelofs, Lei M. Zhang, Kay McKinney, Disha Shrivastava, Cosmin Paduraru, George Tucker, Doina Precup, Feryal M. P. Behbahani, and Aleksandra Faust. 2025 · 2025
Closest in time.
An automated explainable educational assessment system built on llms
Jiazheng Li, Artem Bobrov, David West, Cesare Aloisi, and Yulan He. 2025 · 2025
Closest in time.
Scaling test-time compute optimally can be more effective than scaling LLM parameters
Charlie Snell, Jaehoon Lee, Kelvin Xu, and Aviral Kumar. 2025 · 2025
Closest in time.