Fetching the paper…
Reading the bibliography…
Recent advances in large language models (LLMs) have yielded impressive performance on various tasks, yet they often depend on high-quality feedback that can be costly.
Probability of error of some adaptive pattern-recognition machines
Henry J. Scudder · 1965
Earlier work this paper cites.
Learning from labeled and unlabeled data with label propagation
Xiaojin Zhu and Zoubin Ghahramani · 2002
Earlier work this paper cites.
Curriculum learning
Yoshua Bengio, Jérôme Louradour, Ronan Collobert, and Jason Weston · 2009
Earlier work this paper cites.
Self-paced learning for latent variable models
M. Kumar, Benjamin Packer, and Daphne Koller · 2010
Earlier work this paper cites.
Pseudo-label: The simple and efficient semi-supervised learning method for deep neural networks
Dong-Hyun Lee · 2013
Earlier work this paper cites.
Big data: A review
Seref Sagiroglu and Duygu Sinanc · 2013
Earlier work this paper cites.
Classification with asymmetric label noise: Consistency and maximal denoising
Clayton Scott, Gilles Blanchard, and Gregory Handy · 2013
Earlier work this paper cites.
Learning with pseudo-ensembles
Philip Bachman, Ouais Alsharif, and Doina Precup · 2014
Earlier work this paper cites.
Self-paced curriculum learning
Lu Jiang, Deyu Meng, Qian Zhao, Shiguang Shan, and Alexander Hauptmann · 2015
Earlier work this paper cites.
Learning from corrupted binary labels via class-probability estimation
Aditya Menon, Brendan Van Rooyen, Cheng Soon Ong, and Bob Williamson · 2015
Earlier work this paper cites.
Classification with noisy labels by importance reweighting
Tongliang Liu and Dacheng Tao · 2016
Earlier work this paper cites.
Regularization with stochastic transformations and perturbations for deep semi-supervised learning
Mehdi Sajjadi, Mehran Javanmardi, and Tolga Tasdizen · 2016
Earlier work this paper cites.
Temporal ensembling for semi-supervised learning
Samuli Laine and Timo Aila · 2017
Earlier work this paper cites.
Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results
Antti Tarvainen and Harri Valpola · 2017
Earlier work this paper cites.
Detecting opinion spams and fake news using text classification
Hadeer Ahmed, Issa Traore, and Sherif Saad · 2018
Earlier work this paper cites.
Virtual adversarial training: A regularization method for supervised and semi-supervised learning
Takeru Miyato, Shin-Ichi Maeda, Masanori Koyama, and Shin Ishii · 2018
Earlier work this paper cites.
On the minimal supervision for training any binary classifier from only unlabeled data
Nan Lu, Gang Niu, Aditya K Menon, and Masashi Sugiyama · 2019
Earlier work this paper cites.
Probabilistic end-to-end noise correction for learning with noisy labels
Kun Yi and Jianxin Wu · 2019
Earlier work this paper cites.
Learning from positive and unlabeled data: a survey
Jessa Bekker and Jesse Davis · 2020
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Earlier work this paper cites.
Mitigating overfitting in supervised classification from two unlabeled datasets: A consistent risk correction approach
Nan Lu, Tianyi Zhang, Gang Niu, and Masashi Sugiyama · 2020
Earlier work this paper cites.
Fixmatch: Simplifying semi-supervised learning with consistency and confidence
Kihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang, Han Zhang, Colin A Raffel, Ekin Dogus Cubuk, Alexey Kurakin, and Chun-Liang Li · 2020
Cited alongside, same era.
Repetitive reprediction deep decipher for semi-supervised learning
Guo-Hua Wang and Jianxin Wu · 2020
Cited alongside, same era.
Iterative label improvement: Robust training by confidence based filtering and dataset partitioning
Christian Haase-Schutz, Rainer Stal, Heinz Hertlein, and Bernhard Sick · 2021
Cited alongside, same era.
Meta pseudo labels
Hieu Pham, Zihang Dai, Qizhe Xie, and Quoc V. Le · 2021
Cited alongside, same era.
In defense of pseudo-labeling: An uncertainty-aware pseudo-label selection framework for semi-supervised learning
Mamshad Nayeem Rizve, Kevin Duarte, Yogesh S Rawat, and Mubarak Shah · 2021
Cited alongside, same era.
PIEClass: Weakly-supervised text classification with prompting and noise-robust iterative ensemble training
Yunyi Zhang, Minhao Jiang, Yu Meng, Yu Zhang, and Jiawei Han · 2023
Later among the works it cites.
Solving math word problems via cooperative reasoning induced language models
Xinyu Zhu, Junjie Wang, Lin Zhang, Yuxiang Zhang, Yongfeng Huang, Ruyi Gan, Jiaxing Zhang, and Yujiu Yang · 2023
Later among the works it cites.
Best-of-venom: Attacking RLHF by injecting poisoned preference data
Tim Baumgärtner, Yang Gao, Dana Alon, and Donald Metzler · 2024
Later among the works it cites.
Safe RLHF: Safe reinforcement learning from human feedback
Josef Dai, Xuehai Pan, Ruiyang Sun, Jiaming Ji, Xinbo Xu, Mickel Liu, Yizhou Wang, and Yaodong Yang · 2024
Later among the works it cites.
Improving factuality and reasoning in language models through multiagent debate
Yilun Du, Shuang Li, Antonio Torralba, Joshua B Tenenbaum, and Igor Mordatch · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ana-Cristina Rogoz, Gaman Mihaela, and Radu Tudor Ionescu · 2021
Cited alongside, same era.
Constitutional ai: Harmlessness from ai feedback, 2022
Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, Carol Chen, Catherine Olsson, Christopher Olah, Danny Hernandez, Dawn Drain, Deep Ganguli, Dustin Li, Eli Tran-Johnson, Ethan Perez, Jamie Kerr, Jared Mueller, Jeffrey Ladish, Joshua Landau, Kamal Ndousse, Kamile Lukosuite, Liane Lovitt, Michael Sellitto, Nelson Elhage, Nicholas Schiefer, Noemi Mercado, Nova DasSarma, Robert Lasenby, Robin Larson, Sam Ringer, Scott Johnston, Shauna Kravec, Sheer El Showk, Stanislav Fort, Tamera Lanham, Timothy Telleen-Lawton, Tom Conerly, Tom Henighan, Tristan Hume, Samuel R. Bowman, Zac Hatfield-Dodds, Ben Mann, Dario Amodei, Nicholas Joseph, Sam McCandlish, Tom Brown, and Jared Kaplan · 2022
Cited alongside, same era.
Language models for the prediction of SARS-CoV-2 inhibitors
Andrew E Blanchard, John Gounley, Debsindhu Bhowmik, Mayanka Chandra Shekar, Isaac Lyngaas, Shang Gao, Junqi Yin, Aristeidis Tsaris, Feiyi Wang, and Jens Glaser · 2022
Cited alongside, same era.
SAT: Improving semi-supervised text classification with simple instance-adaptive self-training
Hui Chen, Wei Han, and Soujanya Poria · 2022
Cited alongside, same era.
Inner monologue: Embodied reasoning through planning with language models
Wenlong Huang, Fei Xia, Ted Xiao, Harris Chan, Jacky Liang, Pete Florence, Andy Zeng, Jonathan Tompson, Igor Mordatch, Yevgen Chebotar, Pierre Sermanet, Tomas Jackson, Noah Brown, Linda Luu, Sergey Levine, Karol Hausman, and Brian Ichter · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F Christiano, Jan Leike, and Ryan Lowe · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc V Le, and Denny Zhou · 2022
Cited alongside, same era.
Multi-LLM debate: Framework, principals, and interventions
Andrew Estornell and Yang Liu · 2024
Later among the works it cites.
Large language models cannot self-correct reasoning yet
Jie Huang, Xinyun Chen, Swaroop Mishra, Huaixiu Steven Zheng, Adams Wei Yu, Xinying Song, and Denny Zhou · 2024
Later among the works it cites.
Benchmarking cognitive biases in large language models as evaluators
Ryan Koo, Minhwa Lee, Vipul Raheja, Jong Inn Park, Zae Myung Kim, and Dongyeop Kang · 2024
Later among the works it cites.
Contrastive credibility propagation for reliable semi-supervised learning
Brody Kutt, Pralay Ramteke, Xavier Mignot, Pamela Toman, Nandini Ramanan, Sujit Rokka Chhetri, Shan Huang, Min Du, and William Hewlett · 2024
Later among the works it cites.
RLAIF vs. RLHF: scaling reinforcement learning from human feedback with AI feedback
Harrison Lee, Samrat Phatale, Hassan Mansoor, Thomas Mesnard, Johan Ferret, Kellie Lu, Colton Bishop, Ethan Hall, Victor Carbune, Abhinav Rastogi, and Sushant Prakash · 2024
Later among the works it cites.
When hindsight is not 20/20: Testing limits on reflective thinking in large language models
Yanhong Li, Chenghao Yang, and Allyson Ettinger · 2024
Later among the works it cites.
LLMs as narcissistic evaluators: When ego inflates evaluation scores
Yiqi Liu, Nafise Moosavi, and Chenghua Lin · 2024
Later among the works it cites.
EHRAgent: Code empowers large language models for few-shot complex tabular reasoning on electronic health records
Wenqi Shi, Ran Xu, Yuchen Zhuang, Yue Yu, Jieyu Zhang, Hang Wu, Yuanda Zhu, Joyce C Ho, Carl Yang, and May Dongmei Wang · 2024
Later among the works it cites.
Should we be going MAD? a look at multi-agent debate strategies for LLMs
Andries Smit, Nathan Grinsztajn, Paul Duckworth, Thomas D Barrett, and Arnu Pretorius · 2024
Later among the works it cites.
Generalized preference optimization: A unified approach to offline alignment
Yunhao Tang, Zhaohan Daniel Guo, Zeyu Zheng, Daniele Calandriello, Remi Munos, Mark Rowland, Pierre Harvey Richemond, Michal Valko, Bernardo Avila Pires, and Bilal Piot · 2024
Later among the works it cites.
Large language models are not fair evaluators
Peiyi Wang, Lei Li, Liang Chen, Zefan Cai, Dawei Zhu, Binghuai Lin, Yunbo Cao, Lingpeng Kong, Qi Liu, Tianyu Liu, and Zhifang Sui · 2024
Later among the works it cites.
AvaTaR: Optimizing LLM agents for tool usage via contrastive reasoning
Shirley Wu, Shiyu Zhao, Qian Huang, Kexin Huang, Michihiro Yasunaga, Kaidi Cao, Vassilis N Ioannidis, Karthik Subbian, Jure Leskovec, and James Zou · 2024
Later among the works it cites.
Starling-7B: Improving helpfulness and harmlessness with RLAIF
Banghua Zhu, Evan Frick, Tianhao Wu, Hanlin Zhu, Karthik Ganesan, Wei-Lin Chiang, Jian Zhang, and Jiantao Jiao · 2024
Later among the works it cites.
DeepSeek-R1: Incentivizing reasoning capability in LLMs via reinforcement learning
DeepSeek-AI · 2025
Closest in time.
Foundations of large language models, 2025
Tong Xiao and Jingbo Zhu · 2025
Closest in time.