Fetching the paper…
Reading the bibliography…
Modern large language model (LLM) alignment techniques rely on human feedback, but it is unclear whether these techniques fundamentally limit the capabilities of aligned LLMs.
A mixture of experts classifier with learning based on both labelled and unlabelled data
David J Miller and Hasan Uyar · 1996
Earlier work this paper cites.
Combining labeled and unlabeled data with co-training
Avrim Blum and Tom Mitchell · 1998
Earlier work this paper cites.
Learning from labeled and unlabeled data using graph mincuts
Avrim Blum and Shuchi Chawla · 2001
Earlier work this paper cites.
Learning with local and global consistency
Dengyong Zhou, Olivier Bousquet, Thomas Lal, Jason Weston, and Bernhard Schölkopf · 2003
Earlier work this paper cites.
Semi-supervised learning using Gaussian fields and harmonic functions
Xiaojin Zhu, Zoubin Ghahramani, and John Lafferty · 2003
Earlier work this paper cites.
Semi-supervised learning by entropy minimization
Yves Grandvalet and Yoshua Bengio · 2004
Earlier work this paper cites.
Semi-Supervised Learning
Olivier Chapelle, Bernhard Schölkopf, and Alexander Zien (eds.) · 2006
Earlier work this paper cites.
Correcting sample selection bias by unlabeled data
Jiayuan Huang, Arthur Gretton, Karsten Borgwardt, Bernhard Schölkopf, and Alex Smola · 2006
Earlier work this paper cites.
Boosting for transfer learning
Wenyuan Dai, Qiang Yang, Gui-Rong Xue, and Yong Yu · 2007
Earlier work this paper cites.
A survey on transfer learning
Sinno Jialin Pan and Qiang Yang · 2009
Earlier work this paper cites.
A review of multi-instance learning assumptions
James Foulds and Eibe Frank · 2010
Earlier work this paper cites.
Convex and scalable weakly labeled svms, 2013
Yu-Feng Li, Ivor W. Tsang, James T. Kwok, and Zhi-Hua Zhou · 2013
Earlier work this paper cites.
Multi-source domain adaptation: A causal view
Kun Zhang, Mingming Gong, and Bernhard Scholkopf · 2015
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul Christiano, Jan Leike, Tom B. Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Earlier work this paper cites.
Temporal ensembling for semi-supervised learning
Samuli Laine and Timo Aila · 2017
Earlier work this paper cites.
Self-ensembling for visual domain adaptation
Geoffrey French, Michal Mackiewicz, and Mark Fisher · 2018
Earlier work this paper cites.
Co-teaching: Robust training of deep neural networks with extremely noisy labels, 2018
Bo Han, Quanming Yao, Xingrui Yu, Gang Niu, Miao Xu, Weihua Hu, Ivor Tsang, and Masashi Sugiyama · 2018
Earlier work this paper cites.
Marginal Singularity, and the Benefits of Labels in Covariate-Shift
Samory Kpotufe and Guillaume Martinet · 2018
Earlier work this paper cites.
Detecting and Correcting for Label Shift with Black Box Predictors
Zachary Lipton, Yu-Xiang Wang, and Alexander Smola · 2018
Earlier work this paper cites.
A DIRT-T Approach to Unsupervised Domain Adaptation
Rui Shu, Hung Bui, Hirokazu Narui, and Stefano Ermon · 2018
Earlier work this paper cites.
Generalized cross entropy loss for training deep neural networks with noisy labels, 2018
Zhilu Zhang and Mert R. Sabuncu · 2018
Earlier work this paper cites.
A brief introdution to weakly supervised learning
Zhi-Hua Zou · 2018
Earlier work this paper cites.
Transfer Learning for Nonparametric Classification: Minimax Rate and Adaptive Classifier
T. Tony Cai and Hongji Wei · 2019
Earlier work this paper cites.
Using trusted data to train deep networks on labels corrupted by severe noise, 2019
Dan Hendrycks, Mantas Mazeika, Duncan Wilson, and Kevin Gimpel · 2019
Earlier work this paper cites.
Probabilistic end-to-end noise correction for learning with noisy labels, 2019
Kun Yi and Jianxin Wu · 2019
Earlier work this paper cites.
Language Models are Few-Shot Learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Earlier work this paper cites.
Dividemix: Learning with noisy labels as semi-supervised learning, 2020
Junnan Li, Richard Socher, and Steven C. H. Hoi · 2020
Earlier work this paper cites.
A computationally efficient classification algorithm in posterior drift model: Phase transition and minimax adaptivity, 2020
Ruiqi Liu, Kexuan Li, and Zuofeng Shang · 2020
Cited alongside, same era.
Normalized loss functions for deep learning with noisy labels, 2020
Xingjun Ma, Hanxun Huang, Yisen Wang, Simone Romano, Sarah Erfani, and James Bailey · 2020
Cited alongside, same era.
Minimax optimal approaches to the label shift problem
Subha Maity, Yuekai Sun, and Moulinath Banerjee · 2020
Cited alongside, same era.
Snorkel: Rapid training data creation with weak supervision
Alexander Ratner, Stephen H. Bach, Henry Ehrenberg, Jason Fries, Sen Wu, and Christopher Ré · 2020
Cited alongside, same era.
Self-training with noisy student improves imagenet classification
Qizhe Xie, Minh-Thang Luong, Eduard Hovy, and Quoc V. Le · 2020
Cited alongside, same era.
A survey on in-context learning, 2023
Qingxiu Dong, Lei Li, Damai Dai, Ce Zheng, Zhiyong Wu, Baobao Chang, Xu Sun, Jingjing Xu, Lei Li, and Zhifang Sui · 2023
Later among the works it cites.
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed · 2023
Later among the works it cites.
A survey of reinforcement learning from human feedback, 2023
Timo Kaufmann, Paul Weng, Viktor Bengs, and Eyke Hüllermeier · 2023
Later among the works it cites.
Gpteval: Nlg evaluation using gpt-4 with better human alignment
Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu · 2023
Later among the works it cites.
Automatically correcting large language models: Surveying the landscape of diverse self-correction strategies, 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A comprehensive survey on transfer learning, 2020
Fuzhen Zhuang, Zhiyuan Qi, Keyu Duan, Dongbo Xi, Yongchun Zhu, Hengshu Zhu, Hui Xiong, and Qing He · 2020
Cited alongside, same era.
Eliciting latent knowledge: How to tell if your eyes deceive you
Paul Christiano, Ajeya Cotra, and Mark Xu · 2021
Cited alongside, same era.
Training verifiers to solve math word problems, 2021
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman · 2021
Cited alongside, same era.
Truthful ai: Developing and governing ai that does not lie, 2021
Owain Evans, Owen Cotton-Barratt, Lukas Finnveden, Adam Bales, Avital Balwit, Peter Wills, Luca Righetti, and William Saunders · 2021
Cited alongside, same era.
Measuring mathematical problem solving with the math dataset, 2021
Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt · 2021
Cited alongside, same era.
A linear adjustment based approach to posterior drift in transfer learning
Subha Maity, Diptavo Dutta, Jonathan Terhorst, Yuekai Sun, and Moulinath Banerjee · 2021
Cited alongside, same era.
An Explanation of In-context Learning as Implicit Bayesian Inference
Sang Michael Xie, Aditi Raghunathan, Percy Liang, and Tengyu Ma · 2021
Cited alongside, same era.
Liangming Pan, Michael Saxon, Wenda Xu, Deepak Nathani, Xinyi Wang, and William Yang Wang · 2023
Later among the works it cites.
Universalizing weak supervision, 2023
Changho Shin, Winfred Li, Harit Vishwakarma, Nicholas Roberts, and Frederic Sala · 2023
Later among the works it cites.
Bayesian transfer learning, 2023
Piotr M. Suder, Jason Xu, and David B. Dunson · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models, 2023
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom · 2023
Later among the works it cites.
Least-to-most prompting enables complex reasoning in large language models
Denny Zhou, Nathanael Schärli, Le Hou, Jason Wei, Nathan Scales, Xuezhi Wang, Dale Schuurmans, Claire Cui, Olivier Bousquet, Quoc V Le, and Ed H. Chi · 2023
Later among the works it cites.
Gpt-4 technical report, 2024
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, and Others · 2024
Closest in time.
Quantifying the gain in weak-to-strong generalization, 2024
Moses Charikar, Chirag Pabbaraju, and Kirankumar Shiragur · 2024
Closest in time.
On giant’s shoulders: Effortless weak to strong by dynamic logits fusion, 2024
Chenghao Fan, Zhenyi Lu, Wei Wei, Jie Tian, Xiaoye Qu, Dangyang Chen, and Yu Cheng · 2024
Closest in time.
The unreasonable effectiveness of easy training data for hard tasks, 2024
Peter Hase, Mohit Bansal, Peter Clark, and Sarah Wiegreffe · 2024
Closest in time.
Aligner: Achieving efficient alignment through weak-to-strong correction, 2024
Jiaming Ji, Boyuan Chen, Hantao Lou, Donghai Hong, Borong Zhang, Xuehai Pan, Juntao Dai, and Yaodong Yang · 2024
Closest in time.
Theoretical analysis of weak-to-strong generalization, 2024
Hunter Lang, David Sontag, and Aravindan Vijayaraghavan · 2024
Closest in time.
Introducing superalignment
Jan Leike and Ilya Sutskever · 2024
Closest in time.
Tuning language models by proxy, 2024
Alisa Liu, Xiaochuang Han, Yizhong Wang, Yulia Tsvetkov, Yejin Choi, and Noah A. Smith · 2024
Closest in time.
tinybenchmarks: evaluating llms with fewer examples
Felipe Maia Polo, Lucas Weber, Leshem Choshen, Yuekai Sun, Gongjun Xu, and Mikhail Yurochkin · 2024
Closest in time.
Aligners: Decoupling llms and alignment, 2024
Lilian Ngweta, Mayank Agarwal, Subha Maity, Alex Gittens, Yuekai Sun, and Mikhail Yurochkin · 2024
Closest in time.
Gpt-4o mini: advancing cost-efficient intelligence
OpenAI · 2024
Closest in time.
Transformers can optimally learn regression mixture models
Reese Pathak, Rajat Sen, Weihao Kong, and Abhimanyu Das · 2024
Closest in time.
Easy-to-hard generalization: Scalable alignment beyond human supervision, 2024
Zhiqing Sun, Longhui Yu, Yikang Shen, Weiyang Liu, Yiming Yang, Sean Welleck, and Chuang Gan · 2024
Closest in time.
Gemma: Open models based on gemini research and technology, 2024
Gemma Team and Others · 2024
Closest in time.
Xinyi Wang, Wanrong Zhu, Michael Saxon, Mark Steyvers, and William Yang Wang · 2024
Closest in time.
Provable weak-to-strong generalization via benign overfitting, 2024
David X. Wu and Anant Sahai · 2024
Closest in time.
Transcendence: Generative models can outperform the experts that train them, 2024
Edwin Zhang, Vincent Zhu, Naomi Saphra, Anat Kleiman, Benjamin L. Edelman, Milind Tambe, Sham M. Kakade, and Eran Malach · 2024
Closest in time.
Living on the edge: Phase transitions in convex programs with random data
Dennis Amelunxen, Martin Lotz, Michael B. McCoy, and Joel A. Tropp · 2049
Closest in time.