Fetching the paper…
Reading the bibliography…
Recent work in language modeling has raised the possibility of self-improvement, where a language models evaluates and refines its own generations to achieve higher performance without external feedback.
The complexity of theorem-proving procedures
Stephen A Cook · 1971
Earlier work this paper cites.
Reducibility among combinatorial problems
Richard M Karp · 1972
Earlier work this paper cites.
Universal sequential search problems
Leonid Anatolevich Levin · 1973
Earlier work this paper cites.
On the computational complexity of ising spin glass models
Francisco Barahona · 1982
Earlier work this paper cites.
Optimization by simulated annealing
Scott Kirkpatrick, C Daniel Gelatt Jr, and Mario P Vecchi · 1983
Earlier work this paper cites.
Problem complexity and method efficiency in optimization
Arkadii Nemirovski, David Borisovich Yudin, and Edgar Ronald Dawson · 1983
Earlier work this paper cites.
Information-based complexity
Joseph F Traub, Grzegorz W Wasilkowski, and Henryk Woźniakowski · 1988
Earlier work this paper cites.
Weakly learning DNF and characterizing statistical query learning using Fourier analysis
Avrim Blum, Merrick Furst, Jeffrey Jackson, Michael Kearns, Yishay Mansour, and Steven Rudich · 1994
Earlier work this paper cites.
Controlling the false discovery rate: a practical and powerful approach to multiple testing
Yoav Benjamini and Yosef Hochberg · 1995
Earlier work this paper cites.
Probability inequalities for likelihood ratios and convergence rates of sieve mles
Wing Hung Wong and Xiaotong Shen · 1995
Earlier work this paper cites.
Efficient noise-tolerant learning from statistical queries
Michael Kearns · 1998
Earlier work this paper cites.
Elements of information theory
Thomas M Cover · 1999
Earlier work this paper cites.
Empirical Processes in M-Estimation
S. A. van de Geer · 2000
Earlier work this paper cites.
Variational algorithms for approximate Bayesian inference
Matthew James Beal · 2003
Earlier work this paper cites.
Semi-supervised learning by entropy minimization
Yves Grandvalet and Yoshua Bengio · 2004
Earlier work this paper cites.
Model compression
Cristian Buciluǎ, Rich Caruana, and Alexandru Niculescu-Mizil · 2006
Earlier work this paper cites.
Fast algorithms for logconcave functions: Sampling, rounding, integration and optimization
László Lovász and Santosh Vempala · 2006
Earlier work this paper cites.
From ϵ \epsilon -entropy to KL-entropy: Analysis of minimum information complexity density estimation
Tong Zhang · 2006
Earlier work this paper cites.
Error propagation for approximate policy and value iteration
Amir-massoud Farahmand, Csaba Szepesvári, and Rémi Munos · 2010
Earlier work this paper cites.
Information-based complexity, feedback and dynamics in convex programming
Maxim Raginsky and Alexander Rakhlin · 2011
Earlier work this paper cites.
Information-theoretic lower bounds on the oracle complexity of stochastic convex optimization
Alekh Agarwal, Peter L Bartlett, Pradeep Ravikumar, and Martin J Wainwright · 2012
Earlier work this paper cites.
A complete characterization of statistical query learning with applications to evolvability
Vitaly Feldman · 2012
Earlier work this paper cites.
Eluder dimension and the sample complexity of optimistic exploration
Daniel Russo and Benjamin Van Roy · 2013
Earlier work this paper cites.
Taming the monster: A fast and simple algorithm for contextual bandits
Alekh Agarwal, Daniel Hsu, Satyen Kale, John Langford, Lihong Li, and Robert Schapire · 2014
Earlier work this paper cites.
Amortized inference in probabilistic reasoning
Samuel Gershman and Noah Goodman · 2014
Earlier work this paper cites.
Entropy, optimization and counting
Mohit Singh and Nisheeth K Vishnoi · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2015
Earlier work this paper cites.
f f -divergence inequalities
Igal Sason and Sergio Verdú · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Earlier work this paper cites.
A general characterization of the statistical query complexity
Vitaly Feldman · 2017
Earlier work this paper cites.
Contextual decision processes with low Bellman rank are PAC-learnable
Nan Jiang, Akshay Krishnamurthy, Alekh Agarwal, John Langford, and Robert E Schapire · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
The simulator: Understanding adaptive sampling in the moderate-confidence regime
Max Simchowitz, Kevin Jamieson, and Benjamin Recht · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin · 2018
Earlier work this paper cites.
Born again neural networks
Tommaso Furlanello, Zachary Lipton, Michael Tschannen, Laurent Itti, and Anima Anandkumar · 2018
Earlier work this paper cites.
Bin Dong, Jikai Hou, Yiping Lu, and Zhihua Zhang · 2019
Earlier work this paper cites.
A closer look at deep learning heuristics: Learning rate restarts, warmup and distillation
Akhilesh Gotmare, Nitish Shirish Keskar, Caiming Xiong, and Richard Socher · 2019
Earlier work this paper cites.
Sampling can be faster than optimization
Yi-An Ma, Yuansi Chen, Chi Jin, Nicolas Flammarion, and Michael I Jordan · 2019
Earlier work this paper cites.
Computational separations between sampling and optimization
Kunal Talwar · 2019
Cited alongside, same era.
Transferring inductive biases through knowledge distillation
Samira Abnar, Mostafa Dehghani, and Willem Zuidema · 2020
Cited alongside, same era.
Towards understanding ensemble, knowledge distillation and self-distillation in deep learning
Zeyuan Allen-Zhu and Yuanzhi Li · 2020
Cited alongside, same era.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Cited alongside, same era.
OpenAI · 2023
Later among the works it cites.
Language model self-improvement by reinforcement learning contemplation
Jing-Cheng Pang, Pengyuan Wang, Kaiyuan Li, Xiong-Hui Chen, Jiacheng Xu, Zongzhang Zhang, and Yang Yu · 2023
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn · 2023
Later among the works it cites.
Language models are greedy reasoners: A systematic formal analysis of chain-of-thought
Abulhair Saparov and He He · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, Brian Fuller, Cynthia Gao, Vedanuj Goswami, Naman Goyal, Anthony Hartshorn, Saghar Hosseini, Rui Hou, Hakan Inan, Marcin Kardas, Viktor Kerkez, Madian Khabsa, Isabel Kloumann, Artem Korenev, Punit Singh Koura, Marie-Anne Lachaux, Thibaut Lavril, Jenya Lee, Diana Liskovich, Yinghai Lu, Yuning Mao, Xavier Martinet, Todor Mihaylov, Pushkar Mishra, Igor Molybog, Yixin Nie, Andrew Poulton, Jeremy Reizenstein, Rashi Rungta, Kalyan Saladi, Alan Schelten, Ruan Silva, Eric Michael Smith, Ranjan Subramanian, Xiaoqing Ellen Tan, Binh Tang, Ross Taylor, Adina Williams, Jian Xiang Kuan, Puxin Xu, Zheng Yan, Iliyan Zarov, Yuchen Zhang, Angela Fan, Melanie Kambadur, Sharan Narang, Aurelien Rodriguez, Robert Stojnic, Sergey Edunov, and Thomas Scialom · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt · 2020
Cited alongside, same era.
Bandit algorithms
Tor Lattimore and Csaba Szepesvári · 2020
Cited alongside, same era.
If beam search is the answer, what was the question?
Clara Meister, Tim Vieira, and Ryan Cotterell · 2020
Cited alongside, same era.
Self-distillation amplifies regularization in hilbert space
Hossein Mobahi, Mehrdad Farajtabar, and Peter Bartlett · 2020
Cited alongside, same era.
Learning to summarize with human feedback
Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F Christiano · 2020
Cited alongside, same era.
Amortized bayesian optimization over discrete spaces
Kevin Swersky, Yulia Rubanova, David Dohan, and Kevin Murphy · 2020
Cited alongside, same era.
Tent: Fully test-time adaptation by entropy minimization
Dequan Wang, Evan Shelhamer, Shaoteng Liu, Bruno Olshausen, and Trevor Darrell · 2020
Cited alongside, same era.
Q* approximation schemes for batch reinforcement learning: A theoretical comparison
Tengyang Xie and Nan Jiang · 2020
Cited alongside, same era.
Later among the works it cites.
The role of coverage in online reinforcement learning
Tengyang Xie, Dylan J Foster, Yu Bai, Nan Jiang, and Sham M Kakade · 2023
Later among the works it cites.
Gibbs sampling from human feedback: A provable KL-constrained framework for RLHF
Wei Xiong, Hanze Dong, Chenlu Ye, Han Zhong, Nan Jiang, and Tong Zhang · 2023
Later among the works it cites.
Principled reinforcement learning with human feedback from pairwise or k-wise comparisons
Banghua Zhu, Michael Jordan, and Jiantao Jiao · 2023
Later among the works it cites.
Phi-3 technical report: A highly capable language model locally on your phone
Marah Abdin, Jyoti Aneja, Hany Awadalla, Ahmed Awadallah, Ammar Ahmad Awan, Nguyen Bach, Amit Bahree, Arash Bakhtiari, Jianmin Bao, Harkirat Behl, et al · 2024
Closest in time.
Variational best-of-n alignment
Afra Amini, Tim Vieira, and Ryan Cotterell · 2024
Closest in time.
Scalable online exploration via coverability
Philip Amortila, Dylan J Foster, and Akshay Krishnamurthy · 2024
Closest in time.
Towards a theory of model distillation
Enric Boix-Adsera · 2024
Closest in time.
Large language monkeys: Scaling inference compute with repeated sampling
Bradley Brown, Jordan Juravsky, Ryan Ehrlich, Ronald Clark, Quoc V Le, Christopher Ré, and Azalia Mirhoseini · 2024
Closest in time.
Self-play fine-tuning converts weak language models to strong language models
Zixiang Chen, Yihe Deng, Huizhuo Yuan, Kaixuan Ji, and Quanquan Gu · 2024
Closest in time.
Retraining with predicted hard labels provably increases model accuracy
Rudrajit Das, Inderjit S Dhillon, Alessandro Epasto, Adel Javanmard, Jieming Mao, Vahab Mirrokni, Sujay Sanghavi, and Peilin Zhong · 2024
Closest in time.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al · 2024
Closest in time.
REBEL: Reinforcement learning via regressing relative rewards
Zhaolin Gao, Jonathan D Chang, Wenhao Zhan, Owen Oertell, Gokul Swamy, Kianté Brantley, Thorsten Joachims, J Andrew Bagnell, Jason D Lee, and Wen Sun · 2024
Closest in time.
BoNBoN alignment for large language models and the sweetness of best-of-n sampling
Lin Gui, Cristina Gârbacea, and Victor Veitch · 2024
Closest in time.
Audrey Huang, Wenhao Zhan, Tengyang Xie, Jason D Lee, Wen Sun, Akshay Krishnamurthy, and Dylan J Foster · 2024
Closest in time.
Chain of thought empowers transformers to solve inherently serial problems
Zhiyuan Li, Hong Liu, Denny Zhou, and Tengyu Ma · 2024
Closest in time.
Provably mitigating overoptimization in RLHF: Your SFT loss is implicitly an adversarial regularizer
Zhihan Liu, Miao Lu, Shenao Zhang, Boyi Liu, Hongyi Guo, Yingxiang Yang, Jose Blanchet, and Zhaoran Wang · 2024
Closest in time.
West-of-n: Synthetic preference generation for improved reward modeling
Alizée Pace, Jonathan Mallinson, Eric Malmi, Sebastian Krause, and Aliaksei Severyn · 2024
Closest in time.
Understanding the gains from repeated self-distillation
Divyansh Pareek, Simon S Du, and Sewoong Oh · 2024
Closest in time.
The entropy enigma: Success and failure of entropy minimization
Ori Press, Ravid Shwartz-Ziv, Yann LeCun, and Matthias Bethge · 2024
Closest in time.
Recursive introspection: Teaching language model agents how to self-improve
Yuxiao Qu, Tianjun Zhang, Naman Garg, and Aviral Kumar · 2024
Closest in time.
Bond: Aligning LLMs with Best-of-N distillation
Pier Giuseppe Sessa, Robert Dadashi, Léonard Hussenot, Johan Ferret, Nino Vieillard, Alexandre Ramé, Bobak Shariari, Sarah Perrin, Abe Friesen, Geoffrey Cideron, et al · 2024
Closest in time.
Scaling LLM test-time compute optimally can be more effective than scaling model parameters
Charlie Snell, Jaehoon Lee, Kelvin Xu, and Aviral Kumar · 2024
Closest in time.
Understanding preference fine-tuning through the lens of coverage
Yuda Song, Gokul Swamy, Aarti Singh, J Andrew Bagnell, and Wen Sun · 2024
Closest in time.
Alphazero-like tree-search can guide large language model decoding and training
Ziyu Wan, Xidong Feng, Muning Wen, Stephen Marcus McAleer, Ying Wen, Weinan Zhang, and Jun Wang · 2024
Closest in time.
Tianlu Wang, Ilia Kulikov, Olga Golovneva, Ping Yu, Weizhe Yuan, Jane Dwivedi-Yu, Richard Yuanzhe Pang, Maryam Fazel-Zarandi, Jason Weston, and Xian Li · 2024
Closest in time.
Chain-of-thought reasoning without prompting
Xuezhi Wang and Denny Zhou · 2024
Closest in time.
Exploratory preference optimization: Harnessing implicit Q*-approximation for sample-efficient RLHF
Tengyang Xie, Dylan J Foster, Akshay Krishnamurthy, Corby Rosset, Ahmed Awadallah, and Alexander Rakhlin · 2024
Closest in time.
Asymptotics of language model alignment
Joy Qiping Yang, Salman Salamatian, Ziteng Sun, Ananda Theertha Suresh, and Ahmad Beirami · 2024
Closest in time.
Tree of thoughts: Deliberate problem solving with large language models
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Tom Griffiths, Yuan Cao, and Karthik Narasimhan · 2024
Closest in time.
A theoretical analysis of Nash learning from human feedback under general KL-regularized preference
Chenlu Ye, Wei Xiong, Yuheng Zhang, Nan Jiang, and Tong Zhang · 2024
Closest in time.
Self-rewarding language models
Weizhe Yuan, Richard Yuanzhe Pang, Kyunghyun Cho, Sainbayar Sukhbaatar, Jing Xu, and Jason Weston · 2024
Closest in time.
Probabilistic inference in language models via twisted sequential monte carlo
Stephen Zhao, Rob Brekelmans, Alireza Makhzani, and Roger Baker Grosse · 2024
Closest in time.
Judging LLM-as-a-judge with MT-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al · 2024
Closest in time.