Fetching the paper…
Reading the bibliography…
Learning dynamics, which describes how the learning of specific training examples influences the model's predictions on other examples, gives us a powerful tool for understanding the behavior of deep learning systems.
“On Exact Computation with an Infinitely Wide Neural Net”
Sanjeev Arora, Simon. Du, Wei Hu, Zhiyuan Li, Ruslan Salakhutdinov and Ruosong Wang · 1904
Earlier work this paper cites.
“On Exact Computation with an Infinitely Wide Neural Net”
Sanjeev Arora, Simon. Du, Wei Hu, Zhiyuan Li, Ruslan Salakhutdinov and Ruosong Wang · 1904
Earlier work this paper cites.
“Gradient-based learning applied to document recognition”
Yann LeCun, Léon Bottou, Yoshua Bengio and Patrick Haffner · 1998
Earlier work this paper cites.
“Gradient-based learning applied to document recognition”
Yann LeCun, Léon Bottou, Yoshua Bengio and Patrick Haffner · 1998
Earlier work this paper cites.
“Active learning literature survey”
Burr Settles · 2009
Earlier work this paper cites.
“Active learning literature survey”
Burr Settles · 2009
Earlier work this paper cites.
“Introduction to core-sets: an updated survey”, 2020
Dan Feldman · 2011
Earlier work this paper cites.
“Introduction to core-sets: an updated survey”, 2020
Dan Feldman · 2011
Earlier work this paper cites.
“Distilling the knowledge in a neural network”, 2015
Geoffrey Hinton, Oriol Vinyals and Jeff Dean · 2015
Earlier work this paper cites.
“Distilling the knowledge in a neural network”, 2015
Geoffrey Hinton, Oriol Vinyals and Jeff Dean · 2015
Earlier work this paper cites.
“Deep reinforcement learning from human preferences”
Paul Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg and Dario Amodei · 2017
Earlier work this paper cites.
“Proximal policy optimization algorithms”
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford and Oleg Klimov · 2017
Earlier work this paper cites.
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan. Gomez, Lukasz Kaiser and Illia Polosukhin · 2017
Earlier work this paper cites.
“Deep reinforcement learning from human preferences”
Paul Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg and Dario Amodei · 2017
Earlier work this paper cites.
“Proximal policy optimization algorithms”
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford and Oleg Klimov · 2017
Earlier work this paper cites.
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan. Gomez, Lukasz Kaiser and Illia Polosukhin · 2017
Earlier work this paper cites.
“Neural tangent kernel: Convergence and generalization in neural networks”
Arthur Jacot, Franck Gabriel and Clément Hongler · 2018
Earlier work this paper cites.
Tongzhou Wang, Jun-Yan Zhu, Antonio Torralba and Alexei Efros · 2018
Earlier work this paper cites.
“Neural tangent kernel: Convergence and generalization in neural networks”
Arthur Jacot, Franck Gabriel and Clément Hongler · 2018
Earlier work this paper cites.
Tongzhou Wang, Jun-Yan Zhu, Antonio Torralba and Alexei Efros · 2018
Earlier work this paper cites.
“Stiffness: A new perspective on generalization in neural networks”
Stanislav Fort, Paweł Nowak, Stanislaw Jastrzebski and Srini Narayanan · 2019
Earlier work this paper cites.
“Stiffness: A new perspective on generalization in neural networks”
Stanislav Fort, Paweł Nowak, Stanislaw Jastrzebski and Srini Narayanan · 2019
Earlier work this paper cites.
“Language GANs Falling Short”
Massimo Caccia, Lucas Caccia, William Fedus, Hugo Larochelle, Joelle Pineau and Laurent Charlin · 2020
Earlier work this paper cites.
“On the Weaknesses of Reinforcement Learning for Neural Machine Translation”
Leshem Choshen, Lior Fox, Zohar Aizenbud and Omri Abend · 2020
Earlier work this paper cites.
“The Local Elasticity of Neural Networks”
Hangfeng He and Weijie Su · 2020
Earlier work this paper cites.
“The Curious Case of Neural Text Degeneration”
Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes and Yejin Choi · 2020
Earlier work this paper cites.
“Early-learning regularization prevents memorization of noisy labels”
Sheng Liu, Jonathan Niles-Weed, Narges Razavian and Carlos Fernandez-Granda · 2020
Earlier work this paper cites.
“Estimating training data influence by tracing gradient descent”
Garima Pruthi, Frederick Liu, Satyen Kale and Mukund Sundararajan · 2020
Earlier work this paper cites.
“Compositional languages emerge in a neural iterated learning model”
Yi Ren, Shangmin Guo, Matthieu Labeau, Shay. Cohen and Simon Kirby · 2020
Earlier work this paper cites.
“Language GANs Falling Short”
Massimo Caccia, Lucas Caccia, William Fedus, Hugo Larochelle, Joelle Pineau and Laurent Charlin · 2020
Earlier work this paper cites.
“On the Weaknesses of Reinforcement Learning for Neural Machine Translation”
Leshem Choshen, Lior Fox, Zohar Aizenbud and Omri Abend · 2020
Earlier work this paper cites.
“The Local Elasticity of Neural Networks”
Hangfeng He and Weijie Su · 2020
Earlier work this paper cites.
“The Curious Case of Neural Text Degeneration”
Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes and Yejin Choi · 2020
Earlier work this paper cites.
“Early-learning regularization prevents memorization of noisy labels”
Sheng Liu, Jonathan Niles-Weed, Narges Razavian and Carlos Fernandez-Granda · 2020
Earlier work this paper cites.
“Estimating training data influence by tracing gradient descent”
Garima Pruthi, Frederick Liu, Satyen Kale and Mukund Sundararajan · 2020
Earlier work this paper cites.
“Compositional languages emerge in a neural iterated learning model”
Yi Ren, Shangmin Guo, Matthieu Labeau, Shay. Cohen and Simon Kirby · 2020
Earlier work this paper cites.
“Toward better generalization bounds with locally elastic stability”
Zhun Deng, Hangfeng He and Weijie Su · 2021
Cited alongside, same era.
“Toward better generalization bounds with locally elastic stability”
Zhun Deng, Hangfeng He and Weijie Su · 2021
Cited alongside, same era.
“Training a helpful and harmless assistant with reinforcement learning from human feedback”, 2022
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli and Tom Henighan · 2022
Cited alongside, same era.
“Fine-Tuning can Distort Pretrained Features and Underperform Out-of-Distribution”
Ananya Kumar, Aditi Raghunathan, Robbie Jones, Tengyu Ma and Percy Liang · 2022
Cited alongside, same era.
“Making Look-Ahead Active Learning Strategies Feasible with Neural Tangent Kernels”
Mohamad Mohamadi, Wonho Bae and Danica. Sutherland · 2022
“Direct language model alignment from online AI feedback”, 2024
Shangmin Guo, Biao Zhang, Tianlin Liu, Tianqi Liu, Misha Khalman, Felipe Llinares, Alexandre Rame, Thomas Mesnard, Yao Zhao and Bilal Piot · 2024
Closest in time.
“Towards Efficient and Exact Optimization of Language Model Alignment”, 2024
Haozhe Ji, Cheng Lu, Yilin Niu, Pei Ke, Hongning Wang, Jun Zhu, Jie Tang and Minlie Huang · 2024
Closest in time.
“Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive”, 2024
Arka Pal, Deep Karkhanis, Samuel Dooley, Manley Roberts, Siddartha Naidu and Colin White · 2024
Closest in time.
“Emergence of Hidden Capabilities: Exploring Learning Dynamics in Concept Space”
Core Park, Maya Okawa, Andrew Lee, Ekdeep Lubana and Hidenori Tanaka · 2024
Closest in time.
“From r r to Q ∗ Q^{*} : Your Language Model is Secretly a Q-Function”, 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Training language models to follow instructions with human feedback”
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama and Alex Ray · 2022
Cited alongside, same era.
“Better Supervisory Signals by Observing Learning Paths”
Yi Ren, Shangmin Guo and Danica. Sutherland · 2022
Cited alongside, same era.
“Finetuned Language Models are Zero-Shot Learners”
Jason Wei, Maarten Bosma, Vincent Zhao, Kelvin Guu, Adams Yu, Brian Lester, Nan Du, Andrew. Dai and Quoc Le · 2022
Cited alongside, same era.
“Training a helpful and harmless assistant with reinforcement learning from human feedback”, 2022
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli and Tom Henighan · 2022
Cited alongside, same era.
“Fine-Tuning can Distort Pretrained Features and Underperform Out-of-Distribution”
Ananya Kumar, Aditi Raghunathan, Robbie Jones, Tengyu Ma and Percy Liang · 2022
Cited alongside, same era.
“Making Look-Ahead Active Learning Strategies Feasible with Neural Tangent Kernels”
Mohamad Mohamadi, Wonho Bae and Danica. Sutherland · 2022
Cited alongside, same era.
“Training language models to follow instructions with human feedback”
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama and Alex Ray · 2022
Cited alongside, same era.
Rafael Rafailov, Joey Hejna, Ryan Park and Chelsea Finn · 2024
Closest in time.
“Bias Amplification in Language Model Evolution: An Iterated Learning Perspective”
Yi Ren, Shangmin Guo, Linlu Qiu, Bailin Wang and Danica Sutherland · 2024
Closest in time.
“Understanding Simplicity Bias towards Compositional Mappings via Learning Dynamics”
Yi Ren and Danica. Sutherland · 2024
Closest in time.
“Preference Fine-Tuning of LLMs Should Leverage Suboptimal, On-Policy Data”, 2024
Fahim Tajwar, Anikait Singh, Archit Sharma, Rafael Rafailov, Jeff Schneider, Tengyang Xie, Stefano Ermon, Chelsea Finn and Aviral Kumar · 2024
Closest in time.
“Understanding the performance gap between online and offline alignment algorithms”
Yunhao Tang, Daniel Guo, Zeyu Zheng, Daniele Calandriello, Yuan Cao, Eugene Tarassov, Rémi Munos, BernardoÁvila Pires, Michal Valko and Yong Cheng · 2024
Closest in time.
“Self-play preference optimization for language model alignment”, 2024
Yue Wu, Zhiqing Sun, Huizhuo Yuan, Kaixuan Ji, Yiming Yang and Quanquan Gu · 2024
Closest in time.
“Less: Selecting influential data for targeted instruction tuning”, 2024
Mengzhou Xia, Sadhika Malladi, Suchin Gururangan, Sanjeev Arora and Danqi Chen · 2024
Closest in time.
“Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-constraint”
Wei Xiong, Hanze Dong, Chenlu Ye, Ziqi Wang, Han Zhong, Heng Ji, Nan Jiang and Tong Zhang · 2024
Closest in time.
“Self-Distillation Bridges Distribution Gap in Language Model Fine-Tuning”, 2024
Zhaorui Yang, Qian Liu, Tianyu Pang, Han Wang, Haozhe Feng, Minfeng Zhu and Wei Chen · 2024
Closest in time.
“Self-rewarding language models”, 2024
Weizhe Yuan, Richard Pang, Kyunghyun Cho, Sainbayar Sukhbaatar, Jing Xu and Jason Weston · 2024
Closest in time.
“Negative preference optimization: From catastrophic collapse to effective unlearning”
Ruiqi Zhang, Licong Lin, Yu Bai and Song Mei · 2024
Closest in time.
“A general theoretical paradigm to understand learning from human preferences”
Mohammad Azar, Zhaohan Guo, Bilal Piot, Remi Munos, Mark Rowland, Michal Valko and Daniele Calandriello · 2024
Closest in time.
“Self-play fine-tuning converts weak language models to strong language models”, 2024
Zixiang Chen, Yihe Deng, Huizhuo Yuan, Kaixuan Ji and Quanquan Gu · 2024
Closest in time.
“Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?”, 2024
Zorik Gekhman, Gal Yona, Roee Aharoni, Matan Eyal, Amir Feder, Roi Reichart and Jonathan Herzig · 2024
Closest in time.
“lpNTK: Better Generalisation with Less Data via Sample Interaction During Learning”
Shangmin Guo, Yi Ren, Stefano Albrecht and Kenny Smith · 2024
Closest in time.
“Direct language model alignment from online AI feedback”, 2024
Shangmin Guo, Biao Zhang, Tianlin Liu, Tianqi Liu, Misha Khalman, Felipe Llinares, Alexandre Rame, Thomas Mesnard, Yao Zhao and Bilal Piot · 2024
Closest in time.
“Towards Efficient and Exact Optimization of Language Model Alignment”, 2024
Haozhe Ji, Cheng Lu, Yilin Niu, Pei Ke, Hongning Wang, Jun Zhu, Jie Tang and Minlie Huang · 2024
Closest in time.
“Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive”, 2024
Arka Pal, Deep Karkhanis, Samuel Dooley, Manley Roberts, Siddartha Naidu and Colin White · 2024
Closest in time.
“Emergence of Hidden Capabilities: Exploring Learning Dynamics in Concept Space”
Core Park, Maya Okawa, Andrew Lee, Ekdeep Lubana and Hidenori Tanaka · 2024
Closest in time.
“From r r to Q ∗ Q^{*} : Your Language Model is Secretly a Q-Function”, 2024
Rafael Rafailov, Joey Hejna, Ryan Park and Chelsea Finn · 2024
Closest in time.
“Bias Amplification in Language Model Evolution: An Iterated Learning Perspective”
Yi Ren, Shangmin Guo, Linlu Qiu, Bailin Wang and Danica Sutherland · 2024
Closest in time.
“Understanding Simplicity Bias towards Compositional Mappings via Learning Dynamics”
Yi Ren and Danica. Sutherland · 2024
Closest in time.
“Preference Fine-Tuning of LLMs Should Leverage Suboptimal, On-Policy Data”, 2024
Fahim Tajwar, Anikait Singh, Archit Sharma, Rafael Rafailov, Jeff Schneider, Tengyang Xie, Stefano Ermon, Chelsea Finn and Aviral Kumar · 2024
Closest in time.
“Understanding the performance gap between online and offline alignment algorithms”
Yunhao Tang, Daniel Guo, Zeyu Zheng, Daniele Calandriello, Yuan Cao, Eugene Tarassov, Rémi Munos, BernardoÁvila Pires, Michal Valko and Yong Cheng · 2024
Closest in time.
“Self-play preference optimization for language model alignment”, 2024
Yue Wu, Zhiqing Sun, Huizhuo Yuan, Kaixuan Ji, Yiming Yang and Quanquan Gu · 2024
Closest in time.
“Less: Selecting influential data for targeted instruction tuning”, 2024
Mengzhou Xia, Sadhika Malladi, Suchin Gururangan, Sanjeev Arora and Danqi Chen · 2024
Closest in time.
“Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-constraint”
Wei Xiong, Hanze Dong, Chenlu Ye, Ziqi Wang, Han Zhong, Heng Ji, Nan Jiang and Tong Zhang · 2024
Closest in time.
“Self-Distillation Bridges Distribution Gap in Language Model Fine-Tuning”, 2024
Zhaorui Yang, Qian Liu, Tianyu Pang, Han Wang, Haozhe Feng, Minfeng Zhu and Wei Chen · 2024
Closest in time.
“Self-rewarding language models”, 2024
Weizhe Yuan, Richard Pang, Kyunghyun Cho, Sainbayar Sukhbaatar, Jing Xu and Jason Weston · 2024
Closest in time.
“Negative preference optimization: From catastrophic collapse to effective unlearning”
Ruiqi Zhang, Licong Lin, Yu Bai and Song Mei · 2024
Closest in time.
“Unintentional Unalignment: Likelihood Displacement in Direct Preference Optimization”
Noam Razin, Sadhika Malladi, Adithya Bhaskar, Danqi Chen, Sanjeev Arora and Boris Hanin · 2025
Closest in time.
“Unintentional Unalignment: Likelihood Displacement in Direct Preference Optimization”
Noam Razin, Sadhika Malladi, Adithya Bhaskar, Danqi Chen, Sanjeev Arora and Boris Hanin · 2025
Closest in time.