Fetching the paper…
Reading the bibliography…
With the widespread adoption of Large Language Models (LLMs), the prevalence of iterative interactions among these models is anticipated to increase.
“Markov chain Monte Carlo in practice”
Walter Gilks, Sylvia Richardson and David Spiegelhalter · 1995
Earlier work this paper cites.
“The stochastic EM algorithm: estimation and asymptotic results”
Søren Nielsen · 2000
Earlier work this paper cites.
“Asymptotic statistics”
Aad Van · 2000
Earlier work this paper cites.
“Innateness and culture in the evolution of language”
Simon Kirby, Mike Dowman and Thomas Griffiths · 2007
Earlier work this paper cites.
“The Corpus of Contemporary American English.”, 2008
Mark Davies · 2008
Earlier work this paper cites.
“Using category structures to test iterated learning as a method for identifying inductive biases”
Thomas Griffiths, Brian Christian and Michael Kalish · 2008
Earlier work this paper cites.
“Cumulative cultural evolution in the laboratory: An experimental approach to the origins of structure in human language”
Simon Kirby, Hannah Cornish and Kenny Smith · 2008
Earlier work this paper cites.
“Convention: A philosophical study”
David Lewis · 2008
Earlier work this paper cites.
“Iterated learning and the cultural ratchet”
Aaron Beppu and Thomas Griffiths · 2009
Earlier work this paper cites.
“The evolution of frequency distributions: Relating regularization to inductive biases through iterated learning”
Florencia Reali and Thomas Griffiths · 2009
Earlier work this paper cites.
“The interactive evolution of human communication systems”
Nicolas Fay, Simon Garrod, Leo Roberts and Nik Swoboda · 2010
Earlier work this paper cites.
“A tutorial introduction to Bayesian models of cognitive development”
Amy Perfors, Joshua Tenenbaum, Thomas Griffiths and Fei Xu · 2011
Earlier work this paper cites.
“Probabilistic models, learning algorithms, and response variability: sampling in cognitive development”
Elizabeth Bonawitz, Stephanie Denison, Thomas Griffiths and Alison Gopnik · 2014
Earlier work this paper cites.
“Regularization in language evolution: On the joint contribution of domain-specific biases and domain-general frequency learning”
Vanessa Ferdinand, Simon Kirby and Kenny Smith · 2014
Earlier work this paper cites.
“Compression and communication in the cultural evolution of linguistic structure”
Simon Kirby, Monica Tamariz, Hannah Cornish and Kenny Smith · 2015
Earlier work this paper cites.
“When extremists win: Cultural transmission via iterated learning when populations are heterogeneous”
Danielle Navarro, Andrew Perfors, Arthur Kary, Scott Brown and Chris Donkin · 2018
Earlier work this paper cites.
“The cognitive roots of regularization in language”
Vanessa Ferdinand, Simon Kirby and Kenny Smith · 2019
Earlier work this paper cites.
“The emergence of compositional languages for numeric concepts through iterated learning in neural agents”
Shangmin Guo, Yi Ren, Serhii Havrylov, Stella Frank, Ivan Titov and Kenny Smith · 2019
Earlier work this paper cites.
“Evolving artificial sign languages in the lab: From improvised gesture to systematic sign”
Yasamin Motamedi, Marieke Schouwstra, Kenny Smith, Jennifer Culbertson and Simon Kirby · 2019
Cited alongside, same era.
“Countering language drift with seeded iterated learning”
Yuchen Lu, Soumye Singhal, Florian Strub, Aaron Courville and Olivier Pietquin · 2020
Cited alongside, same era.
“Self-distillation amplifies regularization in Hilbert space”
Hossein Mobahi, Mehrdad Farajtabar and Peter Bartlett · 2020
Cited alongside, same era.
“Compositional languages emerge in a neural iterated learning model”
Yi Ren, Shangmin Guo, Matthieu Labeau, Shay. Cohen and Simon Kirby · 2020
Cited alongside, same era.
“Iterated learning for emergent systematicity in vqa”
Ankit Vani, Max Schwarzer, Yuchen Lu, Eeshan Dhekane and Aaron Courville · 2021
Cited alongside, same era.
“Acre: Abstract causal reasoning beyond covariation”
R McCoy, Shunyu Yao, Dan Friedman, Matthew Hardy and Thomas Griffiths · 2023
Later among the works it cites.
“Improving Systematic Generalization using Iterated Learning and Simplicial Embeddings”
Yi Ren, Samuel Lavoie, Mikhail Galkin, Danica. Sutherland and Aaron Courville · 2023
Later among the works it cites.
“Model Dementia: Generated Data Makes Models Forget”
Ilia Shumailov, Zakhar Shumaylov, Yiren Zhao, Yarin Gal, Nicolas Papernot and Ross Anderson · 2023
Later among the works it cites.
“Llama 2: Open foundation and fine-tuned chat models”
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava and Shruti Bhosale · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chi Zhang, Baoxiong Jia, Mark Edmonds, Song-Chun Zhu and Yixin Zhu · 2021
Cited alongside, same era.
“Training a helpful and harmless assistant with reinforcement learning from human feedback”
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli and Tom Henighan · 2022
Cited alongside, same era.
“From improvisation to learning: How naturalness and systematicity shape language evolution”
Yasamin Motamedi, Lucie Wolters, Danielle Naegeli, Simon Kirby and Marieke Schouwstra · 2022
Cited alongside, same era.
“Training language models to follow instructions with human feedback”
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama and Alex Ray · 2022
Cited alongside, same era.
“Self-instruct: Aligning language model with self generated instructions”
Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah Smith, Daniel Khashabi and Hannaneh Hajishirzi · 2022
Cited alongside, same era.
“An Explanation of In-context Learning as Implicit Bayesian Inference”
Sang Xie, Aditi Raghunathan, Percy Liang and Tengyu Ma · 2022
Cited alongside, same era.
“Weak-to-strong generalization: Eliciting strong capabilities with weak supervision”
Collin Burns, Pavel Izmailov, Jan Kirchner, Bowen Baker, Leo Gao, Leopold Aschenbrenner, Yining Chen, Adrien Ecoffet, Manas Joglekar and Jan Leike · 2023
Cited alongside, same era.
Zhaofeng Wu, Linlu Qiu, Alexis Ross, Ekin Akyürek, Boyuan Chen, Bailin Wang, Najoung Kim, Jacob Andreas and Yoon Kim · 2023
Later among the works it cites.
“Iterative preference learning from human feedback: Bridging theory and practice for rlhf under kl-constraint”
Wei Xiong, Hanze Dong, Chenlu Ye, Ziqi Wang, Han Zhong, Heng Ji, Nan Jiang and Tong Zhang · 2023
Later among the works it cites.
“Gibbs sampling from human feedback: A provable kl-constrained framework for rlhf”
Wei Xiong, Hanze Dong, Chenlu Ye, Han Zhong, Nan Jiang and Tong Zhang · 2023
Later among the works it cites.
“Baize: An open-source chat model with parameter-efficient tuning on self-chat data”
Canwen Xu, Daya Guo, Nan Duan and Julian McAuley · 2023
Later among the works it cites.
“Self-play fine-tuning converts weak language models to strong language models”
Zixiang Chen, Yihe Deng, Huizhuo Yuan, Kaixuan Ji and Quanquan Gu · 2024
Closest in time.
“Iterated Learning Improves Compositionality in Large Vision-Language Models”
Zheng Chenhao, Zhang Jieyu, Kembhavi Aniruddha and Krishna Ranjay · 2024
Closest in time.
“Length-controlled alpacaeval: A simple way to debias automatic evaluators”
Yann Dubois, Balázs Galambosi, Percy Liang and Tatsunori Hashimoto · 2024
Closest in time.
“Direct language model alignment from online ai feedback”
Shangmin Guo, Biao Zhang, Tianlin Liu, Tianqi Liu, Misha Khalman, Felipe Llinares, Alexandre Rame, Thomas Mesnard, Yao Zhao and Bilal Piot · 2024
Closest in time.
“Phenomenal Yet Puzzling: Testing Inductive Reasoning Capabilities of Language Models with Hypothesis Refinement”
Linlu Qiu, Liwei Jiang, Ximing Lu, Melanie Sclar, Valentina Pyatkin, Chandra Bhagavatula, Bailin Wang, Yoon Kim, Yejin Choi and Nouha Dziri · 2024
Closest in time.
“Direct preference optimization: Your language model is secretly a reward model”
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher Manning, Stefano Ermon and Chelsea Finn · 2024
Closest in time.
“Learning Dynamics of LLM Finetuning”
Yi Ren and Danica Sutherland · 2024
Closest in time.
“Self-Rewarding Language Models”
Yuan Weizhe, Pang Richard Yuanzhe, Cho Kyunghyun, Sukhbaatar Sainbayar, Xu Jing and Weston Jason · 2024
Closest in time.
“Perils of Self-Feedback: Self-Bias Amplifies in Large Language Models”
Wenda Xu, Guanglei Zhu, Xuandong Zhao, Liangming Pan, Lei Li and William Wang · 2024
Closest in time.