Fetching the paper…
Reading the bibliography…
In-Context Reinforcement Learning (ICRL) is a frontier paradigm to solve Reinforcement Learning (RL) problems in the foundation model era.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Ralph Allan Bradley and Milton E Terry. 1952 · 1952
Earlier work this paper cites.
The analysis of permutations
Robin L Plackett. 1975 · 1975
Earlier work this paper cites.
Probability inequalities for sums of bounded random variables
Wassily Hoeffding. 1994 · 1994
Earlier work this paper cites.
Interactively optimizing information retrieval systems as a dueling bandits problem
Yisong Yue and Thorsten Joachims. 2009 · 2009
Earlier work this paper cites.
Beat the mean bandit
Yisong Yue and Thorsten Joachims. 2011 · 2011
Earlier work this paper cites.
The k-armed dueling bandits problem
Yisong Yue, Josef Broder, Robert Kleinberg, and Thorsten Joachims. 2012 · 2012
Earlier work this paper cites.
Generic exploration and k-armed voting bandits
Tanguy Urvoy, Fabrice Clerot, Raphael Féraud, and Sami Naamane. 2013 · 2013
Earlier work this paper cites.
Reducing dueling bandits to cardinal bandits
Nir Ailon, Zohar Karnin, and Thorsten Joachims. 2014 · 2014
Earlier work this paper cites.
Contextual dueling bandits
Miroslav Dudík, Katja Hofmann, Robert E Schapire, Aleksandrs Slivkins, and Masrour Zoghi. 2015 · 2015
Earlier work this paper cites.
Regret lower bound and optimal algorithm in dueling bandit problem
Junpei Komiyama, Junya Honda, Hisashi Kashima, and Hiroshi Nakagawa. 2015 · 2015
Earlier work this paper cites.
Double thompson sampling for dueling bandits
Huasen Wu and Xin Liu. 2016 · 2016
Earlier work this paper cites.
Multi-dueling bandits with dependent arms
Yanan Sui, Vincent Zhuang, Joel W Burdick, and Yisong Yue. 2017 · 2017
Earlier work this paper cites.
A survey of preference-based reinforcement learning methods
Christian Wirth, Riad Akrour, Gerhard Neumann, and Johannes Fürnkranz. 2017 · 2017
Cited alongside, same era.
Adversarial bandits with corruptions: Regret lower bound and no-regret algorithm
Mohammad Hajiesmaili, Mohammad Sadegh Talebi, John Lui, Wing Shing Wong, et al. 2020 · 2020
Cited alongside, same era.
Learning to summarize with human feedback
Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F Christiano. 2020 · 2020
Cited alongside, same era.
Preference-based learning for exoskeleton gait optimization
Maegan Tucker, Ellen Novoseller, Claudia Kann, Yanan Sui, Yisong Yue, Joel W Burdick, and Aaron D Ames. 2020 · 2020
Cited alongside, same era.
Large language models are zero-shot reasoners
Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022 · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Inner monologue: Embodied reasoning through planning with language models
Wenlong Huang, Fei Xia, Ted Xiao, Harris Chan, Jacky Liang, Pete Florence, Andy Zeng, Jonathan Tompson, Igor Mordatch, Yevgen Chebotar, et al. 2023 · 2023
Later among the works it cites.
Exploring the sensitivity of llms’ decision-making capabilities: Insights from prompt variations and hyperparameters
Manikanta Loya, Divya Sinha, and Richard Futrell. 2023 · 2023
Later among the works it cites.
Eureka: Human-level reward design via coding large language models
Yecheng Jason Ma, William Liang, Guanzhi Wang, De-An Huang, Osbert Bastani, Dinesh Jayaraman, Yuke Zhu, Linxi Fan, and Anima Anandkumar. 2023 · 2023
Later among the works it cites.
Large language models as general pattern machines
Suvir Mirchandani, Fei Xia, Pete Florence, Brian Ichter, Danny Driess, Montserrat Gonzalez Arenas, Kanishka Rao, Dorsa Sadigh, and Andy Zeng. 2023 · 2023
Later among the works it cites.
Dueling rl: Reinforcement learning with trajectory preferences
Aadirupa Saha, Aldo Pacchiano, and Jonathan Lee. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022 · 2022
Cited alongside, same era.
Versatile dueling bandits: Best-of-both-world analyses for online learning from preferences
Aadirupa Saha and Pierre Gaillard. 2022 · 2022
Cited alongside, same era.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023 · 2023
Cited alongside, same era.
Llms-augmented contextual bandit
Ali Baheri and Cecilia Alm. 2023 · 2023
Cited alongside, same era.
Do as i can, not as i say: Grounding language in robotic affordances
Anthony Brohan, Yevgen Chebotar, Chelsea Finn, Karol Hausman, Alexander Herzog, Daniel Ho, Julian Ibarz, Alex Irpan, Eric Jang, Ryan Julian, et al. 2023 · 2023
Cited alongside, same era.
Reasoning with language model is planning with world model
Shibo Hao, Yi Gu, Haodi Ma, Joshua Hong, Zhen Wang, Daisy Zhe Wang, and Zhiting Hu · 2023
Cited alongside, same era.
In-context reinforcement learning with algorithm distillation
Michael Laskin, Luyu Wang, Junhyuk Oh, Emilio Parisotto, Stephen Spencer, Richie Steigerwald, DJ Strouse, Steven Stenberg Hansen, Angelos Filos, Ethan Brooks, et al
Cited in the paper.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023 · 2023
Later among the works it cites.
Efficient sequential decision making with large language models
Dingyang Chen, Qi Zhang, and Yinglun Zhu. 2024 · 2024
Closest in time.
Can large language models explore in-context?
Akshay Krishnamurthy, Keegan Harris, Dylan J Foster, Cyril Zhang, and Aleksandrs Slivkins · 2024
Closest in time.
Supervised pretraining can learn in-context reinforcement learning
Jonathan Lee, Annie Xie, Aldo Pacchiano, Yash Chandak, Chelsea Finn, Ofir Nachum, and Emma Brunskill. 2024 · 2024
Closest in time.
Evolve: Evaluating and optimizing llms for exploration
Allen Nie, Yi Su, Bo Chang, Jonathan N Lee, Ed H Chi, Quoc V Le, and Minmin Chen. 2024 · 2024
Closest in time.
Tree of thoughts: Deliberate problem solving with large language models
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Tom Griffiths, Yuan Cao, and Karthik Narasimhan. 2024 · 2024
Closest in time.