Fetching the paper…
Reading the bibliography…
We explore the idea of aligning an AI assistant by inverting a model of users' (unknown) preferences from observed interactions.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
William R Thompson · 1933
Earlier work this paper cites.
A Markovian Decision Process
Richard Bellman · 1957
Earlier work this paper cites.
On Adaptive Control Processes
Richard Bellman and Robert Kalaba · 1959
Earlier work this paper cites.
On the rationality postulates underlying the theory of cooperative games
John C Harsanyi · 1961
Earlier work this paper cites.
Evolution and the theory of games
John Maynard Smith · 1982
Earlier work this paper cites.
Maximum Likelihood Estimation of Discrete Control Processes
John Rust · 1988
Earlier work this paper cites.
Governing the commons: The evolution of institutions for collective action
Elinor Ostrom · 1990
Earlier work this paper cites.
Q Q -learning
Christopher JCH Watkins and Peter Dayan · 1992
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Learning to Achieve Goals
Leslie Pack Kaelbling · 1993
Earlier work this paper cites.
Markov Decision Processes—Discrete Stochastic Dynamic Programming
Martin L. Puterman · 1994
Earlier work this paper cites.
Reinforcement Learning: A Survey
Leslie Pack Kaelbling, Michael L Littman, and Andrew W Moore · 1996
Earlier work this paper cites.
Learning Agents for Uncertain Environments
Stuart Russell · 1998
Earlier work this paper cites.
Introduction to Reinforcement Learning
Richard S Sutton and Andrew G Barto · 1998
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David McAllester, Satinder Singh, and Yishay Mansour · 1999
Earlier work this paper cites.
Algorithms for Inverse Reinforcement Learning
Andrew Y Ng and Stuart Russell · 2000
Earlier work this paper cites.
A Bayesian Framework for Reinforcement Learning
Malcolm JA Strens · 2000
Earlier work this paper cites.
Optimal Learning: Computational Procedures for Bayes-adaptive Markov Decision Processes
Michael O’Gordon Duff · 2002
Earlier work this paper cites.
Apprenticeship Learning via Inverse Reinforcement Learning
Pieter Abbeel and Andrew Y Ng · 2004
Earlier work this paper cites.
Apprenticeship Learning using Inverse Reinforcement Learning and Gradient Methods
Gergely Neu and Csaba Szepesvári · 2007
Earlier work this paper cites.
Bayesian Inverse Reinforcement Learning
Deepak Ramachandran and Eyal Amir · 2007
Earlier work this paper cites.
A Game-Theoretic Approach to Apprenticeship Learning
Umar Syed and Robert E Schapire · 2007
Earlier work this paper cites.
Maximum Entropy Inverse Reinforcement Learning
Brian D Ziebart, Andrew Maas, J Andrew Bagnell, and Anind K Dey · 2008
Earlier work this paper cites.
Aleatory or Epistemic? Does it matter?
Armen Der Kiureghian and Ove Ditlevsen · 2009
Earlier work this paper cites.
Where Do Rewards Come From?
Satinder Singh, Richard Lewis, and Andrew Barto · 2009
Earlier work this paper cites.
Sample Complexity of Multi-Task Reinforcement Learning
Emma Brunskill and Lihong Li · 2013
Earlier work this paper cites.
Identifying social norms using coordination games: Why does dictator game sharing vary?
Erin L Krupka and Roberto A Weber · 2013
Earlier work this paper cites.
(More) Efficient Reinforcement Learning via Posterior Sampling
Ian Osband, Daniel Russo, and Benjamin Van Roy · 2013
Cited alongside, same era.
PAC Optimal Exploration in Continuous Space Markov Decision Processes
Jason Pazis and Ronald Parr · 2013
Cited alongside, same era.
Concurrent reinforcement learning from customer interactions
David Silver, Leonard Newnham, David Barker, Suzanne Weller, and Jason McFall · 2013
Cited alongside, same era.
Unwritten rules: virtual bargaining underpins social interaction, culture, and society
Jennifer B Misyak, Tigran Melkonyan, Hossam Zeitoun, and Nick Chater · 2014
Cited alongside, same era.
Bayesian Reinforcement Learning: A Survey
Mohammad Ghavamzadeh, Shie Mannor, Joelle Pineau, and Aviv Tamar · 2015
Cited alongside, same era.
Concurrent PAC RL
Zhaohan Guo and Emma Brunskill · 2015
Cited alongside, same era.
VariBAD: A Very Good Method for Bayes-Adaptive Deep RL via Meta-Learning
Luisa Zintgraf, Kyriacos Shiarlis, Maximilian Igl, Sebastian Schulze, Yarin Gal, Katja Hofmann, and Shimon Whiteson · 2020
Later among the works it cites.
On the Expressivity of Markov Reward
David Abel, Will Dabney, Anna Harutyunyan, Mark K Ho, Michael Littman, Doina Precup, and Satinder Singh · 2021
Later among the works it cites.
Inverse Reinforcement Learning in Contextual MDPs
Stav Belogolovsky, Philip Korsunsky, Shie Mannor, Chen Tessler, and Tom Zahavy · 2021
Later among the works it cites.
Systematic inequalities in language technology performance across the world’s languages
Damián Blasi, Antonios Anastasopoulos, and Graham Neubig · 2021
Later among the works it cites.
Decoupling exploration and exploitation for meta-reinforcement learning without sacrifices
Evan Z Liu, Aditi Raghunathan, Percy Liang, and Chelsea Finn · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Contextual Markov Decision Processes
Assaf Hallak, Dotan Di Castro, and Shie Mannor · 2015
Cited alongside, same era.
Reinforcement Learning Improves Behaviour from Evaluative Feedback
Michael L Littman · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Cited alongside, same era.
Universal Value Function Approximators
Tom Schaul, Daniel Horgan, Karol Gregor, and David Silver · 2015
Cited alongside, same era.
Concrete problems in ai safety
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané · 2016
Cited alongside, same era.
Asynchronous Methods for Deep Reinforcement Learning
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Cited alongside, same era.
Gati Aher, Rosa I Arriaga, and Adam Tauman Kalai · 2022
Later among the works it cites.
Deciding What to Model: Value-Equivalent Sampling for Reinforcement Learning
Dilip Arumugam and Benjamin Van Roy · 2022
Later among the works it cites.
Measuring progress on scalable oversight for large language models
Samuel R Bowman, Jeeyoon Hyun, Ethan Perez, Edwin Chen, Craig Pettit, Scott Heiner, Kamile Lukosuite, Amanda Askell, Andy Jones, Anna Chen, et al · 2022
Later among the works it cites.
Society of Agents: Regret Bounds of Concurrent Thompson Sampling
Yan Chen, Perry Dong, Qinxun Bai, Maria Dimakopoulou, Wei Xu, and Zhengyuan Zhou · 2022
Later among the works it cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Later among the works it cites.
The effects of reward misspecification: Mapping and mitigating misaligned models
Alexander Pan, Kush Bhatia, and Jacob Steinhardt · 2022
Later among the works it cites.
Task ambiguity in humans and language models
Alex Tamkin, Kunal Handa, Avash Shrestha, and Noah Goodman · 2022
Later among the works it cites.
An Explanation of In-Context Learning as Implicit Bayesian Inference
Sang Michael Xie, Aditi Raghunathan, Percy Liang, and Tengyu Ma · 2022
Later among the works it cites.
How could we make a social robot? a virtual bargaining approach
Nick Chater · 2023
Closest in time.
Strategic reasoning with language models
Kanishk Gandhi, Dorsa Sadigh, and Noah D Goodman · 2023
Closest in time.
Meta-prompt: A simple self-improving language agent
Noah Goodman · 2023
Closest in time.
A proposal for importing society’s values: Building towards coherent extrapolated volition with language models
Jan Leike · 2023
Closest in time.
Resource-rational contractualism: A triple theory of moral cognition
Sydney Levine, Nick Chater, Joshua Tenenbaum, and Fiery Cushman · 2023
Closest in time.
Training socially aligned language models in simulated human society
Ruibo Liu, Ruixin Yang, Chenyan Jia, Ge Zhang, Denny Zhou, Andrew M Dai, Diyi Yang, and Soroush Vosoughi · 2023
Closest in time.
OpenAI · 2023
Closest in time.
Generative agents: Interactive simulacra of human behavior
Joon Sung Park, Joseph C O’Brien, Carrie J Cai, Meredith Ringel Morris, Percy Liang, and Michael S Bernstein · 2023
Closest in time.
Cultural Reinforcement Learning: A Framework for Modeling Cumulative Culture on a Limited Channel
Ben Prystawski, Dilip Arumugam, and Noah D Goodman · 2023
Closest in time.
Reflexion: Language agents with verbal reinforcement learning, 2023
Noah Shinn, Federico Cassano, Beck Labash, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao · 2023
Closest in time.
Large language models as optimizers
Chengrun Yang, Xuezhi Wang, Yifeng Lu, Hanxiao Liu, Quoc V Le, Denny Zhou, and Xinyun Chen · 2023
Closest in time.
Retroformer: Retrospective large language agents with policy gradient optimization, 2023
Weiran Yao, Shelby Heinecke, Juan Carlos Niebles, Zhiwei Liu, Yihao Feng, Le Xue, Rithesh Murthy, Zeyuan Chen, Jianguo Zhang, Devansh Arpit, Ran Xu, Phil Mui, Huan Wang, Caiming Xiong, and Silvio Savarese · 2023
Closest in time.
Democratic inputs to ai, 5 2023
Wojciech Zaremba, Arka Dhar, Lama Ahmad, Tyna Eloundou, Shibani Santurkar, Sandhini Agarwal, and Jade Leung · 2023
Closest in time.