Fetching the paper…
Reading the bibliography…
We introduce the framework of performative reinforcement learning where the policy chosen by the learner affects the underlying reward and transition dynamics of the environment.
“Asynchronous methods for deep reinforcement learning”
Volodymyr Mnih et al · 1937
Earlier work this paper cites.
“A further generalization of the Kakutani fixed point theorem, with application to Nash equilibrium points”
Irving Glicksberg · 1952
Earlier work this paper cites.
“Stochastic games”
Lloyd Shapley · 1953
Earlier work this paper cites.
“Game theory, on-line prediction and boosting”
Yoav Freund and Robert Schapire · 1996
Earlier work this paper cites.
“Experts in a Markov decision process”
Eyal Even-Dar, Sham Kakade and Yishay Mansour · 2004
Earlier work this paper cites.
“Finite-Time Bounds for Fitted Value Iteration.”
Rémi Munos and Csaba Szepesvári · 2008
Earlier work this paper cites.
“Convex optimization theory”
Dimitri Bertsekas · 2009
Earlier work this paper cites.
“Error propagation for approximate policy and value iteration”
Amir-massoud Farahmand, Csaba Szepesvári and Rémi Munos · 2010
Earlier work this paper cites.
“Introduction to the non-asymptotic analysis of random matrices”
Roman Vershynin · 2010
Earlier work this paper cites.
“Market structure and equilibrium”
Heinrich Von · 2010
Earlier work this paper cites.
“Competitive Markov decision processes”
Jerzy Filar and Koos Vrieze · 2012
Earlier work this paper cites.
“Computing optimal strategies to commit to in stochastic games”
Joshua Letchford et al · 2012
Earlier work this paper cites.
“The adversarial stochastic shortest path problem with unknown transition probabilities”
Gergely Neu, Andras Gyorgy and Csaba Szepesvári · 2012
Earlier work this paper cites.
“Computing stackelberg equilibria in discounted stochastic games”
Yevgeniy Vorobeychik and Satinder Singh · 2012
Earlier work this paper cites.
“Better rates for any adversarial deterministic MDP”
Ofer Dekel and Elad Hazan · 2013
Earlier work this paper cites.
“Learning Adversary Behavior in Security Games: A PAC Model Perspective”
Arunesh Sinha, Debarun Kar and Milind Tambe · 2016
Earlier work this paper cites.
“Multi-view decision processes: the helper-ai problem”
Christos Dimitrakakis, David Parkes, Goran Radanovic and Paul Tylkin · 2017
Earlier work this paper cites.
“Cloudy with a chance of poaching: Adversary behavior modeling and forecasting with real-world poaching data”
Debarun Kar et al · 2017
Earlier work this paper cites.
“Game-theoretic modeling of human adaptation in human-robot collaboration”
Stefanos Nikolaidis, Swaprava Nath, Ariel Procaccia and Siddhartha Srinivasa · 2017
Earlier work this paper cites.
“Game-theoretic modeling of human adaptation in human-robot collaboration”
Stefanos Nikolaidis, Swaprava Nath, Ariel Procaccia and Siddhartha Srinivasa · 2017
Cited alongside, same era.
“A unified view of entropy-regularized markov decision processes”
Gergely Neu, Anders Jonsson and Vicenç Gómez · 2017
Cited alongside, same era.
“Mastering the game of go without human knowledge”
David Silver et al · 2017
Cited alongside, same era.
“Cooperating with machines”
Jacob Crandall et al · 2018
Cited alongside, same era.
“How algorithmic confounding in recommendation systems increases homogeneity and decreases utility”
Allison Chaney, Brandon Stewart and Barbara Engelhardt · 2018
Cited alongside, same era.
“Stochastic optimization for performative prediction”
Celestine Mendler-Dünner, Juan Perdomo, Tijana Zrnic and Moritz Hardt · 2020
Later among the works it cites.
“Model-free reinforcement learning for stochastic stackelberg security games”
Rajesh Mishra, Deepanshu Vasal and Sriram Vishwanath · 2020
Later among the works it cites.
“Performative prediction”
Juan Perdomo, Tijana Zrnic, Celestine Mendler-Dünner and Moritz Hardt · 2020
Later among the works it cites.
“A game theoretic framework for model based reinforcement learning”
Aravind Rajeswaran, Igor Mordatch and Vikash Kumar · 2020
Later among the works it cites.
“Learning to play sequential games versus unknown opponents”
Pier Sessa, Ilija Bogunovic, Maryam Kamgarpour and Andreas Krause · 2020
Later among the works it cites.
“Sample-efficient learning of stackelberg equilibria in general-sum games”
Yu Bai, Chi Jin, Huan Wang and Caiming Xiong · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Pratik Gajane, Ronald Ortner and Peter Auer · 2018
Cited alongside, same era.
“Superhuman AI for multiplayer poker”
Noam Brown and Tuomas Sandholm · 2019
Cited alongside, same era.
“On the utility of learning about humans for human-ai coordination”
Micah Carroll et al · 2019
Cited alongside, same era.
“A guide to deep learning in healthcare”
Andre Esteva et al · 2019
Cited alongside, same era.
“Learning to collaborate in markov decision processes”
Goran Radanovic, Rati Devidze, David Parkes and Adish Singla · 2019
Cited alongside, same era.
“Online convex optimization in adversarial markov decision processes”
Aviv Rosenberg and Yishay Mansour · 2019
Cited alongside, same era.
“Grandmaster level in StarCraft II using multi-agent reinforcement learning”
Oriol Vinyals et al · 2019
Cited alongside, same era.
Later among the works it cites.
“Reinforcement Learning in Newcomblike Environments”
James Bell, Linda Linsefors, Caspar Oesterheld and Joar Skalse · 2021
Later among the works it cites.
“A kernel-based approach to non-stationary reinforcement learning in metric spaces”
Omar Domingues et al · 2021
Later among the works it cites.
“Corruption-robust exploration in episodic reinforcement learning”
Thodoris Lykouris, Max Simchowitz, Alex Slivkins and Wen Sun · 2021
Later among the works it cites.
“Outside the echo chamber: Optimizing the performative risk”
John Miller, Juan Perdomo and Tijana Zrnic · 2021
Later among the works it cites.
“On Blame Attribution for Accountable Multi-Agent Sequential Decision Making”
Stelios Triantafyllou, Adish Singla and Goran Radanovic · 2021
Later among the works it cites.
“Non-stationary reinforcement learning without prior knowledge: An optimal black-box approach”
Chen-Yu Wei and Haipeng Luo · 2021
Later among the works it cites.
“Batch value-function approximation with only realizability”
Tengyang Xie and Nan Jiang · 2021
Later among the works it cites.
Han Zhong, Zhuoran Yang, Zhaoran Wang and Michael Jordan · 2021
Later among the works it cites.
“Multi-agent reinforcement learning: A selective overview of theories and algorithms”
Kaiqing Zhang, Zhuoran Yang and Tamer Başar · 2021
Later among the works it cites.
Peide Huang, Mengdi Xu, Fei Fang and Ding Zhao · 2022
Closest in time.
“Multiplayer Performative Prediction: Learning in Decision-Dependent Games”
Adhyyan Narang et al · 2022
Closest in time.
“Offline reinforcement learning with realizability and single-policy concentrability”
Wenhao Zhan et al · 2022
Closest in time.