Online reinforcement learning in stochastic games
Original
Wei, C.-Y · 2017
Later among the works it cites.
Approximation methods for bilevel programming
Original
Ghadimi, S · 2018
Later among the works it cites.
Is multiagent deep reinforcement learning the answer or the question? a brief survey
Hernandez-Leal, P · 2018
Later among the works it cites.
Is Q-learning provably efficient?
Original
Jin, C · 2018
Later among the works it cites.
Breaking the curse of horizon: Infinite-horizon off-policy estimation
Original
Liu, Q · 2018
Later among the works it cites.
A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play
Silver, D · 2018
Later among the works it cites.
Reinforcement Learning: An Introduction
Sutton, R. S · 2018
Later among the works it cites.
Solving the rubik’s cube with deep reinforcement learning and search
Agostinelli, F · 2019
Later among the works it cites.
On the value iteration method for dynamic strong Stackelberg equilibria
Bucarey, V · 2019
Later among the works it cites.
Information-theoretic considerations in batch reinforcement learning
Chen, J · 2019
Later among the works it cites.
Off-policy deep reinforcement learning without exploration
Fujimoto, S · 2019
Later among the works it cites.
A survey and critique of multiagent deep reinforcement learning
Hernandez-Leal, P · 2019
Later among the works it cites.
Safe policy improvement with baseline bootstrapping
Laroche, R · 2019
Later among the works it cites.
Learning optimal strategies to commit to
Peng, B · 2019
Later among the works it cites.
Sample-optimal parametric Q-learning using linearly additive features
Yang, L · 2019
Later among the works it cites.
Tighter problem-dependent regret bounds in reinforcement learning without domain knowledge using value function bounds
Zanette, A · 2019
Later among the works it cites.
Model-based reinforcement learning with a generative model is minimax optimal
Agarwal, A · 2020
Later among the works it cites.
Model-based reinforcement learning with value-targeted regression
Ayoub, A · 2020
Later among the works it cites.
Provable self-play algorithms for competitive reinforcement learning
Bai, Y · 2020
Later among the works it cites.
Provably efficient exploration in policy optimization
Cai, Q · 2020
Later among the works it cites.
Computing a pessimistic Stackelberg equilibrium with multiple followers: The mixed-pure case
Coniglio, S · 2020
Later among the works it cites.
A theoretical analysis of deep Q-learning
Fan, J · 2020
Later among the works it cites.
Bandit Algorithms
Lattimore, T · 2020
Later among the works it cites.
A game theoretic framework for model based reinforcement learning
Rajeswaran, A · 2020
Later among the works it cites.
Adaptive incentive design
Ratliff, L. J · 2020
Later among the works it cites.
Solving discounted stochastic two-player games with near-optimal time and sample complexity
Sidford, A · 2020
Later among the works it cites.
Learning zero-sum simultaneous-move Markov games using function approximation and correlated equilibrium
Xie, Q · 2020
Later among the works it cites.
Reinforcement learning in feature space: Matrix bandit, kernels, and regret bound
Yang, L · 2020
Later among the works it cites.
Sample-efficient learning of Stackelberg equilibria in general-sum games
Original
Bai, Y · 2021
Closest in time.
Independent policy gradient methods for competitive reinforcement learning
Original
Daskalakis, C · 2021
Closest in time.
V-learning–a simple, efficient, decentralized algorithm for multiagent rl
Original
Jin, C · 2021
Closest in time.
Dynamic Stackelberg duopoly with sticky prices and a myopic follower
Kańska, K · 2021
Closest in time.
Provably efficient reinforcement learning in decentralized general-sum Markov games
Original
Mao, W · 2021
Closest in time.
Bridging offline reinforcement learning and imitation learning: A tale of pessimism
Original
Rashidinejad, P · 2021
Closest in time.
When can we learn general-sum Markov games with a large number of players sample-efficiently?
Original
Song, Z · 2021
Closest in time.
Provably efficient policy gradient methods for two-player zero-sum Markov games
Original
Zhao, Y · 2021
Closest in time.