Fetching the paper…
Reading the bibliography…
Offline reinforcement learning (RL) suffers from the distribution shift between the offline dataset and the online environment.
Non-Cooperative Games
John Nash. 1951 · 1951
Earlier work this paper cites.
Information and Information Stability of Random Variables and Processes
Mark S. Pinsker. 1964 · 1964
Earlier work this paper cites.
Self-Confirming Equilibrium
Drew Fudenberg and David K Levine. 1993 · 1993
Earlier work this paper cites.
A strategy of win-stay, lose-shift that outperforms tit-for-tat in the Prisoner’s Dilemma game
Martin Nowak and Karl Sigmund. 1993 · 1993
Earlier work this paper cites.
Human cooperation in the simultaneous and the alternating Prisoner’s Dilemma: Pavlov versus Generous Tit-for-Tat
C Wedekind and M Milinski. 1996 · 1996
Earlier work this paper cites.
Correlated-Q Learning. In Proceedings of the Twentieth International Conference on Machine Learning (ICML’03, Vol. 20) . 242
Amy Greenwald and Keith Hall. 2003 · 2003
Earlier work this paper cites.
Nash Q-learning for general-sum stochastic games
Junling Hu and Michael P Wellman. 2003 · 2003
Earlier work this paper cites.
Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems
Sergey Levine, Aviral Kumar, George Tucker, and Justin Fu. 2020 · 2005
Earlier work this paper cites.
Causality and batch reinforcement learning: Complementary approaches to planning in unknown domains
James Bannon, Brad Windsor, Wenbo Song, and Tao Li. 2020 · 2006
Earlier work this paper cites.
Sequence to Sequence Learning with Neural Networks. In Advances in Neural Information Processing Systems , Z. Ghahramani, M. Welling, C. Cortes, N. Lawrence, and K.Q. Weinberger (Eds.), Vol. 27. Curran Associates, Inc
Ilya Sutskever, Oriol Vinyals, and Quoc V Le. 2014 · 2014
Earlier work this paper cites.
Dynamic Games With Asymmetric Information: Common Information Based Perfect Bayesian Equilibria and Sequential Decomposition
Yi Ouyang, Hamidreza Tavafoghi, and Demosthenis Teneketzis. 2016 · 2016
Earlier work this paper cites.
Imitation Learning: A Survey of Learning Methods
Ahmed Hussein, Mohamed Medhat Gaber, Eyad Elyan, and Chrisina Jayne. 2017 · 2017
Earlier work this paper cites.
Multi-Agent Actor-Critic for Mixed Cooperative-Competitive Environments. In Advances in Neural Information Processing Systems 30 (Advances in Neural Information Processing Systems) . Curran Associates, Inc., 6379–6390
Ryan Lowe, Yi Wu, Aviv Tamar, Jean Harb, OpenAI Pieter Abbeel, and Igor Mordatch. 2017 · 2017
Earlier work this paper cites.
New Winning Strategies for the Iterated Prisoner’s Dilemma
Philippe Mathieu and Jean-Paul Delahaye. 2017 · 2017
Earlier work this paper cites.
Attention is All you Need. In Advances in Neural Information Processing Systems (NeuriPS, Vol. 30)
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Autonomous agents modelling other agents: A comprehensive survey and open problems
Stefano V. Albrecht and Peter Stone. 2018 · 2018
Earlier work this paper cites.
Speech-Transformer: A No-Recurrence Sequence-to-Sequence Model for Speech Recognition. In 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . 5884–5888
Linhao Dong, Shuang Xu, and Bo Xu. 2018 · 2018
Earlier work this paper cites.
Learning Policy Representations in Multiagent Systems. In Proceedings of the 35th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 80) , Jennifer Dy and Andreas Krause (Eds.). PMLR, 1802–1811
Aditya Grover, Maruan Al-Shedivat, Jayesh Gupta, Yuri Burda, and Harrison Edwards. 2018 · 2018
Earlier work this paper cites.
On First-Order Meta-Learning Algorithms
Alex Nichol, Joshua Achiam, and John Schulman. 2018 · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018 · 2018
Earlier work this paper cites.
Reducing overestimation bias in multi-agent domains using double centralized critics
Johannes Ackermann, Volker Gabler, Takayuki Osa, and Masashi Sugiyama. 2019 · 2019
Cited alongside, same era.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) , Jill Burstein, Christy Doran, and Thamar Solorio (Eds.). Association for Computational Linguistics, Minneapolis, Minnesota, 4171–4186
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
On Convergence Rate of Adaptive Multiscale Value Function Approximation for Reinforcement Learning. In 2019 IEEE 29th International Workshop on Machine Learning for Signal Processing (MLSP) . 1–6
Tao Li and Quanyan Zhu. 2019 · 2019
Cited alongside, same era.
Offline Multi-Agent Reinforcement Learning with Knowledge Distillation. In Advances in Neural Information Processing Systems , Vol. 35. Curran Associates, Inc., 226–237
Wei-Cheng Tseng, Tsun-Hsuan Johnson Wang, Yen-Chen Lin, and Phillip Isola. 2022 · 2022
Later among the works it cites.
Bootstrapped Transformer for Offline Reinforcement Learning. In Advances in Neural Information Processing Systems , Vol. 35. 34748–34761
Kerong Wang, Hanye Zhao, Xufang Luo, Kan Ren, Weinan Zhang, and Dongsheng Li. 2022 · 2022
Later among the works it cites.
Multi-Agent Reinforcement Learning is a Sequence Modeling Problem. In Advances in Neural Information Processing Systems , Vol. 35. Curran Associates, Inc., 16509–16521
Muning Wen, Jakub Kuba, Runji Lin, Weinan Zhang, Ying Wen, Jun Wang, and Yaodong Yang. 2022 · 2022
Later among the works it cites.
Self-Adaptive Driving in Nonstationary Environments through Conjectural Online Lookahead Adaptation
Tao Li, Haozhe Lei, and Quanyan Zhu. 2023b · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Justin Fu, Aviral Kumar, Ofir Nachum, George Tucker, and Sergey Levine. 2020 · 2020
Cited alongside, same era.
Conservative q-learning for offline reinforcement learning
Aviral Kumar, Aurick Zhou, George Tucker, and Sergey Levine. 2020 · 2020
Cited alongside, same era.
Deep Learning in Robotics: Survey on Model Structures and Training Strategies
Artúr István Károly, Péter Galambos, József Kuti, and Imre J. Rudas. 2021 · 2020
Cited alongside, same era.
Decision transformer: Reinforcement learning via sequence modeling
Lili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee, Aditya Grover, Misha Laskin, Pieter Abbeel, Aravind Srinivas, and Igor Mordatch. 2021 · 2021
Cited alongside, same era.
A minimalist approach to offline reinforcement learning
Scott Fujimoto and Shixiang Shane Gu. 2021 · 2021
Cited alongside, same era.
Generalized Decision Transformer for Offline Hindsight Information Matching. In International Conference on Learning Representations (ICLR)
Hiroki Furuta, Yutaka Matsuo, and Shixiang Shane Gu. 2021 · 2021
Cited alongside, same era.
Offline Reinforcement Learning as One Big Sequence Modeling Problem. In Advances in Neural Information Processing Systems , Vol. 34. Curran Associates, Inc., 1273–1286
Michael Janner, Qiyang Li, and Sergey Levine. 2022 · 2021
Cited alongside, same era.
Blackwell online learning for Markov decision processes. In 2021 55th Annual Conference on Information Sciences and Systems (CISS) . 1–6
Tao Li, Guanze Peng, and Quanyan Zhu. 2021 · 2021
Cited alongside, same era.
Agent Modelling under Partial Observability for Deep Reinforcement Learning. In Advances in Neural Information Processing Systems , Vol. 34. 19210–19222
Georgios Papoudakis, Filippos Christianos, and Stefano Albrecht. 2021 · 2021
Cited alongside, same era.
On the Price of Transparency: A Comparison Between Overt Persuasion and Covert Signaling. In 2023 62nd IEEE Conference on Decision and Control (CDC) . 4267–4272
Tao Li and Quanyan Zhu. 2023 · 2023
Closest in time.
Game-Theoretic Distributed Empirical Risk Minimization With Strategic Network Design
Shutian Liu, Tao Li, and Quanyan Zhu. 2023 · 2023
Closest in time.
Offline Pre-trained Multi-agent Decision Transformer
Linghui Meng, Muning Wen, Chenyang Le, Xiyun Li, Dengpeng Xing, Weinan Zhang, Ying Wen, Haifeng Zhang, Jun Wang, Yaodong Yang, and Bo Xu. 2023 · 2023
Closest in time.
Is Stochastic Mirror Descent Vulnerable to Adversarial Delay Attacks? A Traffic Assignment Resilience Study. In 2023 62nd IEEE Conference on Decision and Control (CDC) . 8328–8333
Yunian Pan, Tao Li, and Quanyan Zhu. 2023a · 2023
Closest in time.
On the Resilience of Traffic Networks under Non-Equilibrium Learning. In 2023 American Control Conference (ACC) . 3484–3489
Yunian Pan, Tao Li, and Quanyan Zhu. 2023b · 2023
Closest in time.
Automated Security Response through Online Learning with Adaptive Conjectures
Kim Hammar, Tao Li, Rolf Stadler, and Quanyan Zhu. 2024 · 2024
Closest in time.
Opponent Modeling with In-context Search. In Advances in Neural Information Processing Systems , Vol. 37. 61549–61591
Yuheng Jing, Bingyun Liu, Kai Li, Yifan Zang, Haobo Fu, Qiang Fu, Junliang Xing, and Jian Cheng. 2024 · 2024
Closest in time.
Tao Li, Zilin Bian, Haozhe Lei, Fan Zuo, Ya-Ting Yang, Quanyan Zhu, Zhenning Li, Zhibin Chen, and Kaan Ozbay. 2024b · 2024
Closest in time.
Multi-level traffic-responsive tilt camera surveillance through predictive correlated online learning
Tao Li, Zilin Bian, Haozhe Lei, Fan Zuo, Ya-Ting Yang, Quanyan Zhu, Zhenning Li, and Kaan Ozbay. 2024a · 2024
Closest in time.
Conjectural online learning with first-order beliefs in asymmetric information stochastic games. In 2024 63rd IEEE Conference on Decision and Control (CDC)
Tao Li, Kim Hammar, Rolf Stadler, and Quanyan Zhu. 2023a · 2024
Closest in time.
Meta stackelberg game: Robust federated learning against adaptive and mixed poisoning attacks
Tao Li, Henger Li, Yunian Pan, Tianyi Xu, Zizhan Zheng, and Quanyan Zhu. 2024c · 2024
Closest in time.
On the Variational Interpretation of Mirror Play in Monotone Games
Yunian Pan, Tao Li, and Quanyan Zhu. 2024 · 2024
Closest in time.
Opponent Transformer: Modeling Opponent Policies as a Sequence Problem. In Coordination and Cooperation in Multi-Agent Reinforcement Learning Workshop (RLC)
Conor Wallace, Umer Siddique, and Yongcan Cao. 2024 · 2024
Closest in time.
Xinhong Xie, Tao Li, and Quanyan Zhu. 2024 · 2024
Closest in time.
Off-policy deep reinforcement learning without exploration. In International conference on machine learning . PMLR, 2052–2062
Scott Fujimoto, David Meger, and Doina Precup. 2019 · 2062
Closest in time.