Fetching the paper…
Reading the bibliography…
The difficulty of appropriately assigning credit is particularly heightened in cooperative MARL with sparse reward, due to the concurrent time and structural scales involved.
Td-gammon, a self-teaching backgammon program, achieves master-level play
Gerald Tesauro · 1994
Earlier work this paper cites.
Rope: Role oriented programming environment for multiagent systems
Michael Becht, Thorsten Gurzki, Jürgen Klarmann, and Matthias Muscholl · 1999
Earlier work this paper cites.
Learning to cooperate via policy search
Leonid Peshkin, Kee-Eung Kim, Nicolas Meuleau, and Leslie Pack Kaelbling · 2000
Earlier work this paper cites.
The complexity of decentralized control of markov decision processes
Daniel S Bernstein, Robert Givan, Neil Immerman, and Shlomo Zilberstein · 2002
Earlier work this paper cites.
Unifying temporal and structural credit assignment problems
Adrian K Agogino and Kagan Tumer · 2004
Earlier work this paper cites.
Human instruction-following with deep reinforcement learning via transfer-learning from text
Felix Hill, Sona Mokra, Nathaniel Wong, and Tim Harley · 2005
Earlier work this paper cites.
Learning to win by reading manuals in a monte-carlo framework
SRK Branavan, David Silver, and Regina Barzilay · 2012
Earlier work this paper cites.
Representation learning: A review and new perspectives
Yoshua Bengio, Aaron Courville, and Pascal Vincent · 2013
Earlier work this paper cites.
Potential-based difference rewards for multiagent reinforcement learning
Sam Devlin, Logan Yliniemi, Daniel Kudenko, and Kagan Tumer · 2014
Earlier work this paper cites.
Unifying count-based exploration and intrinsic motivation
Marc Bellemare, Sriram Srinivasan, Georg Ostrovski, Tom Schaul, David Saxton, and Remi Munos · 2016
Earlier work this paper cites.
Hierarchical deep reinforcement learning: Integrating temporal abstraction and intrinsic motivation
Tejas D Kulkarni, Karthik Narasimhan, Ardavan Saeedi, and Josh Tenenbaum · 2016
Earlier work this paper cites.
Population based training of neural networks
Max Jaderberg, Valentin Dalibard, Simon Osindero, Wojciech M Czarnecki, Jeff Donahue, Ali Razavi, Oriol Vinyals, Tim Green, Iain Dunning, Karen Simonyan, et al · 2017
Earlier work this paper cites.
Curiosity-driven exploration by self-supervised prediction
Deepak Pathak, Pulkit Agrawal, Alexei A Efros, and Trevor Darrell · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Counterfactual multi-agent policy gradients
Jakob Foerster, Gregory Farquhar, Triantafyllos Afouras, Nantas Nardelli, and Shimon Whiteson · 2018
Earlier work this paper cites.
Representation learning for grounded spatial reasoning
Michael Janner, Karthik Narasimhan, and Regina Barzilay · 2018
Earlier work this paper cites.
Role-based modeling for designing agent behavior in self-organizing multi-agent systems
Kemas M Lhaksmana, Yohei Murakami, and Toru Ishida · 2018
Earlier work this paper cites.
Role-based automatic programming framework for interworking a drone and wireless sensor networks
Hong Min, Jinman Jung, Seoyeon Kim, Bongjae Kim, and Junyoung Heo · 2018
Earlier work this paper cites.
Grounding language for transfer in deep reinforcement learning
Karthik Narasimhan, Regina Barzilay, and Tommi Jaakkola · 2018
Earlier work this paper cites.
Credit assignment for collective multiagent rl with global rewards
Duc Thien Nguyen, Akshat Kumar, and Hoong Chuin Lau · 2018
Earlier work this paper cites.
Value-decomposition networks for cooperative multi-agent learning based on team reward
Peter Sunehag, Guy Lever, Audrunas Gruslys, Wojciech Marian Czarnecki, Vinicius Zambaldi, Max Jaderberg, Marc Lanctot, Nicolas Sonnerat, Joel Z Leibo, Karl Tuyls, et al · 2018
Earlier work this paper cites.
Hierarchical deep multiagent reinforcement learning with temporal abstraction
Hongyao Tang, Jianye Hao, Tangjie Lv, Yingfeng Chen, Zongzhang Zhang, Hangtian Jia, Chunxu Ren, Yan Zheng, Zhaopeng Meng, Changjie Fan, et al · 2018
Earlier work this paper cites.
Exploration by random network distillation
Yuri Burda, Harrison Edwards, Amos Storkey, and Oleg Klimov · 2019
Earlier work this paper cites.
On the utility of learning about humans for human-ai coordination
Micah Carroll, Rohin Shah, Mark K Ho, Tom Griffiths, Sanjit Seshia, Pieter Abbeel, and Anca Dragan · 2019
Earlier work this paper cites.
Meta-learning language-guided policy learning
John D Co-Reyes, Abhishek Gupta, Suvansh Sanjeev, Nick Altieri, John DeNero, Pieter Abbeel, and Sergey Levine · 2019
Earlier work this paper cites.
Liir: Learning individual intrinsic reward in multi-agent reinforcement learning
Yali Du, Lei Han, Meng Fang, Ji Liu, Tianhong Dai, and Dacheng Tao · 2019
Earlier work this paper cites.
A survey and critique of multiagent deep reinforcement learning
Pablo Hernandez-Leal, Bilal Kartal, and Matthew E Taylor · 2019
Earlier work this paper cites.
Hierarchical decision making by generating and following natural language instructions
Hengyuan Hu, Denis Yarats, Qucheng Gong, Yuandong Tian, and Mike Lewis · 2019
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin Ming-Wei Chang Kenton and Lee Kristina Toutanova · 2019
Earlier work this paper cites.
Maven: Multi-agent variational exploration
Anuj Mahajan, Tabish Rashid, Mikayel Samvelyan, and Shimon Whiteson · 2019
Earlier work this paper cites.
The starcraft multi-agent challenge
Mikayel Samvelyan, Tabish Rashid, Christian Schroeder de Witt, Gregory Farquhar, Nantas Nardelli, Tim GJ Rudner, Chia-Man Hung, Philip HS Torr, Jakob Foerster, and Shimon Whiteson · 2019
Earlier work this paper cites.
Qtran: Learning to factorize with transformation for cooperative multi-agent reinforcement learning
Kyunghwan Son, Daewoo Kim, Wan Ju Kang, David Earl Hostallero, and Yung Yi · 2019
Earlier work this paper cites.
Reinforced cross-modal matching and self-supervised imitation learning for vision-language navigation
Xin Wang, Qiuyuan Huang, Asli Celikyilmaz, Jianfeng Gao, Dinghan Shen, Yuan-Fang Wang, William Yang Wang, and Lei Zhang · 2019
Earlier work this paper cites.
Learning to map natural language instructions to physical quadcopter control using simulated flight
Valts Blukis, Yannick Terme, Eyvind Niklasson, Ross A Knepper, and Yoav Artzi · 2020
Earlier work this paper cites.
Deep coordination graphs
Wendelin Böhmer, Vitaly Kurin, and Shimon Whiteson · 2020
Cited alongside, same era.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Cited alongside, same era.
Investigating partner diversification methods in cooperative multi-agent deep reinforcement learning
Rujikorn Charakorn, Poramate Manoonpong, and Nat Dilokthanakul · 2020
Cited alongside, same era.
Towards defining role models in advanced systems engineering
Eva-Maria Grote, Stefan Achilles Pfeifer, Daniel Röltgen, Arno Kühn, and Roman Dumitrescu · 2020
Cited alongside, same era.
The nethack learning environment
Heinrich Küttler, Nantas Nardelli, Alexander Miller, Roberta Raileanu, Marco Selvatici, Edward Grefenstette, and Tim Rocktäschel · 2020
Cited alongside, same era.
Reward design in cooperative multi-agent reinforcement learning for packet routing
Byol-explore: Exploration by bootstrapped prediction
Zhaohan Guo, Shantanu Thakoor, Miruna Pîslar, Bernardo Avila Pires, Florent Altché, Corentin Tallec, Alaa Saade, Daniele Calandriello, Jean-Bastien Grill, Yunhao Tang, et al · 2022
Later among the works it cites.
Foundation models for semantic novelty in reinforcement learning
Tarun Gupta, Peter Karkus, Tong Che, Danfei Xu, and Marco Pavone · 2022
Later among the works it cites.
Maser: Multi-agent reinforcement learning with subgoals generated from experience replay buffer
Jeewon Jeon, Woojun Kim, Whiyoung Jung, and Youngchul Sung · 2022
Later among the works it cites.
Vima: General robot manipulation with multimodal prompts
Yunfan Jiang, Agrim Gupta, Zichen Zhang, Guanzhi Wang, Yongqiang Dou, Yanjun Chen, Li Fei-Fei, Anima Anandkumar, Yuke Zhu, and Linxi Fan · 2022
Later among the works it cites.
Non-linear coordination graphs
Yipeng Kang, Tonghan Wang, Qianlan Yang, Xiaoran Wu, and Chongjie Zhang · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hangyu Mao, Zhibo Gong, and Zhen Xiao · 2020
Cited alongside, same era.
Robots that use language
Stefanie Tellex, Nakul Gopalan, Hadas Kress-Gazit, and Cynthia Matuszek · 2020
Cited alongside, same era.
Roma: multi-agent reinforcement learning with emergent roles
Tonghan Wang, Heng Dong, Victor Lesser, and Chongjie Zhang · 2020
Cited alongside, same era.
Influence-based multi-agent exploration
Tonghan Wang*, Jianhao Wang*, Yi Wu, and Chongjie Zhang · 2020
Cited alongside, same era.
Macro-action-based deep multi-agent reinforcement learning
Yuchen Xiao, Joshua Hoffman, and Christopher Amato · 2020
Cited alongside, same era.
Rtfm: Generalising to new environment dynamics via reading
Victor Zhong, Tim Rocktäschel, and Edward Grefenstette · 2020
Cited alongside, same era.
Learning implicit credit assignment for cooperative multi-agent reinforcement learning
Meng Zhou, Ziyu Liu, Pengwei Sui, Yixuan Li, and Yuk Ying Chung · 2020
Cited alongside, same era.
Dung Nguyen, Phuoc Nguyen, Svetha Venkatesh, and Truyen Tran · 2022
Later among the works it cites.
Can wikipedia help offline reinforcement learning?
Machel Reid, Yutaro Yamada, and Shixiang Shane Gu · 2022
Later among the works it cites.
Self-organized group for cooperative multi-agent reinforcement learning
Jianzhun Shao, Zhiqiang Lou, Hongchang Zhang, Yuhang Jiang, Shuncheng He, and Xiangyang Ji · 2022
Later among the works it cites.
Semantic exploration from language abstractions and pretrained representations
Allison Tam, Neil Rabinowitz, Andrew Lampinen, Nicholas A Roy, Stephanie Chan, DJ Strouse, Jane Wang, Andrea Banino, and Felix Hill · 2022
Later among the works it cites.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed H Chi, Quoc V Le, Denny Zhou, et al · 2022
Later among the works it cites.
Asynchronous actor-critic for multi-agent reinforcement learning
Yuchen Xiao, Weihao Tan, and Christopher Amato · 2022
Later among the works it cites.
Grounded reinforcement learning: Learning to win the game under human commands
Shusheng Xu, Huaijie Wang, and Yi Wu · 2022
Later among the works it cites.
The surprising effectiveness of ppo in cooperative multi-agent games
Chao Yu, Akash Velu, Eugene Vinitsky, Jiaxuan Gao, Yu Wang, Alexandre Bayen, and Yi Wu · 2022
Later among the works it cites.
Do as i can, not as i say: Grounding language in robotic affordances
Anthony Brohan, Yevgen Chebotar, Chelsea Finn, Karol Hausman, Alexander Herzog, Daniel Ho, Julian Ibarz, Alex Irpan, Eric Jang, Ryan Julian, et al · 2023
Closest in time.
Grounding large language models in interactive environments with online reinforcement learning
Thomas Carta, Clément Romac, Thomas Wolf, Sylvain Lamprier, Olivier Sigaud, and Pierre-Yves Oudeyer · 2023
Closest in time.
Batch prompting: Efficient inference with large language model apis
Zhoujun Cheng, Jungo Kasai, and Tao Yu · 2023
Closest in time.
Collaborating with language models for embodied reasoning
Ishita Dasgupta, Christine Kaeser-Chen, Kenneth Marino, Arun Ahuja, Sheila Babayan, Felix Hill, and Rob Fergus · 2023
Closest in time.
Entity divider with language grounding in multi-agent reinforcement learning
Ziluo Ding, Wanpeng Zhang, Junpeng Yue, Xiangjun Wang, Tiejun Huang, and Zongqing Lu · 2023
Closest in time.
Palm-e: An embodied multimodal language model
Danny Driess, Fei Xia, Mehdi SM Sajjadi, Corey Lynch, Aakanksha Chowdhery, Brian Ichter, Ayzaan Wahid, Jonathan Tompson, Quan Vuong, Tianhe Yu, et al · 2023
Closest in time.
Guiding pretraining in reinforcement learning with large language models
Yuqing Du, Olivia Watkins, Zihan Wang, Cédric Colas, Trevor Darrell, Pieter Abbeel, Abhishek Gupta, and Jacob Andreas · 2023
Closest in time.
Exploration in deep reinforcement learning: From single-agent to multiagent domain
Jianye Hao, Tianpei Yang, Hongyao Tang, Chenjia Bai, Jinyi Liu, Zhaopeng Meng, Peng Liu, and Zhen Wang · 2023
Closest in time.
Learning optimal” pigovian tax” in sequential social dilemmas
Yun Hua, Shang Gao, Wenhao Li, Bo Jin, Xiangfeng Wang, and Hongyuan Zha · 2023
Closest in time.
Reward design with language models
Minae Kwon, Sang Michael Xie, Kalesha Bullard, and Dorsa Sadigh · 2023
Closest in time.
Self-refine: Iterative refinement with self-feedback
Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, et al · 2023
Closest in time.
Reflexion: an autonomous agent with dynamic memory and self-reflection
Noah Shinn, Beck Labash, and Ashwin Gopinath · 2023
Closest in time.
Neeraj Varshney, Wenlin Yao, Hongming Zhang, Jianshu Chen, and Dong Yu · 2023
Closest in time.
Zihao Wang, Shaofei Cai, Anji Liu, Xiaojian Ma, and Yitao Liang · 2023
Closest in time.
Spring: Gpt-4 out-performs rl algorithms by studying papers and reasoning
Yue Wu, So Yeon Min, Shrimai Prabhumoye, Yonatan Bisk, Ruslan Salakhutdinov, Amos Azaria, Tom Mitchell, and Yuanzhi Li · 2023
Closest in time.
Harnessing the power of llms in practice: A survey on chatgpt and beyond
Jingfeng Yang, Hongye Jin, Ruixiang Tang, Xiaotian Han, Qizhang Feng, Haoming Jiang, Bing Yin, and Xia Hu · 2023
Closest in time.
React: Synergizing reasoning and acting in language models
Shunyu Yao, Jeffrey Zhao, Dian Yu, Izhak Shafran, Karthik R Narasimhan, and Yuan Cao · 2023
Closest in time.
Language to rewards for robotic skill synthesis
Wenhao Yu, Nimrod Gileadi, Chuyuan Fu, Sean Kirmani, Kuang-Huei Lee, Montse Gonzalez Arenas, Hao-Tien Lewis Chiang, Tom Erez, Leonard Hasenclever, Jan Humplik, et al · 2023
Closest in time.
Plan4mc: Skill reinforcement learning and planning for open-world minecraft tasks
Haoqi Yuan, Chi Zhang, Hongcheng Wang, Feiyang Xie, Penglin Cai, Hao Dong, and Zongqing Lu · 2023
Closest in time.
Xizhou Zhu, Yuntao Chen, Hao Tian, Chenxin Tao, Weijie Su, Chenyu Yang, Gao Huang, Bin Li, Lewei Lu, Xiaogang Wang, et al · 2023
Closest in time.
Introducing ChatGPT
OpenAI · 2026
Closest in time.