Fetching the paper…
Reading the bibliography…
Game-playing agents like AlphaGo have achieved superhuman performance through self-play, which is theoretically guaranteed to yield optimal policies in competitive games.
Strategic information transmission
Vincent P Crawford and Joel Sobel · 1982
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Markov games as a framework for multi-agent reinforcement learning
Michael L Littman · 1994
Earlier work this paper cites.
Modeling expert effects and common ground using Questions Under Discussion
Alex Djalali, David Clausen, Sven Lauer, Karl Schultz, and Christopher Potts · 2011
Earlier work this paper cites.
Goal-driven answers in the Cards dialogue corpus
Christopher Potts · 2012
Earlier work this paper cites.
Toward natural turn-taking in a virtual human negotiation agent
David DeVault, Johnathan Mell, and Jonathan Gratch · 2015
Earlier work this paper cites.
Fictitious self-play in extensive-form games
Johannes Heinrich, Marc Lanctot, and David Silver · 2015
Earlier work this paper cites.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Earlier work this paper cites.
Mastering the game of Go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Earlier work this paper cites.
Natural language does not emerge ‘naturally’ in multi-agent dialog
Satwik Kottur, José Moura, Stefan Lee, and Dhruv Batra · 2017
Earlier work this paper cites.
Deal or No Deal? End-to-end learning of negotiation dialogues
Mike Lewis, Denis Yarats, Yann Dauphin, Devi Parikh, and Dhruv Batra · 2017
Earlier work this paper cites.
Mastering the game of Go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, et al · 2017
Earlier work this paper cites.
Superhuman AI for heads-up no-limit poker: Libratus beats top professionals
Noam Brown and Tuomas Sandholm · 2018
Earlier work this paper cites.
Emergent communication through negotiation
Kris Cao, Angeliki Lazaridou, Marc Lanctot, Joel Z Leibo, Karl Tuyls, and Stephen Clark · 2018
Earlier work this paper cites.
Decoupling strategy and generation in negotiation dialogues
He He, Derek Chen, Anusha Balakrishnan, and Percy Liang · 2018
Earlier work this paper cites.
A general reinforcement learning algorithm that masters chess, shogi, and go through self-play
David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, Timothy Lillicrap, Karen Simonyan, and Demis Hassabis · 2018
Earlier work this paper cites.
The Hanabi challenge: A new frontier for AI research
Nolan Bard, Jakob N. Foerster, Sarath Chandar, Neil Burch, Marc Lanctot, H. Francis Song, Emilio Parisotto, Vincent Dumoulin, Subhodeep Moitra, Edward Hughes, Iain Dunning, Shibl Mourad, Hugo Larochelle, Marc G. Bellemare, and Michael Bowling · 2019
Earlier work this paper cites.
On the utility of learning about humans for human-AI coordination
Micah Carroll, Rohin Shah, Mark K Ho, Tom Griffiths, Sanjit Seshia, Pieter Abbeel, and Anca Dragan · 2019
Earlier work this paper cites.
Way off-policy batch deep reinforcement learning of implicit human preferences in dialog
Natasha Jaques, Asma Ghandeharioun, Judy Hanwen Shen, Craig Ferguson, Agata Lapedriza, Noah Jones, Shixiang Gu, and Rosalind Picard · 2019
Earlier work this paper cites.
On the interaction between supervision and self-play in emergent communication
Ryan Lowe, Abhinav Gupta, Jakob Foerster, Douwe Kiela, and Joelle Pineau · 2019
Cited alongside, same era.
Executing instructions in situated collaborative interactions
Alane Suhr, Claudia Yan, Jack Schluger, Stanley Yu, Hadi Khader, Marwa Mouallem, Iris Zhang, and Yoav Artzi · 2019
Cited alongside, same era.
Emergent compositionality in signaling games
Nicholas Tomlin and Ellie Pavlick · 2019
Cited alongside, same era.
A natural language corpus of common grounding under continuous and partially-observable context
Takuma Udagawa and Akiko Aizawa · 2019
Cited alongside, same era.
Provable self-play algorithms for competitive reinforcement learning
Yu Bai and Chi Jin · 2020
Cited alongside, same era.
Near-optimal reinforcement learning with self-play
Yu Bai, Chi Jin, and Tiancheng Yu · 2020
Pragmatics in language grounding: Phenomena, tasks, and modeling approaches
Daniel Fried, Nicholas Tomlin, Jennifer Hu, Roma Patel, and Aida Nematzadeh · 2023
Later among the works it cites.
Improving language model negotiation with self-play and in-context learning from AI feedback
Yao Fu, Hao Peng, Tushar Khot, and Mirella Lapata · 2023
Later among the works it cites.
Strategic reasoning with language models
Kanishk Gandhi, Dorsa Sadigh, and Noah D Goodman · 2023
Later among the works it cites.
Incorporating worker perspectives into MTurk annotation practices for NLP
Olivia Huang, Eve Fleisig, and Dan Klein · 2023
Later among the works it cites.
Beyond static datasets: A deep interaction approach to LLM evaluation
Jiatong Li, Rui Li, and Qi Liu · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Bail: Best-action imitation learning for batch deep reinforcement learning
Xinyue Chen, Zijian Zhou, Zheng Wang, Che Wang, Yanqiu Wu, and Keith Ross · 2020
Cited alongside, same era.
CaSiNo: A corpus of campsite negotiation dialogues for automatic negotiation systems
Kushal Chawla, Jaysa Ramirez, Rene Clever, Gale Lucas, Jonathan May, and Jonathan Gratch · 2021
Cited alongside, same era.
Decision transformer: Reinforcement learning via sequence modeling
Lili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee, Aditya Grover, Misha Laskin, Pieter Abbeel, Aravind Srinivas, and Igor Mordatch · 2021
Cited alongside, same era.
Emergent communication under competition
Michael Noukhovitch, Travis LaCroix, Angeliki Lazaridou, and Aaron Courville · 2021
Cited alongside, same era.
AI-generated characters for supporting personalized learning and well-being
Pat Pataranutaporn, Valdemar Danry, Joanne Leong, Parinya Punpongsanon, Dan Novy, Pattie Maes, and Misha Sra · 2021
Cited alongside, same era.
Collaborating with humans without human data
DJ Strouse, Kevin McKee, Matt Botvinick, Edward Hughes, and Richard Everett · 2021
Cited alongside, same era.
Later among the works it cites.
GPT-4 technical report
OpenAI · 2023
Later among the works it cites.
Generative agents: Interactive simulacra of human behavior
Joon Sung Park, Joseph O’Brien, Carrie Jun Cai, Meredith Ringel Morris, Percy Liang, and Michael S. Bernstein · 2023
Later among the works it cites.
Gameeval: Evaluating LLMs on conversational games
Dan Qiao, Chenfei Wu, Yaobo Liang, Juntao Li, and Nan Duan · 2023
Later among the works it cites.
Reflexion: language agents with verbal reinforcement learning
Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao · 2023
Later among the works it cites.
Tree of thoughts: Deliberate problem solving with large language models
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Tom Griffiths, Yuan Cao, and Karthik Narasimhan · 2023
Later among the works it cites.
States as strings as strategies: Steering language models with game-theoretic solvers
Ian Gemp, Yoram Bachrach, Marc Lanctot, Roma Patel, Vibhavari Dasagi, Luke Marris, Georgios Piliouras, and Karl Tuyls · 2024
Closest in time.
MindAgent: Emergent gaming interaction
Ran Gong, Qiuyuan Huang, Xiaojian Ma, Yusuke Noda, Zane Durante, Zilong Zheng, Demetri Terzopoulos, Li Fei-Fei, Jianfeng Gao, and Hoi Vo · 2024
Closest in time.
Decision-oriented dialogue for human-AI collaboration, 2024
Jessy Lin, Nicholas Tomlin, Jacob Andreas, and Jason Eisner · 2024
Closest in time.
Autonomous evaluation and refinement of digital agents
Jiayi Pan, Yichi Zhang, Nicholas Tomlin, Yifei Zhou, Sergey Levine, and Alane Suhr · 2024
Closest in time.
Beyond human data: Scaling self-training for problem-solving with language models
Avi Singh, John D Co-Reyes, Rishabh Agarwal, Ankesh Anand, Piyush Patil, Xavier Garcia, Peter J Liu, James Harrison, Jaehoon Lee, Kelvin Xu, Aaron T Parisi, Abhishek Kumar, Alexander A Alemi, Alex Rizkowsky, Azade Nova, Ben Adlam, Bernd Bohnet, Gamaleldin Fathy Elsayed, Hanie Sedghi, Igor Mordatch, Isabelle Simpson, Izzeddin Gur, Jasper Snoek, Jeffrey Pennington, Jiri Hron, Kathleen Kenealy, Kevin Swersky, Kshiteej Mahajan, Laura A Culp, Lechao Xiao, Maxwell Bileschi, Noah Constant, Roman Novak, Rosanne Liu, Tris Warkentin, Yamini Bansal, Ethan Dyer, Behnam Neyshabur, Jascha Sohl-Dickstein, and Noah Fiedel · 2024
Closest in time.
Policy learning with a language bottleneck
Megha Srivastava, Cedric Colas, Dorsa Sadigh, and Jacob Andreas · 2024
Closest in time.
SOTOPIA- π \pi
Ruiyi Wang, Haofei Yu, Wenxin Zhang, Zhengyang Qi, Maarten Sap, Yonatan Bisk, Graham Neubig, and Hao Zhu · 2024
Closest in time.
SOTOPIA: Interactive evaluation for social intelligence in language agents
Xuhui Zhou, Hao Zhu, Leena Mathur, Ruohong Zhang, Haofei Yu, Zhengyang Qi, Louis-Philippe Morency, Yonatan Bisk, Daniel Fried, Graham Neubig, and Maarten Sap · 2024
Closest in time.