Fetching the paper…
Reading the bibliography…
When intelligent agents communicate to accomplish shared goals, how do these goals shape the agents' language? We study the dynamics of learning in latent language policies (LLPs), in which instructor agents generate natural-language subgoal descriptions and executor agents map these descriptions to low-level actions.
Convention: A philosophical study
David Lewis. 1969 · 1969
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams. 1992 · 1992
Earlier work this paper cites.
Feudal reinforcement learning
Peter Dayan and Geoffrey E. Hinton. 1993 · 1993
Earlier work this paper cites.
Understanding language change
April MS McMahon and McMahon April. 1994 · 1994
Earlier work this paper cites.
Why do new meanings occur? A cognitive typology of the motivations for lexical semantic change
Andreas Blank. 1999 · 1999
Earlier work this paper cites.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
Sutton Richard, Precup Doina, and Singh Satinder. 1999 · 1999
Earlier work this paper cites.
Hierarchical reinforcement learning with the maxq value function decomposition
Thomas G Dietterich. 2000 · 2000
Earlier work this paper cites.
On the sample complexity of reinforcement learning
Sham Machandranath Kakade et al. 2003 · 2003
Earlier work this paper cites.
Communicative need modulates competition in language change
Andres Karjus, Richard A Blythe, Simon Kirby, and Kenny Smith. 2020 · 2006
Earlier work this paper cites.
Frequency of word-use predicts rates of lexical evolution throughout Indo-European history
Mark Pagel, Quentin D Atkinson, and Andrew Meade. 2007 · 2007
Earlier work this paper cites.
Evolutionary dynamics of Lewis signaling games: signaling systems vs. partial pooling
Simon M Huttegger, Brian Skyrms, Rory Smead, and Kevin JS Zollman. 2010 · 2010
Earlier work this paper cites.
Predicting pragmatic reasoning in language games
Michael C Frank and Noah D Goodman. 2012 · 2012
Earlier work this paper cites.
Sample complexity of multi-task reinforcement learning
Emma Brunskill and Lihong Li. 2013 · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba. 2014 · 2014
Earlier work this paper cites.
Sequence level training with recurrent neural networks
Marc’Aurelio Ranzato, Sumit Chopra, Michael Auli, and Wojciech Zaremba. 2015 · 2015
Earlier work this paper cites.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz. 2015 · 2015
Earlier work this paper cites.
Deep reinforcement learning for robotic manipulation
Shixiang Gu, Ethan Holly, Timothy P. Lillicrap, and Sergey Levine. 2016 · 2016
Cited alongside, same era.
Reinforcement learning with unsupervised auxiliary tasks
Max Jaderberg, Volodymyr Mnih, Wojciech Marian Czarnecki, Tom Schaul, Joel Z Leibo, David Silver, and Koray Kavukcuoglu. 2016 · 2016
Cited alongside, same era.
Deep reinforcement learning for dialogue generation
Jiwei Li, Will Monroe, Alan Ritter, Dan Jurafsky, Michel Galley, and Jianfeng Gao. 2016 · 2016
Cited alongside, same era.
Safe and efficient off-policy reinforcement learning
Rémi Munos, Tom Stepleton, Anna Harutyunyan, and Marc Bellemare. 2016 · 2016
Cited alongside, same era.
Loss is its own reward: Self-supervision for reinforcement learning
Evan Shelhamer, Parsa Mahmoudieh, Max Argus, and Trevor Darrell. 2016 · 2016
Core knowledge, language, and number
Elizabeth S Spelke. 2017 · 2017
Later among the works it cites.
Learning with latent language
Jacob Andreas, Dan Klein, and Sergey Levine. 2018 · 2018
Later among the works it cites.
Hierarchical and interpretable skill acquisition in multi-task reinforcement learning
Tianmin Shu, Caiming Xiong, and Richard Socher. 2018 · 2018
Later among the works it cites.
Community regularization of visually-grounded dialog
Akshat Agarwal, Swaminathan Gurumurthy, Vasu Sharma, Mike Lewis, and Katia Sycara. 2019 · 2019
Later among the works it cites.
Bam! Born-again multi-task networks for natural language understanding
Kevin Clark, Minh-Thang Luong, Urvashi Khandelwal, Christopher D Manning, and Quoc Le. 2019 · 2019
Later among the works it cites.
Sharing knowledge in multi-task deep reinforcement learning
Carlo D’Eramo, Davide Tateo, Andrea Bonarini, Marcello Restelli, and Jan Peters. 2019 · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Mastering the game of go with deep neural networks and tree search
D. Silver, A. Huang, C. Maddison, A. Guez, L. Sifre, G. Van Den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, and M. Lanctot. 2016 · 2016
Cited alongside, same era.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, et al. 2016 · 2016
Cited alongside, same era.
Modular multitask reinforcement learning with policy sketches
Jacob Andreas, Dan Klein, and Sergey Levine. 2017 · 2017
Cited alongside, same era.
The option-critic architecture
Pierre-Luc Bacon, Jean Harb, and Doina Precup. 2017 · 2017
Cited alongside, same era.
Color naming across languages reflects color use
Edward Gibson, Richard Futrell, Julian Jara-Ettinger, Kyle Mahowald, Leon Bergen, Sivalogeswaran Ratnasingam, Mitchell Gibson, Steven T Piantadosi, and Bevil R Conway. 2017 · 2017
Cited alongside, same era.
Natural language does not emerge ‘naturally’in multi-agent dialog
Satwik Kottur, José Moura, Stefan Lee, and Dhruv Batra. 2017 · 2017
Cited alongside, same era.
Deal or no deal? End-to-end learning of negotiation dialogues
Mike Lewis, Denis Yarats, Yann Dauphin, Devi Parikh, and Dhruv Batra. 2017 · 2017
Cited alongside, same era.
Later among the works it cites.
Seeded self-play for language learning
Abhinav Gupta, Ryan Lowe, Jakob Foerster, Douwe Kiela, and Joelle Pineau. 2019 · 2019
Later among the works it cites.
Hierarchical decision making by generating and following natural language instructions
Hengyuan Hu, Denis Yarats, Qucheng Gong, Yuandong Tian, and Mike Lewis. 2019 · 2019
Later among the works it cites.
Language as an abstraction for hierarchical deep reinforcement learning
Yiding Jiang, Shixiang Shane Gu, Kevin P Murphy, and Chelsea Finn. 2019 · 2019
Later among the works it cites.
Countering language drift via visual grounding
Jason Lee, Kyunghyun Cho, and Douwe Kiela. 2019 · 2019
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. 2019 · 2019
Later among the works it cites.
Don’t stop pretraining: Adapt language models to domains and tasks
Suchin Gururangan, Ana Marasović, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A Smith. 2020 · 2020
Later among the works it cites.
Multi-agent communication meets natural language: Synergies between functional and structural language learning
Angeliki Lazaridou, Anna Potapenko, and Olivier Tieleman. 2020 · 2020
Later among the works it cites.
Countering language drift with seeded iterated learning
Yuchen Lu, Soumye Singhal, Florian Strub, Aaron Courville, and Olivier Pietquin. 2020 · 2020
Later among the works it cites.
Learning to summarize with human feedback
Nisan Stiennon, Long Ouyang, Jeffrey Wu, Daniel Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, and Paul F Christiano. 2020 · 2020
Later among the works it cites.
Rate of language evolution is affected by population size
Lindell Bromham, Xia Hua, Thomas G Fitzpatrick, and Simon J Greenhill. 2015 · 2097
Closest in time.