Fetching the paper…
Reading the bibliography…
We propose a novel reinforcement learning algorithm, AlphaNPI, that incorporates the strengths of Neural Programmer-Interpreters (NPI) and AlphaZero.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Łukasz Kaiser and Ilya Sutskever · 2015
Earlier work this paper cites.
Hierarchical Monte-Carlo planning
Ngo Anh Vien and Marc Toussaint · 2015
Earlier work this paper cites.
Neural module networks
Jacob Andreas, Marcus Rohrbach, Trevor Darrell, and Dan Klein · 2016
Earlier work this paper cites.
Markovian state and action abstractions for MDPs via hierarchical MCTS
Aijun Bai, Siddharth Srivastava, and Stuart Russell · 2016
Earlier work this paper cites.
Deepcoder: Learning to write programs
Matej Balog, Alexander L Gaunt, Marc Brockschmidt, Sebastian Nowozin, and Daniel Tarlow · 2016
Earlier work this paper cites.
Adaptive neural compilation
Rudy R Bunel, Alban Desmaison, Pawan K Mudigonda, Pushmeet Kohli, and Philip Torr · 2016
Earlier work this paper cites.
Terpret: A probabilistic programming language for program induction
Alexander L Gaunt, Marc Brockschmidt, Rishabh Singh, Nate Kushman, Pushmeet Kohli, Jonathan Taylor, and Daniel Tarlow · 2016
Earlier work this paper cites.
Hierarchical deep reinforcement learning: Integrating temporal abstraction and intrinsic motivation
Tejas D Kulkarni, Karthik Narasimhan, Ardavan Saeedi, and Josh Tenenbaum · 2016
Earlier work this paper cites.
Neural programmer-interpreters
Scott Reed and Nando de Freitas · 2016
Earlier work this paper cites.
Modular multitask reinforcement learning with policy sketches
Jacob Andreas, Dan Klein, and Sergey Levine · 2017
Earlier work this paper cites.
Programming with a differentiable forth interpreter
Matko Bošnjak, Tim Rocktäschel, Jason Naradowsky, and Sebastian Riedel · 2017
Earlier work this paper cites.
Towards synthesizing complex programs from input-output examples
Xinyun Chen, Chang Liu, and Dawn Song · 2017
Cited alongside, same era.
Misha Denil, Sergio Gomez Colmenarejo, Serkan Cabi, David Saxton, and Nando de Freitas · 2017
Cited alongside, same era.
Neural program lattices
Chengtao Li, Daniel Tarlow, Alexander L Gaunt, Marc Brockschmidt, and Nate Kushman · 2017
Cited alongside, same era.
Mastering the game of Go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, et al · 2017
Cited alongside, same era.
Feudal networks for hierarchical reinforcement learning
Alexander Sasha Vezhnevets, Simon Osindero, Tom Schaul, Nicolas Heess, Max Jaderberg, David Silver, and Koray Kavukcuoglu · 2017
Cited alongside, same era.
Improving neural program synthesis with inferred execution traces
Richard Shin, Illia Polosukhin, and Dawn Song · 2018
Later among the works it cites.
Neural program synthesis from diverse demonstration videos
Shao-Hua Sun, Hyeonwoo Noh, Sriram Somasundaram, and Joseph Lim · 2018
Later among the works it cites.
Programmatically interpretable reinforcement learning
Abhinav Verma, Vijayaraghavan Murali, Rishabh Singh, Pushmeet Kohli, and Swarat Chaudhuri · 2018
Later among the works it cites.
Da Xiao, Jo-Yu Liao, and Xingyuan Yuan · 2018
Later among the works it cites.
Neural task programming: Learning to generalize across hierarchical tasks
Danfei Xu, Suraj Nair, Yuke Zhu, Julian Gao, Animesh Garg, Li Fei-Fei, and Silvio Savarese · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ashley D Edwards, Laura Downs, and James C Davidson · 2018
Cited alongside, same era.
Learning to infer graphics programs from hand-drawn images
Kevin Ellis, Daniel Ritchie, Armando Solar-Lezama, and Josh Tenenbaum · 2018
Cited alongside, same era.
Learning explanatory rules from noisy data
Richard Evans and Edward Grefenstette · 2018
Cited alongside, same era.
Parametrized hierarchical procedures for neural programming
Roy Fox, Richard Shin, Sanjay Krishnan, Ken Goldberg, Dawn Song, and Ion Stoica · 2018
Cited alongside, same era.
Ranked reward: Enabling self-play reinforcement learning for combinatorial optimization
Alexandre Laterre, Yunguan Fu, Mohamed Khalil Jabri, Alain-Sam Cohen, David Kas, Karl Hajjar, Torbjorn S Dahl, Amine Kerkeni, and Karim Beguir · 2018
Cited alongside, same era.
Hierarchical reinforcement learning with hindsight
Andrew Levy, Robert Platt, and Kate Saenko · 2018
Cited alongside, same era.
Learning independent causal mechanisms
Giambattista Parascandolo, Niki Kilbertus, Mateo Rojas-Carulla, and Bernhard Schölkopf · 2018
Cited alongside, same era.
Yoshua Bengio, Tristan Deleu, Nasim Rahaman, Nan Rosemary Ke, Sébastien Lachapelle, Olexa Bilaniuk, Anirudh Goyal, and Christopher J. Pal · 2019
Closest in time.
Sample efficient adaptive text-to-speech
Yutian Chen, Yannis Assael, Brendan Shillingford, David Budden, Scott Reed, Heiga Zen, Quan Wang, Luis C. Cobo, Andrew Trask, Ben Laurie, Caglar Gulcehre, Aaron van den Oord, Oriol Vinyals, and Nando de Freitas · 2019
Closest in time.
Learning to infer program sketches
Maxwell I. Nye, Luke B. Hewitt, Joshua B. Tenenbaum, and Armando Solar-Lezama · 2019
Closest in time.
Hierarchical reinforcement learning via advantage-weighted information maximization
Takayuki Osa, Voot Tangkaratt, and Masashi Sugiyama · 2019
Closest in time.
Fast context adaptation via meta-learning
Luisa M. Zintgraf, Kyriacos Shiarlis, Vitaly Kurin, Katja Hofmann, and Shimon Whiteson · 2019
Closest in time.
Neural program meta-induction
Jacob Devlin, Rudy R Bunel, Rishabh Singh, Matthew Hausknecht, and Pushmeet Kohli · 2088
Closest in time.