Fetching the paper…
Reading the bibliography…
Problems which require both long-horizon planning and continuous control capabilities pose significant challenges to existing reinforcement learning agents.
“On a Test of Whether one of Two Random Variables is Stochastically Larger than the Other”
H.. Mann and D.. Whitney · 1947
Earlier work this paper cites.
“On a Test of Whether one of Two Random Variables is Stochastically Larger than the Other”
H.. Mann and D.. Whitney · 1947
Earlier work this paper cites.
“The Hungarian method for the assignment problem”
Harold Kuhn · 1955
Earlier work this paper cites.
“The Hungarian method for the assignment problem”
Harold Kuhn · 1955
Earlier work this paper cites.
“Algorithms for the assignment and transportation problems”
James Munkres · 1957
Earlier work this paper cites.
“Algorithms for the assignment and transportation problems”
James Munkres · 1957
Earlier work this paper cites.
“Feudal Reinforcement Learning”
Peter Dayan and Geoffrey. Hinton · 1992
Earlier work this paper cites.
“Feudal Reinforcement Learning”
Peter Dayan and Geoffrey. Hinton · 1992
Earlier work this paper cites.
“Between MDPs and Semi-MDPs: A Framework for Temporal Abstraction in Reinforcement Learning”
Richard. Sutton, Doina Precup and Satinder. Singh · 1999
Earlier work this paper cites.
“Between MDPs and Semi-MDPs: A Framework for Temporal Abstraction in Reinforcement Learning”
Richard. Sutton, Doina Precup and Satinder. Singh · 1999
Earlier work this paper cites.
“Hierarchical Reinforcement Learning with the MAXQ Value Function Decomposition”
Thomas. Dietterich · 2000
Earlier work this paper cites.
“Hierarchical Reinforcement Learning with the MAXQ Value Function Decomposition”
Thomas. Dietterich · 2000
Earlier work this paper cites.
“Using Abstract Models of Behaviours to Automatically Generate Reinforcement Learning Hierarchies”
Malcolm.. Ryan · 2002
Earlier work this paper cites.
“Using Abstract Models of Behaviours to Automatically Generate Reinforcement Learning Hierarchies”
Malcolm.. Ryan · 2002
Earlier work this paper cites.
“FF-Replan: A Baseline for Probabilistic Planning”
Sung Yoon, Alan Fern and Robert Givan · 2007
Earlier work this paper cites.
“FF-Replan: A Baseline for Probabilistic Planning”
Sung Yoon, Alan Fern and Robert Givan · 2007
Earlier work this paper cites.
“Physically Embedded Planning Problems: New Challenges for Reinforcement Learning”
Mehdi Mirza, Andrew Jaegle, Jonathan. Hunt, Arthur Guez, Saran Tunyasuvunakool, Alistair Muldal, Théophane Weber, Péter Karkus, Sébastien Racanière, Lars Buesing, Timothy. Lillicrap and Nicolas Heess · 2009
Earlier work this paper cites.
“Physically Embedded Planning Problems: New Challenges for Reinforcement Learning”
Mehdi Mirza, Andrew Jaegle, Jonathan. Hunt, Arthur Guez, Saran Tunyasuvunakool, Alistair Muldal, Théophane Weber, Péter Karkus, Sébastien Racanière, Lars Buesing, Timothy. Lillicrap and Nicolas Heess · 2009
Earlier work this paper cites.
Alper Ahmetoglu, M. Seker, Aysu Sayin, Serkan Bugur, Justus. Piater, Erhan Öztop and Emre Ugur · 2012
Earlier work this paper cites.
Alper Ahmetoglu, M. Seker, Aysu Sayin, Serkan Bugur, Justus. Piater, Erhan Öztop and Emre Ugur · 2012
Earlier work this paper cites.
“Bottom-up learning of object categories, action effects and logical rules: From continuous manipulative exploration to symbolic planning”
Emre Ugur and Justus Piater · 2015
Earlier work this paper cites.
“Adam: A Method for Stochastic Optimization”
Diederik. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
“Bottom-up learning of object categories, action effects and logical rules: From continuous manipulative exploration to symbolic planning”
Emre Ugur and Justus Piater · 2015
Earlier work this paper cites.
“Adam: A Method for Stochastic Optimization”
Diederik. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
“Mastering the Game of Go with Deep Neural Networks and Tree Search”
David Silver, Aja Huang, Chris. Maddison, Arthur Guez, Laurent Sifre, George van Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, Sander Dieleman, Dominik Grewe, John Nham, Nal Kalchbrenner, Ilya Sutskever, Timothy Lillicrap, Madeleine Leach, Koray Kavukcuoglu, Thore Graepel and Demis Hassabis · 2016
Earlier work this paper cites.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang and Wojciech Zaremba · 2016
Earlier work this paper cites.
“Mastering the Game of Go with Deep Neural Networks and Tree Search”
David Silver, Aja Huang, Chris. Maddison, Arthur Guez, Laurent Sifre, George van Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, Sander Dieleman, Dominik Grewe, John Nham, Nal Kalchbrenner, Ilya Sutskever, Timothy Lillicrap, Madeleine Leach, Koray Kavukcuoglu, Thore Graepel and Demis Hassabis · 2016
Earlier work this paper cites.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang and Wojciech Zaremba · 2016
Earlier work this paper cites.
“Variational Intrinsic Control”
Karol Gregor, Danilo Rezende and Daan Wierstra · 2017
Earlier work this paper cites.
“An Efficient Approach to Model-Based Hierarchical Reinforcement Learning”
Zhuoru Li, Akshay Narayan and Tze-Yun Leong · 2017
Earlier work this paper cites.
“Stochastic Neural Networks for Hierarchical Reinforcement Learning”
Carlos Florensa, Yan Duan and Pieter Abbeel · 2017
Earlier work this paper cites.
“The Option-Critic Architecture”
Pierre-Luc Bacon, Jean Harb and Doina Precup · 2017
Earlier work this paper cites.
“Hindsight Experience Replay”
Marcin Andrychowicz, Dwight Crow, Alex Ray, Jonas Schneider, Rachel Fong, Peter Welinder, Bob McGrew, Josh Tobin, Pieter Abbeel and Wojciech Zaremba · 2017
Earlier work this paper cites.
“Kinova Modular Robot Arms for Service Robotics Applications”
Alexandre Campeau-Lecours, Hugo Lamontagne, Simon Latour, Philippe Fauteux, Véronique Maheu, François Boucher, Charles Deguire and Louis-Joseph L’Ecuyer · 2017
Earlier work this paper cites.
“Variational Intrinsic Control”
Karol Gregor, Danilo Rezende and Daan Wierstra · 2017
Earlier work this paper cites.
“The Concrete Distribution: A Continuous Relaxation of Discrete Random Variables”
Chris. Maddison, Andriy Mnih and Yee Teh · 2017
Earlier work this paper cites.
“Categorical Reparameterization with Gumbel-Softmax”
Eric Jang, Shixiang Gu and Ben Poole · 2017
Cited alongside, same era.
“Variational Intrinsic Control”
Karol Gregor, Danilo Rezende and Daan Wierstra · 2017
Cited alongside, same era.
“An Efficient Approach to Model-Based Hierarchical Reinforcement Learning”
Zhuoru Li, Akshay Narayan and Tze-Yun Leong · 2017
Cited alongside, same era.
“Stochastic Neural Networks for Hierarchical Reinforcement Learning”
Carlos Florensa, Yan Duan and Pieter Abbeel · 2017
Cited alongside, same era.
“The Option-Critic Architecture”
Pierre-Luc Bacon, Jean Harb and Doina Precup · 2017
Cited alongside, same era.
“Hindsight Experience Replay”
Marcin Andrychowicz, Dwight Crow, Alex Ray, Jonas Schneider, Rachel Fong, Peter Welinder, Bob McGrew, Josh Tobin, Pieter Abbeel and Wojciech Zaremba · 2017
Cited alongside, same era.
“Diversity is All You Need: Learning Skills without a Reward Function”
Benjamin Eysenbach, Abhishek Gupta, Julian Ibarz and Sergey Levine · 2019
Later among the works it cites.
“Unsupervised Control Through Non-Parametric Discriminative Rewards”
David Warde-Farley, Tom de Wiele, Tejas. Kulkarni, Catalin Ionescu, Steven Hansen and Volodymyr Mnih · 2019
Later among the works it cites.
“Learning Multi-Level Hierarchies with Hindsight”
Andrew Levy, George Konidaris, Robert Jr. and Kate Saenko · 2019
Later among the works it cites.
“SDRL: Interpretable and Data-Efficient Deep Reinforcement Learning Leveraging Symbolic Planning”
Daoming Lyu, Fangkai Yang, Bo Liu and Steven Gustafson · 2019
Later among the works it cites.
“Learning Reward Machines for Partially Observable Reinforcement Learning”
Rodrigo Toro, Ethan Waldie, Toryn Klassen, Rick Valenzano, Margarita Castro and Sheila McIlraith · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Kinova Modular Robot Arms for Service Robotics Applications”
Alexandre Campeau-Lecours, Hugo Lamontagne, Simon Latour, Philippe Fauteux, Véronique Maheu, François Boucher, Charles Deguire and Louis-Joseph L’Ecuyer · 2017
Cited alongside, same era.
“Variational Intrinsic Control”
Karol Gregor, Danilo Rezende and Daan Wierstra · 2017
Cited alongside, same era.
“The Concrete Distribution: A Continuous Relaxation of Discrete Random Variables”
Chris. Maddison, Andriy Mnih and Yee Teh · 2017
Cited alongside, same era.
“Categorical Reparameterization with Gumbel-Softmax”
Eric Jang, Shixiang Gu and Ben Poole · 2017
Cited alongside, same era.
Yuval Tassa, Yotam Doron, Alistair Muldal, Tom Erez, Yazhe Li, Diego de Las, David Budden, Abbas Abdolmaleki, Josh Merel, Andrew Lefrancq, Timothy. Lillicrap and Martin. Riedmiller · 2018
Cited alongside, same era.
“On the necessity of abstraction” Artificial Intelligence
George Konidaris · 2018
Cited alongside, same era.
Archit Sharma, Shixiang Gu, Sergey Levine, Vikash Kumar and Karol Hausman · 2020
Later among the works it cites.
“Option Discovery using Deep Skill Chaining”
Akhil Bagaria and George Konidaris · 2020
Later among the works it cites.
“Sub-policy Adaptation for Hierarchical Reinforcement Learning”
Alexander. Li, Carlos Florensa, Ignasi Clavera and Pieter Abbeel · 2020
Later among the works it cites.
“Symbolic Plans as High-Level Instructions for Reinforcement Learning”
León Illanes, Xi Yan, Rodrigo Icarte and Sheila. McIlraith · 2020
Later among the works it cites.
“Learning Portable Representations for High-Level Planning”
Steven James, Benjamin Rosman and George Konidaris · 2020
Later among the works it cites.
“Hierarchical-Actor-Critc-HAC-”
Andrew Levy · 2020
Later among the works it cites.
“Dynamics-Aware Unsupervised Discovery of Skills”
Archit Sharma, Shixiang Gu, Sergey Levine, Vikash Kumar and Karol Hausman · 2020
Later among the works it cites.
“Option Discovery using Deep Skill Chaining”
Akhil Bagaria and George Konidaris · 2020
Later among the works it cites.
“Sub-policy Adaptation for Hierarchical Reinforcement Learning”
Alexander. Li, Carlos Florensa, Ignasi Clavera and Pieter Abbeel · 2020
Later among the works it cites.
“Symbolic Plans as High-Level Instructions for Reinforcement Learning”
León Illanes, Xi Yan, Rodrigo Icarte and Sheila. McIlraith · 2020
Later among the works it cites.
“Learning Portable Representations for High-Level Planning”
Steven James, Benjamin Rosman and George Konidaris · 2020
Later among the works it cites.
“Hierarchical-Actor-Critc-HAC-”
Andrew Levy · 2020
Later among the works it cites.
“Integrated Task and Motion Planning”
Caelan Garrett, Rohan Chitnis, Rachel Holladay, Beomjoon Kim, Tom Silver, Leslie Kaelbling and Tomás Lozano-Pérez · 2021
Later among the works it cites.
“Relative Variational Intrinsic Control”
Kate Baumli, David Warde-Farley, Steven Hansen and Volodymyr Mnih · 2021
Later among the works it cites.
“Hierarchical Reinforcement Learning By Discovering Intrinsic Options”
Jesse Zhang, Haonan Yu and Wei Xu · 2021
Later among the works it cites.
“RePReL: Integrating Relational Planning and Reinforcement Learning for Effective Abstraction”
Harsha Kokel, Arjun Manoharan, Sriraam Natarajan, Balaraman Ravindran and Prasad Tadepalli · 2021
Later among the works it cites.
“uArm-Python-SDK”
uArm-Developer · 2021
Later among the works it cites.
“pytorch-soft-actor-critic”
Pranjal Tandon · 2021
Later among the works it cites.
“Integrated Task and Motion Planning”
Caelan Garrett, Rohan Chitnis, Rachel Holladay, Beomjoon Kim, Tom Silver, Leslie Kaelbling and Tomás Lozano-Pérez · 2021
Later among the works it cites.
“Relative Variational Intrinsic Control”
Kate Baumli, David Warde-Farley, Steven Hansen and Volodymyr Mnih · 2021
Later among the works it cites.
“Hierarchical Reinforcement Learning By Discovering Intrinsic Options”
Jesse Zhang, Haonan Yu and Wei Xu · 2021
Later among the works it cites.
“RePReL: Integrating Relational Planning and Reinforcement Learning for Effective Abstraction”
Harsha Kokel, Arjun Manoharan, Sriraam Natarajan, Balaraman Ravindran and Prasad Tadepalli · 2021
Later among the works it cites.
“uArm-Python-SDK”
uArm-Developer · 2021
Later among the works it cites.
“pytorch-soft-actor-critic”
Pranjal Tandon · 2021
Later among the works it cites.
“Leveraging Approximate Symbolic Models for Reinforcement Learning via Skill Diversity”
Lin Guan, Sarath Sreedharan and Subbarao Kambhampati · 2022
Closest in time.
“Learning Neuro-Symbolic Relational Transition Models for Bilevel Planning”
Rohan Chitnis, Tom Silver, Joshua. Tenenbaum, Tomas Lozano-Perez and Leslie Kaelbling · 2022
Closest in time.
“Learning Neuro-Symbolic Skills for Bilevel Planning”
Tom Silver, Ashay Athalye, Joshua. Tenenbaum, Tomas Lozano-Perez and Leslie Kaelbling · 2022
Closest in time.
“Leveraging Approximate Symbolic Models for Reinforcement Learning via Skill Diversity”
Lin Guan, Sarath Sreedharan and Subbarao Kambhampati · 2022
Closest in time.
“Learning Neuro-Symbolic Relational Transition Models for Bilevel Planning”
Rohan Chitnis, Tom Silver, Joshua. Tenenbaum, Tomas Lozano-Perez and Leslie Kaelbling · 2022
Closest in time.
“Learning Neuro-Symbolic Skills for Bilevel Planning”
Tom Silver, Ashay Athalye, Joshua. Tenenbaum, Tomas Lozano-Perez and Leslie Kaelbling · 2022
Closest in time.