Fetching the paper…
Reading the bibliography…
Describing skills in natural language has the potential to provide an accessible way to inject human knowledge about decision-making into an AI system.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Ralph Allan Bradley and Milton E. Terry · 1952
Earlier work this paper cites.
An axiomatic basis for computer programming
C. A. R. Hoare · 1969
Earlier work this paper cites.
A heuristic approach to the discovery of macro-operators
Glenn A. Iba · 1989
Earlier work this paper cites.
Made-up minds - a constructivist approach to artificial intelligence
Gary L. Drescher · 1991
Earlier work this paper cites.
Feudal reinforcement learning
Peter Dayan and Geoffrey E Hinton · 1993
Earlier work this paper cites.
Learning and executing generalized robot plans
Richard Fikes, Peter E. Hart, and Nils J. Nilsson · 1993
Earlier work this paper cites.
Python tutorial , volume 620
Guido Van Rossum and Fred L Drake Jr · 1995
Earlier work this paper cites.
Pddl-the planning domain definition language
Drew McDermott, Malik Ghallab, Adele E. Howe, Craig A. Knoblock, Ashwin Ram, Manuela M. Veloso, Daniel S. Weld, and David E. Wilkins · 1998
Earlier work this paper cites.
Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning
Richard S. Sutton, Doina Precup, and Satinder Singh · 1999
Earlier work this paper cites.
Temporal Abstraction in Reinforcement Learning
Doina Precup · 2000
Earlier work this paper cites.
Automatic discovery of subgoals in reinforcement learning using diverse density
Amy McGovern and Andrew G. Barto · 2001
Earlier work this paper cites.
Q-cut - dynamic discovery of sub-goals in reinforcement learning
Ishai Menache, Shie Mannor, and Nahum Shimkin · 2002
Earlier work this paper cites.
Learning options in reinforcement learning
Martin Stolle and Doina Precup · 2002
Earlier work this paper cites.
Using relative novelty to identify useful temporal abstractions in reinforcement learning
Özgür Şimşek and Andrew G. Barto · 2004
Earlier work this paper cites.
Concurrent hierarchical reinforcement learning
Bhaskara Marthi, Stuart Russell, David Latham, and Carlos Guestrin · 2005
Earlier work this paper cites.
Reinforcement learning with human teachers: Evidence of feedback and guidance with implications for learning performance
Andrea Lockerd Thomaz, Cynthia Breazeal, et al · 2006
Earlier work this paper cites.
Keep your options open: An information-based driving principle for sensorimotor systems
Alexander S Klyubin, Daniel Polani, and Chrystopher L Nehaniv · 2008
Earlier work this paper cites.
Interactively shaping agents via human reinforcement: The tamer framework
W Bradley Knox and Peter Stone · 2009
Earlier work this paper cites.
Active reward learning
Christian Daniel, Malte Viering, Jan Metz, Oliver Kroemer, and Jan Peters · 2014
Earlier work this paper cites.
Unifying count-based exploration and intrinsic motivation
Marc G. Bellemare, Sriram Srinivasan, Georg Ostrovski, Tom Schaul, David Saxton, and Rémi Munos · 2016
Earlier work this paper cites.
Modular multitask reinforcement learning with policy sketches
Jacob Andreas, Dan Klein, and Sergey Levine · 2017
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul Francis Christiano, Jan Leike, Tom B. Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Cited alongside, same era.
Variational intrinsic control
Karol Gregor, Danilo Jimenez Rezende, and Daan Wierstra · 2017
Cited alongside, same era.
When waiting is not an option : Learning options with a deliberation cost
Jean Harb, Pierre-Luc Bacon, Martin Klissarov, and Doina Precup · 2017
Cited alongside, same era.
A laplacian framework for option discovery in reinforcement learning
Marlos C Machado, Marc G Bellemare, and Michael Bowling · 2017
Cited alongside, same era.
Film: Visual reasoning with a general conditioning layer
Ethan Perez, Florian Strub, Harm de Vries, Vincent Dumoulin, and Aaron C. Courville · 2017
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al · 2022
Later among the works it cites.
React: Synergizing reasoning and acting in language models
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao · 2022
Later among the works it cites.
Palm-e: An embodied multimodal language model
Danny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch, Aakanksha Chowdhery, Brian Ichter, Ayzaan Wahid, Jonathan Tompson, Quan Vuong, Tianhe Yu, Wenlong Huang, Yevgen Chebotar, Pierre Sermanet, Daniel Duckworth, Sergey Levine, Vincent Vanhoucke, Karol Hausman, Marc Toussaint, Klaus Greff, Andy Zeng, Igor Mordatch, and Pete Florence · 2023
Later among the works it cites.
Deep laplacian-based options for temporally-extended exploration
Martin Klissarov and Marlos C. Machado · 2023
Later among the works it cites.
Efficient memory management for large language model serving with pagedattention
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
Feudal networks for hierarchical reinforcement learning
Alexander Sasha Vezhnevets, Simon Osindero, Tom Schaul, Nicolas Heess, Max Jaderberg, David Silver, and Koray Kavukcuoglu · 2017
Cited alongside, same era.
From skills to symbols: Learning symbolic representations for abstract high-level planning
George Dimitri Konidaris, Leslie Pack Kaelbling, and Tomas Lozano-Perez · 2018
Cited alongside, same era.
Options of interest: Temporal abstraction with interest functions
Khimya Khetarpal, Martin Klissarov, Maxime Chevalier-Boisvert, Pierre-Luc Bacon, and Doina Precup · 2020
Cited alongside, same era.
The NetHack Learning Environment
Heinrich Küttler, Nantas Nardelli, Alexander H. Miller, Roberta Raileanu, Marco Selvatici, Edward Grefenstette, and Tim Rocktäschel · 2020
Cited alongside, same era.
Sample factory: Egocentric 3d control from pixels at 100000 FPS with asynchronous reinforcement learning
Aleksei Petrenko, Zhehui Huang, Tushar Kumar, Gaurav S. Sukhatme, and Vladlen Koltun · 2020
Cited alongside, same era.
Program guided agent
Shao-Hua Sun, Te-Lin Wu, and Joseph J. Lim · 2020
Cited alongside, same era.
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, and Ion Stoica · 2023
Later among the works it cites.
Code as policies: Language model programs for embodied control
Jacky Liang, Wenlong Huang, Fei Xia, Peng Xu, Karol Hausman, Brian Ichter, Pete Florence, and Andy Zeng · 2023
Later among the works it cites.
Llm+p: Empowering large language models with optimal planning proficiency
B. Liu, Yuqian Jiang, Xiaohan Zhang, Qian Liu, Shiqi Zhang, Joydeep Biswas, and Peter Stone · 2023
Later among the works it cites.
Temporal abstraction in reinforcement learning with the successor representation
Marlos C. Machado, André Barreto, and Doina Precup · 2023
Later among the works it cites.
Self-refine: Iterative refinement with self-feedback
Aman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan, Luyu Gao, Sarah Wiegreffe, Uri Alon, Nouha Dziri, Shrimai Prabhumoye, Yiming Yang, Sean Welleck, Bodhisattwa Prasad Majumder, Shashank Gupta, Amir Yazdanbakhsh, and Peter Clark · 2023
Later among the works it cites.
Embodied lifelong learning for task and motion planning
Jorge Mendez-Mendez, Leslie Pack Kaelbling, and Tomás Lozano-Pérez · 2023
Later among the works it cites.
Reflexion: language agents with verbal reinforcement learning
Noah Shinn, Federico Cassano, Beck Labash, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao · 2023
Later among the works it cites.
Generalized planning in pddl domains with pretrained large language models
Tom Silver, Soham Dan, Kavitha Srinivas, Joshua B. Tenenbaum, Leslie Pack Kaelbling, and Michael Katz · 2023
Later among the works it cites.
Does zero-shot reinforcement learning exist?
Ahmed Touati, Jérémy Rapin, and Yann Ollivier · 2023
Later among the works it cites.
Abhimanyu Dubey, Abhinav Jauhri, and Abhinav Pandey et al · 2024
Closest in time.
Open-endedness is essential for artificial superhuman intelligence, 06 2024
Edward Hughes, Michael Dennis, Jack Parker-Holder, Feryal Behbahani, Aditi Mavalankar, Yuge Shi, Tom Schaul, and Tim Rocktaschel · 2024
Closest in time.
Motif: Intrinsic motivation from artificial intelligence feedback
Martin Klissarov, Pierluca D’Oro, Shagun Sodhani, Roberta Raileanu, Pierre-Luc Bacon, Pascal Vincent, Amy Zhang, and Mikael Henaff · 2024
Closest in time.
Practice makes perfect: Planning to learn skill parameter policies
Nishanth Kumar, Tom Silver, Willie McClinton, Linfeng Zhao, Stephen Proulx, Tomás Lozano-Pérez, Leslie Pack Kaelbling, and Jennifer Barry · 2024
Closest in time.
Nethack: an illustrated guide to the mazes of menace, dec 2022
Dion Moult · 2024
Closest in time.
Voyager: An open-ended embodied agent with large language models
Guanzhi Wang, Yuqi Xie, Yunfan Jiang, Ajay Mandlekar, Chaowei Xiao, Yuke Zhu, Linxi Fan, and Anima Anandkumar · 2024
Closest in time.
Fine-tuning reinforcement learning models is secretly a forgetting mitigation problem
Maciej Wolczyk, Bartłomiej Cupiał, Mateusz Ostaszewski, Michal Bortkiewicz, Michal Zajkac, Razvan Pascanu, Lukasz Kuci’nski, and Piotr Milo’s · 2024
Closest in time.