Fetching the paper…
Reading the bibliography…
This paper proposes DeepSynth, a method for effective training of deep Reinforcement Learning (RL) agents when the reward is sparse and non-Markovian, but at the same time progress towards the reward requires achieving an unknown sequence of high-level objectives.
Certified Reinforcement Learning with Logic Guidance
Hasanbeig, M.; Abate, A.; and Kroening, D. 2019a · 1902
Earlier work this paper cites.
Modular Deep Reinforcement Learning with Temporal Logic Specifications
Yuan, L. Z.; Hasanbeig, M.; Abate, A.; and Kroening, D. 2019 · 1909
Earlier work this paper cites.
DeepSynth: Program Synthesis for Automatic Task Segmentation in Deep Reinforcement Learning [Extended Version]
Hasanbeig, M.; Yogananda Jeppu, N.; Abate, A.; Melham, T.; and Kroening, D. 2019b · 1911
Earlier work this paper cites.
Unsupervised Object Segmentation with Explicit Localization Module
Liu, W.; Wei, L.; Sharpnack, J.; and Owens, J. D. 2019 · 1911
Earlier work this paper cites.
Active Task-Inference-Guided Deep Inverse Reinforcement Learning
Memarian, F.; Xu, Z.; Wu, B.; Wen, M.; and Topcu, U. 2020 · 1938
Earlier work this paper cites.
Conflict, Arousal, and Curiosity
Berlyne, D. E. 1960 · 1960
Earlier work this paper cites.
A Computing Procedure for Quantification Theory
Davis, M.; and Putnam, H. 1960 · 1960
Earlier work this paper cites.
Complexity of Automaton Identification from Given Data
Gold, E. M. 1978 · 1978
Earlier work this paper cites.
Learning Regular Sets from Queries and Counterexamples
Angluin, D. 1987 · 1987
Earlier work this paper cites.
Flow: The Psychology of Optimal Experience
Csikszentmihalyi, M. 1990 · 1990
Earlier work this paper cites.
Q-learning
Watkins, C. J.; and Dayan, P. 1992 · 1992
Earlier work this paper cites.
The Rabin Index and Chain Automata, with Applications to Automata and Games
Krishnan, S. C.; Puri, A.; Brayton, R. K.; and Varaiya, P. P. 1995 · 1995
Earlier work this paper cites.
Neuro-dynamic Programming , volume 1
Bertsekas, D. P.; and Tsitsiklis, J. N. 1996 · 1996
Earlier work this paper cites.
Finding Hard Instances of the Satisfiability Problem: A Survey
Cook, S.; and Mitchell, D. 1996 · 1996
Earlier work this paper cites.
Results of the Abbadingo One DFA Learning Competition and a New Evidence-driven State Merging Algorithm
Lang, K. J.; Pearlmutter, B. A.; and Price, R. A. 1998 · 1998
Earlier work this paper cites.
Reinforcement Learning: An Introduction , volume 1
Sutton, R. S.; and Barto, A. G. 1998 · 1998
Earlier work this paper cites.
Policy Invariance Under Reward Transformations: Theory and Application to Reward Shaping
Ng, A. Y.; Harada, D.; and Russell, S. 1999 · 1999
Earlier work this paper cites.
Intrinsic and Extrinsic Motivations: Classic Definitions and New Directions
Ryan, R. M.; and Deci, E. L. 2000 · 2000
Earlier work this paper cites.
Temporal Abstraction in Reinforcement Learning
Precup, D. 2001 · 2001
Earlier work this paper cites.
Learning Non-Markovian Reward Models in MDPs
Rens, G.; and Raskin, J.-F. 2020 · 2001
Earlier work this paper cites.
Near-optimal Reinforcement Learning in Polynomial Time
Kearns, M.; and Singh, S. 2002 · 2002
Earlier work this paper cites.
Stochastic Optimal Control: The Discrete-time Case
Bertsekas, D. P.; and Shreve, S. 2004 · 2004
Earlier work this paper cites.
Formal Policy Synthesis for Continuous-Space Systems via Reinforcement Learning
Kazemi, M.; and Soudjani, S. 2020 · 2005
Earlier work this paper cites.
Neural Fitted Q iteration – First Experiences with a Data Efficient Neural Reinforcement Learning Method
Riedmiller, M. 2005 · 2005
Earlier work this paper cites.
An Application of Reinforcement Learning to Aerobatic Helicopter Flight
Abbeel, P.; Coates, A.; Quigley, M.; and Ng, A. Y. 2007 · 2007
Earlier work this paper cites.
Reverse Engineering State Machines by Interactive Grammar Inference
Walkinshaw, N.; Bogdanov, K.; Holcombe, M.; and Salahuddin, S. 2007 · 2007
Earlier work this paper cites.
Inferring Finite-State Models with Temporal Constraints
Walkinshaw, N.; and Bogdanov, K. 2008 · 2008
Earlier work this paper cites.
Online Learning of Non-Markovian Reward Models
Rens, G.; Raskin, J.-F.; Reynouad, R.; and Marra, G. 2020 · 2009
Cited alongside, same era.
Extended Finite-State Machine Induction Using SAT-Solver
Ulyantsev, V.; and Tsarev, F. 2011 · 2011
Cited alongside, same era.
Hierarchical relative entropy policy search
Daniel, C.; Neumann, G.; and Peters, J. 2012 · 2012
Cited alongside, same era.
Synthesis from Examples
Gulwani, S. 2012 · 2012
Cited alongside, same era.
The Arcade Learning Environment: An Evaluation Platform for General Agents
Bellemare, M. G.; Naddaf, Y.; Veness, J.; and Bowling, M. 2013 · 2013
Cited alongside, same era.
Software Model Synthesis Using Satisfiability Solvers
Heule, M. J. H.; and Verwer, S. 2013 · 2013
Cited alongside, same era.
Logically-Constrained Reinforcement Learning
Hasanbeig, M.; Abate, A.; and Kroening, D. 2018 · 2018
Later among the works it cites.
Teaching Multiple Tasks to an RL Agent using LTL
Toro Icarte, R.; Klassen, T. Q.; Valenzano, R.; and McIlraith, S. A. 2018 · 2018
Later among the works it cites.
Exact Finite-state Machine Identification from Scenarios and Temporal Properties
Ulyantsev, V.; Buzhinsky, I.; and Shalyto, A. 2018 · 2018
Later among the works it cites.
MINT Framework Github Repository
Walkinshaw, N. 2018 · 2018
Later among the works it cites.
LTL and Beyond: Formal Languages for Reward Function Specification in Reinforcement Learning
Camacho, A.; Toro Icarte, R.; Klassen, T. Q.; Valenzano, R.; and McIlraith, S. A. 2019 · 2019
Closest in time.
Foundations for Restraining Bolts: Reinforcement Learning with LTLf/LDLf Restraining Specifications
De Giacomo, G.; Iocchi, L.; Favorito, M.; and Patrizi, F. 2019 · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Probably Approximately Correct MDP Learning and Control With Temporal Logic Constraints
Fu, J.; and Topcu, U. 2014 · 2014
Cited alongside, same era.
Logically-Constrained Neural Fitted Q-Iteration
Hasanbeig, M.; Abate, A.; and Kroening, D. 2019b · 2014
Cited alongside, same era.
A Learning Based Approach to Control Synthesis of Markov Decision Processes for Linear Temporal Logic Specifications
Sadigh, D.; Kim, E. S.; Coogan, S.; Sastry, S. S.; and Seshia, S. A. 2014 · 2014
Cited alongside, same era.
Adam: A Method for Stochastic Optimization
Kingma, D. P.; and Ba, J. 2015 · 2015
Cited alongside, same era.
Human-level Control Through Deep Reinforcement Learning
Mnih, V.; et al. 2015 · 2015
Cited alongside, same era.
OpenAI Gym
Brockman, G.; Cheung, V.; Pettersson, L.; Schneider, J.; Schulman, J.; Tang, J.; and Zaremba, W. 2016 · 2016
Cited alongside, same era.
Closest in time.
Omega-Regular Objectives in Model-Free Reinforcement Learning
Hahn, E. M.; Perez, M.; Schewe, S.; Somenzi, F.; Trivedi, A.; and Wojtczak, D. 2019 · 2019
Closest in time.
SegSort: Segmentation by Discriminative Sorting of Segments
Hwang, J.-J.; Yu, S. X.; Shi, J.; Collins, M. D.; Yang, T.-J.; Zhang, X.; and Chen, L.-C. 2019 · 2019
Closest in time.
Invariant Information Clustering for Unsupervised Image Classification and Segmentation
Ji, X.; Henriques, J. F.; and Vedaldi, A. 2019 · 2019
Closest in time.
Learning Finite State Representations of Recurrent Policy Networks
Koul, A.; Fern, A.; and Greydanus, S. 2019 · 2019
Closest in time.
Learning Reward Machines for Partially Observable Reinforcement Learning
Toro Icarte, R.; Waldie, E.; Klassen, T.; Valenzano, R.; Castro, M.; and McIlraith, S. 2019 · 2019
Closest in time.
Grandmaster Level in StarCraft II Using Multi-agent Reinforcement Learning
Vinyals, O.; Babuschkin, I.; Czarnecki, W. M.; Mathieu, M.; Dudzik, A.; Chung, J.; Choi, D. H.; Powell, R.; Ewalds, T.; Georgiev, P.; Oh, J.; Horgan, D.; Kroiss, M.; Danihelka, I.; Huang, A.; Sifre, L.; Cai, T.; Agapiou, J. P.; Jaderberg, M.; Vezhnevets, A. S.; Leblond, R.; Pohlen, T.; Dalibard, V.; Budden, D.; Sulsky, Y.; Molloy, J.; Paine, T. L.; Gulcehre, C.; Wang, Z.; Pfaff, T.; Wu, Y.; Ring, R.; Yogatama, D.; Wünsch, D.; McKinney, K.; Smith, O.; Schaul, T.; Lillicrap, T.; Kavukcuoglu, K.; Hassabis, D.; Apps, C.; and Silver, D. 2019 · 2019
Closest in time.
Learning the Language of Software Errors
Chockler, H.; Kesseli, P.; Kroening, D.; and Strichman, O. 2020 · 2020
Closest in time.
Imitation Learning over Heterogeneous Agents with Restraining Bolts
De Giacomo, G.; Favorito, M.; Iocchi, L.; and Patrizi, F. 2020 · 2020
Closest in time.
Induction of Subgoal Automata for Reinforcement Learning
Furelos-Blanco, D.; Law, M.; Russo, A.; Broda, K.; and Jonsson, A. 2020 · 2020
Closest in time.
Reinforcement Learning with Non-Markovian Rewards
Gaon, M.; and Brafman, R. 2020 · 2020
Closest in time.
Cautious Reinforcement Learning with Logical Constraints
Hasanbeig, M.; Abate, A.; and Kroening, D. 2020 · 2020
Closest in time.
Deep Reinforcement Learning with Temporal Logics
Hasanbeig, M.; Kroening, D.; and Abate, A. 2020 · 2020
Closest in time.
Trace2Model Github repository
Jeppu, N. Y. 2020 · 2020
Closest in time.
Learning Concise Models from Long Execution Traces
Jeppu, N. Y.; Melham, T.; Kroening, D.; and O’Leary, J. 2020 · 2020
Closest in time.
Formal Controller Synthesis for Continuous-space MDPs via Model-free Reinforcement Learning
Lavaei, A.; Somenzi, F.; Soudjani, S.; Trivedi, A.; and Zamani, M. 2020 · 2020
Closest in time.
Joint Inference of Reward Machines and Policies for Reinforcement Learning
Xu, Z.; Gavran, I.; Ahmad, Y.; Majumdar, R.; Neider, D.; Topcu, U.; and Wu, B. 2020 · 2020
Closest in time.
Receding Horizon Control-Based Motion Planning With Partially Infeasible LTL Constraints
Cai, M.; Peng, H.; Li, Z.; Gao, H.; and Kan, Z. 2021 · 2021
Closest in time.
First Return, Then Explore
Ecoffet, A.; Huizinga, J.; Lehman, J.; Stanley, K. O.; and Clune, J. 2021 · 2021
Closest in time.
Rectifying Pseudo Label Learning via Uncertainty Estimation for Domain Adaptive Semantic Segmentation
Zheng, Z.; and Yang, Y. 2021 · 2021
Closest in time.
Tabu Search
Glover, F.; and Laguna, M. 1998 · 2093
Closest in time.