Fetching the paper…
Reading the bibliography…
Reinforcement Learning (RL) is a popular machine learning paradigm where intelligent agents interact with the environment to fulfill a long-term goal.
H. J. Hermans, “A questionnaire measure of achievement motivation.” JAP , 1970
1970
Earlier work this paper cites.
J. Barwise, “An introduction to first-order logic,” in Studies in Logic and the Foundations of Mathematics , 1977, vol. 90, pp. 5–46
1977
Earlier work this paper cites.
A. E. Roth, “Introduction to the shapley value,” The Shapley Value , pp. 1–27, 1988
1988
Earlier work this paper cites.
A. Feinberg, “Markov decision processes: Discrete stochastic dynamic programming (martin l. puterman),” SIAM Review , 1996
1996
Earlier work this paper cites.
W. T. B. Uther et al. , “Tree based discretization for continuous state space reinforcement learning,” in AAAI , 1998
1998
Earlier work this paper cites.
D. M. Williamson et al. , “‘mental model’comparison of automated and human scoring,” Journal of Educational Measurement , 1999
1999
Earlier work this paper cites.
A. Y. Ng et al. , “Policy invariance under reward transformations: Theory and application to reward shaping,” in ICML , 1999
1999
Earlier work this paper cites.
M. Gevrey et al. , “Review and comparison of methods to study the contribution of variables in artificial neural network models,” Ecol. Modell. , 2003
2003
Earlier work this paper cites.
M. Bilgic et al. , “Explaining recommendations: Satisfaction vs. promotion,” in IUI Workshop , 2005
2005
Earlier work this paper cites.
A. Powers et al. , “The advisor robot: tracing people’s mental model from a robot’s physical attributes,” in HRI , 2006
2006
Earlier work this paper cites.
P. Pu and L. Chen, “Trust building with explanation interfaces,” in IUI , 2006
2006
Earlier work this paper cites.
B. Y. Lim et al. , “Assessing demand for intelligibility in context-aware applications,” in ICUC , 2009
2009
Earlier work this paper cites.
B. Y. Lim et al. , “Why and why not explanations improve the intelligibility of context-aware intelligent systems,” in CHI , 2009
2009
Earlier work this paper cites.
H. Kjellström et al. , “Tracking people interacting with objects,” in CVPR , 2010
2010
Earlier work this paper cites.
P. Lietz, “Research into questionnaire design: A summary of the literature,” IJMR , 2010
2010
Earlier work this paper cites.
H. Hasselt, “Double q-learning,” in NeurIPS , 2010
2010
Earlier work this paper cites.
S. Ross et al. , “A reduction of imitation learning and structured prediction to no-regret online learning,” in AISTATS , 2011
2011
Earlier work this paper cites.
R. V. Yampolskiy et al. , “Artificial general intelligence and the human mental model,” in Singularity Hypotheses , 2012
2012
Earlier work this paper cites.
F. Maes et al. , “Policy search in a space of simple closed-form formulas: Towards interpretability of reinforcement learning,” in DS , 2012
2012
Earlier work this paper cites.
J. Snoek et al. , “Practical bayesian optimization of machine learning algorithms,” in NeurIPS , 2012
2012
Earlier work this paper cites.
2013
Earlier work this paper cites.
T. Kulesza et al. , “Too much, too little, or just right? ways explanations impact end users’ mental models,” in VL/HCC , 2013
2013
Earlier work this paper cites.
J. Yosinski et al. , “How transferable are features in deep neural networks?” in NeurIPS , 2014
2014
Earlier work this paper cites.
D. Silver et al. , “Deterministic policy gradient algorithms,” in ICML , 2014
2014
Earlier work this paper cites.
F. Gedikli et al. , “How should i explain? a comparison of different explanation types for recommender systems,” IJHCI , 2014
2014
Earlier work this paper cites.
F. Nothdurft et al. , “Probabilistic human-computer trust handling,” in SIGDIAL , 2014
2014
Earlier work this paper cites.
J. Zheng et al. , “Robust bayesian inverse reinforcement learning with sparse behavior noise,” in AAAI , 2014
2014
Earlier work this paper cites.
V. Mnih et al. , “Human-level control through deep reinforcement learning,” Nature , 2015
2015
Earlier work this paper cites.
M. Hausknecht et al. , “Deep recurrent q-learning for partially observable mdps,” in AAAI , 2015
2015
Earlier work this paper cites.
T. Schaul et al. , “Prioritized experience replay,” arXiv preprint arXiv:1511.05952 , 2015
2015
Earlier work this paper cites.
J. Schulman et al. , “Trust region policy optimization,” in ICML , 2015
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
E. Rader et al. , “Understanding user beliefs about algorithmic curation in the facebook news feed,” in CHI , 2015
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
W. R. Stauffer et al. , “Components and characteristics of the dopamine reward utility signal,” Journal of Comparative Neurology , 2016
2016
Earlier work this paper cites.
Y. Bengio et al. , Deep Learning , 2016
2016
Earlier work this paper cites.
V. Mnih et al. , “Asynchronous methods for deep reinforcement learning,” in ICML , 2016
2016
Earlier work this paper cites.
K. He et al. , “Deep residual learning for image recognition,” in CVPR , 2016
2016
Earlier work this paper cites.
T. Zahavy et al. , “Graying the black box: Understanding dqns,” in ICML , 2016
2016
Earlier work this paper cites.
H. Van Hasselt et al. , “Deep reinforcement learning with double q-learning,” in AAAI , 2016
2016
Earlier work this paper cites.
Z. Wang et al. , “Dueling network architectures for deep reinforcement learning,” in ICML , 2016
2016
Earlier work this paper cites.
B. Kim et al. , “Examples are not enough, learn to criticize! criticism for interpretability,” in NeurIPS , 2016
2016
Earlier work this paper cites.
M. Kay et al. , “When (ish) is my bus? user-centered visualizations of uncertainty in everyday, mobile predictive systems,” in CHI , 2016
2016
Earlier work this paper cites.
M. T. Ribeiro et al. , “” why should i trust you?” explaining the predictions of any classifier,” in KDD , 2016
2016
Earlier work this paper cites.
H. Lakkaraju et al. , “Interpretable decision sets: A joint framework for description and prediction,” in KDD , 2016
2016
Earlier work this paper cites.
A. Binder et al. , “Analyzing and validating neural networks predictions,” in ICML Workshop , 2016
2016
Earlier work this paper cites.
Q. Lu et al. , “Using genetic programming with prior formula knowledge to solve symbolic regression problem,” Computational Intelligence and Neuroscience , vol. 2016, 2016
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
V. Sze et al. , “Efficient processing of deep neural networks: A tutorial and survey,” Proc. IEEE , 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
D. Silver et al. , “Mastering the game of go without human knowledge,” Nature , 2017
2017
Earlier work this paper cites.
P. Wang et al. , “Formulation of deep reinforcement learning architecture toward autonomous driving for on-ramp merge,” in ITSC , 2017
2017
Earlier work this paper cites.
B. Hayes et al. , “Improving robot controller transparency through autonomous policy explanation,” in HRI , 2017
2017
Earlier work this paper cites.
M. G. Bellemare and Dothers, “A distributional perspective on reinforcement learning,” in ICML , 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
S. Berkovsky et al. , “How to recommend? user trust factors in movie recommender systems,” in IUI , 2017
2017
Earlier work this paper cites.
M. Sundararajan et al. , “Axiomatic attribution for deep networks,” in ICML , 2017
2017
Earlier work this paper cites.
D. Hein et al. , “Particle swarm optimization for generating interpretable fuzzy reinforcement learning policies,” EAAI , 2017
2017
Earlier work this paper cites.
Y. Fukuchi et al. , “Autonomous self-explanation of behavior for interactive reinforcement learning agents,” in HAI , 2017
2017
Earlier work this paper cites.
——, “Application of instruction-based behavior explanation to a reinforcement learning agent with changing policy,” in NeurIPS , 2017
2017
Earlier work this paper cites.
C. Wirth et al. , “A survey of preference-based reinforcement learning methods,” JMLR , 2017
2017
Earlier work this paper cites.
P. F. Christiano et al. , “Deep reinforcement learning from human preferences,” in NeurIPS , 2017
2017
Earlier work this paper cites.
R. S. Sutton et al. , Reinforcement learning: An introduction . MIT press, 2018
2018
Earlier work this paper cites.
T. Haarnoja et al. , “Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor,” in ICML , 2018
2018
Earlier work this paper cites.
S. Fujimoto et al. , “Addressing function approximation error in actor-critic methods,” in ICML , 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
A. R. Fayjie et al. , “Driverless car: Autonomous driving using deep reinforcement learning in urban environment,” in UR , 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
A. Rosenfeld et al. , “Leveraging human knowledge in tabular reinforcement learning: A study of human subjects,” KER , 2018
2018
Earlier work this paper cites.
R. Iyer et al. , “Transparency and explanation in deep reinforcement learning neural networks,” in AIES , 2018
2018
Earlier work this paper cites.
S. Greydanus et al. , “Visualizing and understanding atari agents,” in ICML , 2018
2018
Earlier work this paper cites.
B. RichardWebster et al. , “Visual psychophysics for making face recognition algorithms more explainable,” in ECCV , 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
G. Liu et al. , “Toward interpretable deep reinforcement learning with linear model u-trees,” in ECML PKDD , 2018
2018
Earlier work this paper cites.
M. Hessel et al. , “Rainbow: Combining improvements in deep reinforcement learning,” in AAAI , 2018
2018
Earlier work this paper cites.
T. Shu et al. , “Hierarchical and interpretable skill acquisition in multi-task reinforcement learning,” in ICLR , 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
T. Ribeiro et al. , “Anchors: High-precision model-agnostic explanations,” in AAAI , 2018
2018
Earlier work this paper cites.
B. Nushi, E. Kamar, and E. Horvitz, “Towards accountable ai: Hybrid human-machine analyses for characterizing system failure,” in HCOMP , 2018
2018
Earlier work this paper cites.
B. Kim et al. , “Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (tcav),” in ICML , 2018
2018
Earlier work this paper cites.
S. Coppers et al. , “Intellingo: An intelligible translation environment,” in CHI , 2018
2018
Earlier work this paper cites.
Z. Yang et al. , “Learn to interpret atari agents,” arXiv preprint arXiv:1812.11276 , 2018
2018
Earlier work this paper cites.
A. Verma et al. , “Programmatically interpretable reinforcement learning,” in ICML , 2018
2018
Cited alongside, same era.
O. Bastani et al. , “Verifiable reinforcement learning via policy extraction,” in NeurIPS , 2018
2018
Cited alongside, same era.
2018
Cited alongside, same era.
——, “Sanity checks for saliency maps,” in NeurIPS , 2018
2018
Cited alongside, same era.
D. Hein et al. , “Interpretable policies for reinforcement learning by genetic programming,” EAAI , 2018
2018
Cited alongside, same era.
J. Foerster et al. , “Counterfactual multi-agent policy gradients,” in AAAI , 2018
B. Škrlj et al. , “autobot: evolving neuro-symbolic representations for explainable low resource text classification,” ML , 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
A. Heuillet et al. , “Explainability in deep reinforcement learning,” KBS , 2021
2021
Later among the works it cites.
L. Wells et al. , “Explainable ai and reinforcement learning—a systematic review of current approaches and trends,” Frontiers in Artificial Intelligence , 2021
2021
Later among the works it cites.
2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2018
Cited alongside, same era.
I. Mishra et al. , “Visual sparse bayesian reinforcement learning: a framework for interpreting what an agent has learned,” in SSCI , 2018
2018
Cited alongside, same era.
V. Goel et al. , “Unsupervised video object segmentation for deep reinforcement learning,” in NeurIPS , 2018
2018
Cited alongside, same era.
V. Petsiuk et al. , “Rise: Randomized input sampling for explanation of black-box models,” in BMVC , 2018
2018
Cited alongside, same era.
2018
Cited alongside, same era.
J. Kim et al. , “Textual explanations for self-driving vehicles,” in ECCV , 2018
2018
Cited alongside, same era.
Y. Li et al. , “In the eye of beholder: Joint learning of gaze and actions in first person video,” in ECCV , 2018
2018
Cited alongside, same era.
Later among the works it cites.
A. Chraibi et al. , “Makespan optimisation in cloudlet scheduling with improved DQN algorithm in cloud computing,” Scientific Programming , 2021
2021
Later among the works it cites.
T. Huber et al. , “Local and global explanations of agent behavior: Integrating strategy summaries with saliency maps,” AI , 2021
2021
Later among the works it cites.
S. Mohseni et al. , “A multidisciplinary survey and framework for design and evaluation of explainable ai systems,” TiiS , 2021
2021
Later among the works it cites.
Y. Tang et al. , “The sensory neuron as a transformer: Permutation-invariant neural networks for reinforcement learning,” in NeurIPS , 2021
2021
Later among the works it cites.
N. Topin et al. , “Iterative bounding mdps: Learning interpretable policies via non-interpretable methods,” in AAAI , 2021
2021
Later among the works it cites.
W. Guo et al. , “Edge: Explaining deep reinforcement learning policies,” in NeurIPS , 2021
2021
Later among the works it cites.
D. Trivedi et al. , “Learning to synthesize programs as interpretable and generalizable policies,” in NeurIPS , 2021
2021
Later among the works it cites.
M. Landajuela et al. , “Discovering symbolic policies with deep reinforcement learning,” in ICML , 2021
2021
Later among the works it cites.
M. Zimmer et al. , “Differentiable logic machines,” arXiv preprint arXiv:2102.11529 , 2021
2021
Later among the works it cites.
M. L. Olson et al. , “Counterfactual state explanations for reinforcement learning agents via generative deep learning,” AI , 2021
2021
Later among the works it cites.
G. Stein, “Generating high-quality explanations for navigation in partially-revealed environments,” in NeurIPS , 2021
2021
Later among the works it cites.
J. Li et al. , “Shapley counterfactual credits for multi-agent reinforcement learning,” in KDD , 2021
2021
Later among the works it cites.
H. Wu et al. , “Self-supervised attention-aware reinforcement learning,” in AAAI , 2021
2021
Later among the works it cites.
S. Mirchandani et al. , “Ella: Exploration through learned language abstraction,” in NeurIPS , 2021
2021
Later among the works it cites.
K. Zhang et al. , “Explainable ai in deep reinforcement learning models for power system emergency control,” TCSS , 2021
2021
Later among the works it cites.
S. S. Guo et al. , “Machine versus human attention in deep reinforcement learning tasks,” in NeurIPS , 2021
2021
Later among the works it cites.
S. Sodhani et al. , “Multi-task reinforcement learning with context-based representations,” in ICML , 2021
2021
Later among the works it cites.
V. Chen et al. , “Ask your humans: Using human instructions to improve generalization in reinforcement learning,” in ICLR , 2021
2021
Later among the works it cites.
B. Ghai et al. , “Explainable active learning (xal) toward ai explanations as interfaces for machine teachers,” HCI , 2021
2021
Later among the works it cites.
Y. Shen et al. , “Autopreview: A framework for autopilot behavior understanding,” in CHI , 2021
2021
Later among the works it cites.
D. Franco et al. , “Deep fair models for complex data: Graphs labeling and explainable face recognition,” Neurocomputing , 2022
2022
Closest in time.
G. A. Vouros, “Explainable deep reinforcement learning: State of the art and challenges,” CSUR , 2022
2022
Closest in time.
D. Garg et al. , “Lisa: Learning interpretable skill abstractions from language,” in NeurIPS , 2022
2022
Closest in time.
Q. Chen et al. , “Es-dqn: A learning method for vehicle intelligent speed control strategy under uncertain cut-in scenario,” TVT , 2022
2022
Closest in time.
2022
Closest in time.
L. Chen et al. , “Conditional dqn-based motion planning with fuzzy logic for autonomous driving,” TITS , 2022
2022
Closest in time.
C. Liu et al. , “Forecasting the market with machine learning algorithms: An application of NMC-BERT-LSTM-DQN-X algorithm in quantitative trading,” KDD , 2022
2022
Closest in time.
S. Park et al. , “Applying DQN solutions in fog-based vehicular networks: Scheduling, caching, and collision control,” Vehicular Communications , 2022
2022
Closest in time.
A. Vashist et al. , “Dqn based exit selection in multi-exit deep neural networks for applications targeting situation awareness,” in ICCE , 2022
2022
Closest in time.
B. Fernández-Gauna et al. , “Actor-critic continuous state reinforcement learning for wind-turbine control robust optimization,” Information Sciences , 2022
2022
Closest in time.
X. Gong et al. , “Actor-critic with familiarity-based trajectory experience replay,” Information Sciences , vol. 582, pp. 633–647, 2022
2022
Closest in time.
X. Xin et al. , “Supervised advantage actor-critic for recommender systems,” in CIKM , 2022
2022
Closest in time.
L. L. Custode et al. , “Evolutionary learning of interpretable decision trees,” in IJCAI , 2022
2022
Closest in time.
2022
Closest in time.
Y. Amitai and O. Amir, ““i don’t think so”: Summarizing policy disagreements for agent comparison,” in AAAI , 2022
2022
Closest in time.
2022
Closest in time.
P. Jin et al. , “Trainify: A cegar-driven training and verification framework for safe deep reinforcement learning,” in CAV , 2022
2022
Closest in time.
M. Jin et al. , “Creativity of ai: Automatic symbolic option discovery for facilitating deep reinforcement learning,” in AAAI , 2022
2022
Closest in time.
Z. Ashwood et al. , “Dynamic inverse reinforcement learning for characterizing animal behavior,” in NeurIPS , 2022
2022
Closest in time.
A. Heuillet et al. , “Collective explainable ai: Explaining cooperative strategies and agent contribution in multiagent reinforcement learning with shapley values,” IEEE Computational Intelligence Magazine , 2022
2022
Closest in time.
E. M. Kenny et al. , “Towards interpretable deep reinforcement learning with human-friendly prototypes,” in ICLR , 2022
2022
Closest in time.
R. Ragodos et al. , “Protox: Explaining a reinforcement learning agent via prototyping,” in NeurIPS , 2022
2022
Closest in time.
S. Wäldchen et al. , “Training characteristic functions with reinforcement learning: Xai-methods play connect four,” in ICML , 2022
2022
Closest in time.
D. Bertoin et al. , “Look where you look! saliency-guided q-networks for generalization in visual reinforcement learning,” in NeurIPS , 2022
2022
Closest in time.
X. Peng et al. , “Inherently explainable reinforcement learning in natural language,” in NeurIPS , 2022
2022
Closest in time.
V. Zhong et al. , “Improving policy learning via language dynamics distillation,” in NeurIPS , 2022
2022
Closest in time.
T. Rudolf et al. , “Fuzzy action-masked reinforcement learning behavior planning for highly automated driving,” in CCAR , 2022
2022
Closest in time.
W. Shi et al. , “Efficient hierarchical policy network with fuzzy rules,” IJMLC , 2022
2022
Closest in time.
2022
Closest in time.
——, “Interpretable machine learning: Fundamental principles and 10 grand challenges,” Statistic Surveys , 2022
2022
Closest in time.
S. Liu et al. , “Contrastive identity-aware learning for multi-agent value decomposition,” in AAAI , 2023
2023
Closest in time.
2023
Closest in time.
S. Liu et al. , “Transmission interface power flow adjustment: A deep reinforcement learning approach based on multi-task attribution map,” IEEE TPS , 2023
2023
Closest in time.
S. Milani et al. , “Explainable reinforcement learning: A survey and comparative review,” CSUR , 2023
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
L. Ou et al. , “Fuzzy centered explainable network for reinforcement learning,” TFS , 2023
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
D. Beechey et al. , “Explaining reinforcement learning with shapley values,” in ICML , 2023
2023
Closest in time.
B. Hu et al. , “An explainable and robust motion planning and control approach for autonomous vehicle on-ramping merging task using deep reinforcement learning,” TTE , 2023
2023
Closest in time.
R. Luss et al. , “Local explanations for reinforcement learning,” in IJCAI , 2023
2023
Closest in time.
G. Zhang et al. , “Learning state importance for preference-based reinforcement learning,” ML , 2023
2023
Closest in time.
Y. Qing, S. Liu, J. Cong, K. Chen, Y. Zhou, and M. Song, “A2po: Towards effective offline reinforcement learning from an advantage-aware perspective,” Advances in Neural Information Processing Systems , vol. 37, pp. 29 064–29 090, 2024
2024
Closest in time.
2024
Closest in time.
F. Xu et al. , “Temporal prototype-aware learning for active voltage control on power distribution networks,” in KDD , 2024
2024
Closest in time.
——, “Progressive decision-making framework for power system topology control,” ESWA , 2024
2024
Closest in time.
V. Ivchyk, “Overcoming barriers to artificial intelligence adoption,” Three Seas Economic Journal , 2024
2024
Closest in time.
Q. Delfosse et al. , “Interpretable and explainable logical policies via neurally guided symbolic abstraction,” NeurIPS , 2024
2024
Closest in time.
A. Meulemans et al. , “Would i have gotten that reward? long-term credit assignment by counterfactual contribution analysis,” NeurIPS , 2024
2024
Closest in time.
Y. Amitai et al. , “Explaining reinforcement learning agents through counterfactual action outcomes,” in AAAI , 2024
2024
Closest in time.
H. Sun et al. , “Accountability in offline reinforcement learning: Explaining decisions with a corpus of examples,” NeurIPS , 2024
2024
Closest in time.
C. Wang et al. , “Explainable deep adversarial reinforcement learning approach for robust autonomous driving,” TIV , 2024
2024
Closest in time.
K. Lee et al. , “Refining diffusion planner for reliable behavior synthesis by automatic detection of infeasible plans,” NeurIPS , 2024
2024
Closest in time.