Fetching the paper…
Reading the bibliography…
Deep reinforcement learning (DRL) has demonstrated impressive performance in various gaming simulators and real-world applications.
Adversarial training can hurt generalization
Raghunathan, A.; Xie, S. M.; Yang, F.; Duchi, J. C.; and Liang, P. 2019 · 1906
Earlier work this paper cites.
Behaviour suite for reinforcement learning
Osband, I.; Doron, Y.; Hessel, M.; Aslanides, J.; Sezener, E.; Saraiva, A.; McKinney, K.; Lattimore, T.; Szepezvari, C.; Singh, S.; et al. 2019 · 1908
Earlier work this paper cites.
Off-Policy Evaluation in Partially Observable Environments
Tennenholtz, G.; Mannor, S.; and Shalit, U. 2019 · 1909
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
Mnih, V.; Badia, A. P.; Mirza, M.; Graves, A.; Lillicrap, T.; Harley, T.; Silver, D.; and Kavukcuoglu, K. 2016 · 1937
Earlier work this paper cites.
Estimating causal effects of treatments in randomized and nonrandomized studies
Rubin, D. B. 1974 · 1974
Earlier work this paper cites.
The complexity of Markov decision processes
Papadimitriou, C. H.; and Tsitsiklis, J. N. 1987 · 1987
Earlier work this paper cites.
Reinforcement learning in Markovian and non-Markovian environments
Schmidhuber, J. 1991 · 1991
Earlier work this paper cites.
Learning complex, extended sequences using the principle of history compression
Schmidhuber, J. 1992 · 1992
Earlier work this paper cites.
Analysis of semiparametric regression models for repeated outcomes in the presence of missing data
Robins, J. M.; Rotnitzky, A.; and Zhao, L. P. 1995 · 1995
Earlier work this paper cites.
Digital control of dynamic systems , volume 3
Franklin, G. F.; Powell, J. D.; Workman, M. L.; et al. 1998 · 1998
Earlier work this paper cites.
Planning and acting in partially observable stochastic domains
Kaelbling, L. P.; Littman, M. L.; and Cassandra, A. R. 1998 · 1998
Earlier work this paper cites.
Reinforcement learning: An introduction
Sutton, R. S.; Barto, A. G.; Bach, F.; et al. 1998 · 1998
Earlier work this paper cites.
Causal diagrams for epidemiologic research
Greenland, S.; Pearl, J.; and Robins, J. M. 1999 · 1999
Earlier work this paper cites.
Causation and causal inference in epidemiology
Rothman, K. J.; and Greenland, S. 2005 · 2005
Earlier work this paper cites.
Counterfactually Guided Policy Transfer in Clinical Settings
Killian, T. W.; Ghassemi, M.; and Joshi, S. 2020 · 2006
Earlier work this paper cites.
Off-policy evaluation in infinite-horizon reinforcement learning with latent confounders
Bennett, A.; Kallus, N.; Li, L.; and Mousavi, A. 2021 · 2007
Earlier work this paper cites.
Improving Fair Predictions Using Variational Inference In Causal Models
Helwegen, R.; Louizos, C.; and Forré, P. 2020 · 2008
Earlier work this paper cites.
Causality
Pearl, J. 2009 · 2009
Earlier work this paper cites.
Rubin causal model
Imbens, G. W.; and Rubin, D. B. 2010 · 2010
Earlier work this paper cites.
Multi-task Language Modeling for Improving Speech Recognition of Rare Words
Yang, C.-H. H.; Liu, L.; Gandhe, A.; Gu, Y.; Raju, A.; Filimonov, D.; and Bulyko, I. 2020a · 2011
Earlier work this paper cites.
Fast reinforcement learning with large action sets using error-correcting output codes for mdp factorization
Dulac-Arnold, G.; Denoyer, L.; Preux, P.; and Gallinari, P. 2012 · 2012
Earlier work this paper cites.
Deep learning of representations: Looking forward
Bengio, Y. 2013 · 2013
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. P.; and Welling, M. 2013 · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P.; and Ba, J. 2014 · 2014
Earlier work this paper cites.
Intriguing properties of neural networks
Szegedy, C.; Zaremba, W.; Sutskever, I.; Bruna, J.; Erhan, D.; Goodfellow, I.; and Fergus, R. 2014 · 2014
Earlier work this paper cites.
Bandits with unobserved confounders: A causal approach
Bareinboim, E.; Forney, A.; and Pearl, J. 2015 · 2015
Earlier work this paper cites.
Taming the noise in reinforcement learning via soft updates
Fox, R.; Pakman, A.; and Tishby, N. 2015 · 2015
Earlier work this paper cites.
Explaining and harnessing adversarial examples
Goodfellow, I. J.; Shlens, J.; and Szegedy, C. 2015 · 2015
Earlier work this paper cites.
On estimation of wind velocity, angle-of-attack and sideslip angle of small UAVs using standard sensors
Johansen, T. A.; Cristofaro, A.; Sørensen, K.; Hansen, J. M.; and Fossen, T. I. 2015 · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Mnih, V.; Kavukcuoglu, K.; Silver, D.; Rusu, A. A.; Veness, J.; Bellemare, M. G.; Graves, A.; Riedmiller, M.; Fidjeland, A. K.; Ostrovski, G.; et al. 2015 · 2015
Cited alongside, same era.
Brockman, G.; Cheung, V.; Pettersson, L.; Schneider, J.; Schulman, J.; Tang, J.; and Zaremba, W. 2016 · 2016
Cited alongside, same era.
Q lamda with Off-Policy Corrections
Harutyunyan, A.; Bellemare, M. G.; Stepleton, T.; and Munos, R. 2016 · 2016
Cited alongside, same era.
Learning to navigate in complex environments
Mirowski, P.; Pascanu, R.; Viola, F.; Soyer, H.; Ballard, A. J.; Banino, A.; Denil, M.; Goroshin, R.; Sifre, L.; Kavukcuoglu, K.; et al. 2016 · 2016
Unity: A general platform for intelligent agents
Juliani, A.; Berges, V.-P.; Vckay, E.; Gao, Y.; Henry, H.; Mattar, M.; and Lange, D. 2018 · 2018
Later among the works it cites.
Qt-opt: Scalable deep reinforcement learning for vision-based robotic manipulation
Kalashnikov, D.; Irpan, A.; Pastor, P.; Ibarz, J.; Herzog, A.; Jang, E.; Quillen, D.; Holly, E.; Kalakrishnan, M.; Vanhoucke, V.; et al. 2018 · 2018
Later among the works it cites.
Bayesian policy optimization for model uncertainty
Lee, G.; Hou, B.; Mandalika, A.; Lee, J.; Choudhury, S.; and Srinivasa, S. S. 2018 · 2018
Later among the works it cites.
Deconfounding reinforcement learning in observational settings
Lu, C.; Schölkopf, B.; and Hernández-Lobato, J. M. 2018 · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Causal inference in statistics: A primer
Pearl, J.; Glymour, M.; and Jewell, N. P. 2016 · 2016
Cited alongside, same era.
Deep reinforcement learning with double q-learning
Van Hasselt, H.; Guez, A.; and Silver, D. 2016 · 2016
Cited alongside, same era.
Can you trust autonomous vehicles: Contactless attacks against sensors of self-driving vehicle
Yan, C.; Xu, W.; and Liu, J. 2016 · 2016
Cited alongside, same era.
OpenAI Baselines
Dhariwal, P.; Hesse, C.; Klimov, O.; Nichol, A.; Plappert, M.; Radford, A.; Schulman, J.; Sidor, S.; Wu, Y.; and Zhokhov, P. 2017 · 2017
Cited alongside, same era.
Counterfactual data-fusion for online reinforcement learners
Forney, A.; Pearl, J.; and Bareinboim, E. 2017 · 2017
Cited alongside, same era.
Deep reinforcement learning for robotic manipulation with asynchronous off-policy updates
Gu, S.; Holly, E.; Lillicrap, T.; and Levine, S. 2017 · 2017
Cited alongside, same era.
Adversarial attacks on neural network policies
Huang, S.; Papernot, N.; Goodfellow, I.; Duan, Y.; and Abbeel, P. 2017 · 2017
Cited alongside, same era.
Moreno, P.; Humplik, J.; Papamakarios, G.; Pires, B. A.; Buesing, L.; Heess, N.; and Weber, T. 2018 · 2018
Later among the works it cites.
Learning to adapt in dynamic, real-world environments through meta-reinforcement learning
Nagabandi, A.; Clavera, I.; Liu, S.; Fearing, R. S.; Abbeel, P.; Levine, S.; and Finn, C. 2018 · 2018
Later among the works it cites.
Is robustness the cost of accuracy?–a comprehensive study on the robustness of 18 deep image classification models
Su, D.; Zhang, H.; Chen, H.; Yi, J.; Chen, P.-Y.; and Gao, Y. 2018 · 2018
Later among the works it cites.
Evaluating the robustness of neural networks: An extreme value theory approach
Weng, T.-W.; Zhang, H.; Chen, P.-Y.; Yi, J.; Su, D.; Gao, Y.; Hsieh, C.-J.; and Daniel, L. 2018 · 2018
Later among the works it cites.
Transfer in Deep Reinforcement Learning Using Knowledge Graphs
Ammanabrolu, P.; and Riedl, M. 2019 · 2019
Later among the works it cites.
Causal identification under markov equivalence: Completeness results
Jaber, A.; Zhang, J.; and Bareinboim, E. 2019 · 2019
Later among the works it cites.
The seven tools of causal inference, with reflections on machine learning
Pearl, J. 2019 · 2019
Later among the works it cites.
DoWhy A Python package for causal inference
Sharma, A.; Kiciman, E.; et al. 2019 · 2019
Later among the works it cites.
When causal intervention meets adversarial examples and image masking for deep neural networks
Yang, C.-H. H.; Liu, Y.-C.; Chen, P.-Y.; Ma, X.; and Tsai, Y.-C. J. 2019 · 2019
Later among the works it cites.
A survey of deep learning techniques for autonomous driving
Grigorescu, S.; Trasnea, B.; Cocias, T.; and Macesanu, G. 2020 · 2020
Later among the works it cites.
Learning latent plans from play
Lynch, C.; Khansari, M.; Xiao, T.; Kumar, V.; Tompson, J.; Levine, S.; and Sermanet, P. 2020 · 2020
Later among the works it cites.
Explainable reinforcement learning through a causal lens
Madumal, P.; Miller, T.; Sonenberg, L.; and Vetere, F. 2020 · 2020
Later among the works it cites.
Enhanced Adversarial Strategically-Timed Attacks Against Deep Reinforcement Learning
Yang, C.-H. H.; Qi, J.; Chen, P.-Y.; Ouyang, Y.; Hung, I.-T. D.; Lee, C.-H.; and Ma, X. 2020b · 2020
Later among the works it cites.
A survey of autonomous driving: Common practices and emerging technologies
Yurtsever, E.; Lambert, J.; Carballo, A.; and Takeda, K. 2020 · 2020
Later among the works it cites.
Invariant causal prediction for block mdps
Zhang, A.; Lyle, C.; Sodhani, S.; Filos, A.; Kwiatkowska, M.; Pineau, J.; Gal, Y.; and Precup, D. 2020 · 2020
Later among the works it cites.
A Causal View on Robustness of Neural Networks
Zhang, C.; Zhang, K.; and Li, Y. 2020 · 2020
Later among the works it cites.
Designing optimal dynamic treatment regimes: A causal reinforcement learning approach
Zhang, J.; and Bareinboim, E. 2020 · 2020
Later among the works it cites.
Causal imitation learning with unobserved confounders
Zhang, J.; Kumor, D.; and Bareinboim, E. 2020 · 2020
Later among the works it cites.
Neural Network Verification in Control
Everett, M. 2021 · 2021
Closest in time.
Estimating identifiable causal effects through double machine learning
Jung, Y.; Tian, J.; and Bareinboim, E. 2021 · 2021
Closest in time.
Causal autoregressive flows
Khemakhem, I.; Monti, R.; Leech, R.; and Hyvarinen, A. 2021 · 2021
Closest in time.
Bounding Causal Effects on Continuous Outcome
Zhang, J.; and Bareinboim, E. 2021 · 2021
Closest in time.
Trial without error: Towards safe reinforcement learning via human intervention
Saunders, W.; Sastry, G.; Stuhlmueller, A.; and Evans, O. 2018 · 2069
Closest in time.