Fetching the paper…
Reading the bibliography…
As AI becomes more prevalent throughout society, effective methods of integrating humans and AI systems that leverage their respective strengths and mitigate risk have become an important priority.
Medical heuristics: the silent adjudicators of clinical practice
McDonald, C. J. (1996) · 1996
Earlier work this paper cites.
Predictive representations of state
Littman, M. and R. S. Sutton (2001) · 2001
Earlier work this paper cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Levine, S., A. Kumar, G. Tucker, and J. Fu (2020) · 2005
Earlier work this paper cites.
Causal inference using potential outcomes: Design, modeling, decisions
Rubin, D. B. (2005) · 2005
Earlier work this paper cites.
Minimax estimation of conditional moment models
Dikkala, N., G. Lewis, L. Mackey, and V. Syrgkanis (2020) · 2006
Earlier work this paper cites.
An introduction to proximal causal learning
Tchetgen Tchetgen, E. J., A. Ying, Y. Cui, X. Shi, and W. Miao (2020) · 2009
Earlier work this paper cites.
Algorithmic Trading & DMA: An introduction to direct access trading strategies
Johnson, B. (2010) · 2010
Earlier work this paper cites.
Hilbert space embeddings of hidden markov models
Song, L., B. Boots, S. Siddiqi, G. J. Gordon, and A. Smola (2010) · 2010
Earlier work this paper cites.
Dynamic reward shaping: training a robot by voice
Tenorio-Gonzalez, A. C., E. F. Morales, and L. Villasenor-Pineda (2010) · 2010
Earlier work this paper cites.
Closing the learning-planning loop with predictive state representations
Boots, B., S. M. Siddiqi, and G. J. Gordon (2011) · 2011
Earlier work this paper cites.
On the completeness condition in nonparametric instrumental problems
D’Haultfoeuille, X. (2011) · 2011
Earlier work this paper cites.
A spectral algorithm for learning hidden Markov models
Hsu, D., S. M. Kakade, and T. Zhang (2012) · 2012
Earlier work this paper cites.
Predictive state representations: A new theory for modeling dynamical systems
Singh, S., M. James, and M. Rudary (2012) · 2012
Earlier work this paper cites.
Policy shaping: Integrating human feedback with reinforcement learning
Griffith, S., K. Subramanian, J. Scholz, C. L. Isbell, and A. L. Thomaz (2013) · 2013
Earlier work this paper cites.
Algorithmic trading review
Treleaven, P., M. Galas, and V. Lalchand (2013) · 2013
Earlier work this paper cites.
Tensor decompositions for learning latent variable models
Anandkumar, A., R. Ge, D. Hsu, S. M. Kakade, and M. Telgarsky (2014) · 2014
Earlier work this paper cites.
Local identification of nonparametric and semiparametric models
Chen, X., V. Chernozhukov, S. Lee, and W. K. Newey (2014) · 2014
Earlier work this paper cites.
Reinforcement learning from demonstration through shaping
Brys, T., A. Harutyunyan, H. B. Suay, S. Chernova, M. E. Taylor, and A. Nowé (2015) · 2015
Earlier work this paper cites.
Delay to admission to critical care and mortality among deteriorating ward patients in uk hospitals: a multicentre, prospective, observational cohort study
Harris, S., M. Singer, K. Rowan, and C. Sanderson (2015) · 2015
Earlier work this paper cites.
Markov decision processes with unobserved confounders: A causal approach
Zhang, J. and E. Bareinboim (2016) · 2016
Earlier work this paper cites.
Learning to teach reinforcement learning agents
Fachantidis, A., M. E. Taylor, and I. Vlahavas (2017) · 2017
Cited alongside, same era.
Trial without error: Towards safe reinforcement learning via human intervention
Saunders, W., G. Sastry, A. Stuhlmueller, and O. Evans (2017) · 2017
Cited alongside, same era.
Human decisions and machine predictions
Kleinberg, J., H. Lakkaraju, J. Leskovec, J. Ludwig, and S. Mullainathan (2018) · 2018
Cited alongside, same era.
Identifying causal effects with proxy variables of an unmeasured confounder
Miao, W., Z. Geng, and E. J. Tchetgen Tchetgen (2018) · 2018
Cited alongside, same era.
A confounding bridge approach for double negative control inference on causal effects
Miao, W., X. Shi, and E. T. Tchetgen (2018) · 2018
Cited alongside, same era.
Estimating and improving dynamic treatment regimes with a time-varying instrumental variable
Chen, S. and B. Zhang (2021) · 2021
Later among the works it cites.
Causal reinforcement learning: An instrumental variable approach
Li, J., Y. Luo, and X. Zhang (2021) · 2021
Later among the works it cites.
Instrumental variable value iteration for causal offline reinforcement learning
Liao, L., Z. Fu, Z. Yang, Y. Wang, M. Kolar, and Z. Wang (2021) · 2021
Later among the works it cites.
A spectral approach to off-policy evaluation for POMDPs
Nair, Y. and N. Jiang (2021) · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Financial risk the fall of knight capital group
Saltapidas, C. and R. Maghsood (2018) · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
Sutton, R. S. and A. G. Barto (2018) · 2018
Cited alongside, same era.
Deep tamer: Interactive agent shaping in high-dimensional state spaces
Warnell, G., N. Waytowich, V. Lawhern, and P. Stone (2018) · 2018
Cited alongside, same era.
Precision medicine
Kosorok, M. R. and E. B. Laber (2019) · 2019
Cited alongside, same era.
Deep brain stimulation: current challenges and future directions
Lozano, A. M., N. Lipsman, H. Bergman, P. Brown, S. Chabardes, J. W. Chang, K. Matthews, C. C. McIntyre, T. E. Schlaepfer, M. Schulder, et al. (2019) · 2019
Cited alongside, same era.
Sample-efficient reinforcement learning of undercomplete pomdps
Jin, C., S. Kakade, A. Krishnamurthy, and Q. Liu (2020) · 2020
Cited alongside, same era.
Confounding-robust policy evaluation in infinite-horizon reinforcement learning
Kallus, N. and A. Zhou (2020) · 2020
Cited alongside, same era.
Shi, C., M. Uehara, and N. Jiang (2021) · 2021
Later among the works it cites.
Provably efficient causal reinforcement learning with confounded observational data
Wang, L., Z. Yang, and Z. Wang (2021) · 2021
Later among the works it cites.
Proximal causal inference for complex longitudinal studies
Ying, A., W. Miao, X. Shi, and E. J. Tchetgen Tchetgen (2021) · 2021
Later among the works it cites.
Sample-efficient reinforcement learning for pomdps with linear function approximations
Cai, Q., Z. Yang, and Z. Wang (2022) · 2022
Closest in time.
Offline reinforcement learning with instrumental variables in confounded markov decision processes
Fu, Z., Z. Qi, Z. Wang, Z. Yang, Y. Xu, and M. R. Kosorok (2022) · 2022
Closest in time.
Finrl-meta: Market environments and benchmarks for data-driven financial reinforcement learning
Liu, X.-Y., Z. Xia, J. Rui, J. Gao, H. Yang, M. Zhu, C. D. Wang, Z. Wang, and J. Guo (2022) · 2022
Closest in time.
Lu, M., Y. Min, Z. Wang, and Z. Yang (2022) · 2022
Closest in time.
Off-policy confidence interval estimation with confounded Markov decision process
Shi, C., J. Zhu, Y. Shen, S. Luo, H. Zhu, and R. Song (2022) · 2022
Closest in time.
Optimal regimes for algorithm-assisted human decision-making
Stensrud, M. J. and A. L. Sarvet (2022) · 2022
Closest in time.
Future-dependent value-based off-policy evaluation in pomdps
Uehara, M., H. Kiyohara, A. Bennett, V. Chernozhukov, N. Jiang, N. Kallus, C. Shi, and W. Sun (2022) · 2022
Closest in time.
Provably efficient reinforcement learning in partially observable dynamical systems
Uehara, M., A. Sekhari, J. D. Lee, N. Kallus, and W. Sun (2022) · 2022
Closest in time.
A review of off-policy evaluation in reinforcement learning
Uehara, M., C. Shi, and N. Kallus (2022) · 2022
Closest in time.
Yu, M., Z. Yang, and J. Fan (2022) · 2022
Closest in time.
Proximal learning for individualized treatment regimes under unmeasured confounding
Qi, Z., R. Miao, and X. Zhang (2023) · 2023
Closest in time.
An instrumental variable approach to confounded off-policy evaluation
Xu, Y., J. Zhu, C. Shi, S. Luo, and R. Song (2023) · 2023
Closest in time.
Two-way deconfounder for off-policy evaluation in causal reinforcement learning
Yu, S., S. Fang, R. Peng, Z. Qi, F. Zhou, and C. Shi (2024) · 2024
Closest in time.