Fetching the paper…
Reading the bibliography…
Observational longitudinal studies are a common means to study treatment efficacy and safety in chronic mental illness.
Hyperbolic discounting and learning over multiple horizons
Fedus, W., C. Gelada, Y. Bengio, M. G. Bellemare, and H. Larochelle (2019) · 1902
Earlier work this paper cites.
Off-policy estimation of long-term average outcomes with applications to mobile health
Liao, P., P. Klasnja, and S. Murphy (2019) · 1912
Earlier work this paper cites.
Tree-based reinforcement learning for estimating optimal dynamic treatment regimes
Tao, Y., L. Wang, and D. Almirall (2018) · 1914
Earlier work this paper cites.
Estimating the infinitesimal generator of a continuous time, finite state markov process
Albert, A. (1962) · 1962
Earlier work this paper cites.
Duality theorem in markovian decision problems
Yamada, K. (1975) · 1975
Earlier work this paper cites.
State of the art—a survey of partially observable markov decision processes: theory, models, and algorithms
Monahan, G. E. (1982) · 1982
Earlier work this paper cites.
Markov decision processes with a borel measurable cost function?the average case
Kurano, M. (1986) · 1986
Earlier work this paper cites.
Necessary and sufficient conditions for a bounded solution to the optimality equation in average reward markov decision chains
Cavazos-Cadena, R. (1988) · 1988
Earlier work this paper cites.
A tutorial on hidden markov models and selected applications in speech recognition
Rabiner, L. R. (1989) · 1989
Earlier work this paper cites.
Recurrence conditions for markov decision processes with borel state space: a survey
Hernández-Lerma, O., R. Montes-de Oca, and R. Cavazos-Cadena (1991) · 1991
Earlier work this paper cites.
Maximum-likelihood estimation for hidden markov models
Leroux, B. G. (1992) · 1992
Earlier work this paper cites.
P values maximized over a confidence set for the nuisance parameter
Berger, R. L. and D. D. Boos (1994) · 1994
Earlier work this paper cites.
Weak convergence
van der Vaart, A. W. and J. A. Wellner (1996) · 1996
Earlier work this paper cites.
Asymptotic normality of the maximum-likelihood estimator for general hidden markov models
Bickel, P. J., Y. Ritov, and T. Ryden (1998) · 1998
Earlier work this paper cites.
A survey of pomdp applications
Cassandra, A. R. (1998) · 1998
Earlier work this paper cites.
Solving pomdps by searching in policy space
Hansen, E. A. (1998) · 1998
Earlier work this paper cites.
Planning and acting in partially observable stochastic domains
Kaelbling, L. P., M. L. Littman, and A. R. Cassandra (1998) · 1998
Earlier work this paper cites.
Asymptotic normality of the maximum likelihood estimator in state space models
Jensen, J. L. and N. V. Petersen (1999) · 1999
Earlier work this paper cites.
Testing and estimation of direct effects by reparameterizing directed acyclic graphs with structural nested models
Robins, J. M. (1999) · 1999
Earlier work this paper cites.
Exponential forgetting and geometric ergodicity in hidden markov models
Le Gland, F. and L. Mevel (2000) · 2000
Earlier work this paper cites.
Asymptotics of the maximum likelihood estimator for general hidden markov models
Douc, R. and C. Matias (2001) · 2001
Earlier work this paper cites.
Marginal mean models for dynamic regimes
Murphy, S. A., M. J. van der Laan, J. M. Robins, and C. P. P. R. Group (2001) · 2001
Earlier work this paper cites.
Point-based value iteration: An anytime algorithm for pomdps
J Pineau, G Gordon, S. T. (2003) · 2003
Earlier work this paper cites.
Optimal dynamic treatment regimes
Murphy, S. A. (2003) · 2003
Earlier work this paper cites.
Rationale, design, and methods of the systematic treatment enhancement program for bipolar disorder (step-bd)
Sachs, G. S., M. E. Thase, M. W. Otto, M. Bauer, D. Miklowitz, S. R. Wisniewski, P. Lavori, B. Lebowitz, M. Rudorfer, E. Frank, et al. (2003) · 2003
Earlier work this paper cites.
Asymptotic properties of the maximum likelihood estimator in autoregressive models with markov regime
Douc, R., E. Moulines, and T. Rydén (2004) · 2004
Earlier work this paper cites.
Optimal structural nested models for optimal sequential decisions
Robins, J. M. (2004) · 2004
Earlier work this paper cites.
Multiple imputation for nonresponse in surveys
Rubin, D. B. (2004) · 2004
Earlier work this paper cites.
A generalization error for q-learning
Murphy, S. A. (2005) · 2005
Earlier work this paper cites.
Measure theory and probability theory
Athreya, K. B. and S. N. Lahiri (2006) · 2006
Earlier work this paper cites.
Causal effect models for intention to treat and realistic individualized treatment rules
van der Laan, M. J. (2006) · 2006
Cited alongside, same era.
Point-based policy iteration
Ji, S., R. Parr, H. Li, X. Liao, and L. Carin (2007) · 2007
Cited alongside, same era.
Demystifying optimal dynamic treatment regimes
Moodie, E. E., T. S. Richardson, and D. A. Stephens (2007) · 2007
Cited alongside, same era.
Approximate Dynamic Programming: Solving the curses of dimensionality
Powell, W. B. (2007) · 2007
Cited alongside, same era.
Who will benefit from antidepressants in the acute treatment of bipolar depression? a reanalysis of the step-bd study by sachs et al. 2007, using q-learning
Wu, F., E. B. Laber, I. A. Lipkovich, and E. Severus (2015) · 2007
Cited alongside, same era.
Introduction to empirical processes and semiparametric inference
Bipolar disorder dynamics: affective instabilities, relaxation oscillations and noise
Bonsall, M. B., J. R. Geddes, G. M. Goodwin, and E. A. Holmes (2015) · 2015
Later among the works it cites.
Treatment decisions based on scalar and functional baseline covariates
Ciarleglio, A., E. Petkova, R. T. Ogden, and T. Tarpey (2015) · 2015
Later among the works it cites.
Doubly robust off-policy value evaluation for reinforcement learning
Jiang, N. and L. Li (2015) · 2015
Later among the works it cites.
Tree-based methods for individualized treatment regimes
Laber, E. and Y. Zhao (2015) · 2015
Later among the works it cites.
Efficient learning of continuous-time hidden markov models for disease progression
Liu, Y.-Y., S. Li, F. Li, L. Song, and J. M. Rehg (2015) · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kosorok, M. R. (2008) · 2008
Cited alongside, same era.
Estimation and extrapolation of optimal treatment and testing strategies
Robins, J., L. Orellana, and A. Rotnitzky (2008) · 2008
Cited alongside, same era.
Identifiability of parameters in latent structure models with many observed variables
Allman, E. S., C. Matias, J. A. Rhodes, et al. (2009) · 2009
Cited alongside, same era.
Mathematical models of bipolar disorder
Daugherty, D., T. Roque-Urrea, J. Urrea-Roque, J. Troyer, S. Wirkus, and M. A. Porter (2009) · 2009
Cited alongside, same era.
Reinforcement learning design for cancer clinical trials
Zhao, Y., M. R. Kosorok, and D. Zeng (2009) · 2009
Cited alongside, same era.
Regret-regression for optimal dynamic treatment regimes
Henderson, R., P. Ansell, and D. Alshibani (2010) · 2010
Cited alongside, same era.
Algorithms for reinforcement learning
Szepesvári, C. (2010) · 2010
Cited alongside, same era.
Do antidepressants increase the risk of mania and bipolar disorder in people with depression? a retrospective electronic case register cohort study
Patel, R., P. Reiss, H. Shetty, M. Broadbent, R. Stewart, P. McGuire, and M. Taylor (2015) · 2015
Later among the works it cites.
Penalized q-learning for dynamic treatment regimens
Song, R., W. Wang, D. Zeng, and M. R. Kosorok (2015) · 2015
Later among the works it cites.
New statistical learning methods for estimating optimal dynamic treatment regimes
Zhao, Y.-Q., D. Zeng, E. B. Laber, and M. R. Kosorok (2015) · 2015
Later among the works it cites.
Recursive partitioning for heterogeneous causal effects
Athey, S. and G. Imbens (2016) · 2016
Later among the works it cites.
Flexible functional regression methods for estimating individualized treatment rules
Ciarleglio, A., E. Petkova, T. Tarpey, and R. T. Ogden (2016) · 2016
Later among the works it cites.
Applications of time-series analysis to mood fluctuations in bipolar disorder to promote treatment innovation: a case series
Holmes, E., M. Bonsall, S. Hales, H. Mitchell, F. Renner, S. Blackwell, P. Watson, G. Goodwin, and M. Di Simplicio (2016) · 2016
Later among the works it cites.
Super-learning of an optimal dynamic treatment rule
Luedtke, A. R. and M. J. van der Laan (2016) · 2016
Later among the works it cites.
A batch, off-policy, actor-critic algorithm for optimizing the average reward
Murphy, S. A., Y. Deng, E. B. Laber, H. R. Maei, R. S. Sutton, and K. Witkiewitz (2016) · 2016
Later among the works it cites.
Data-efficient off-policy policy evaluation for reinforcement learning
Thomas, P. and E. Brunskill (2016) · 2016
Later among the works it cites.
Bayesian nonparametric estimation for dynamic treatment regimes with sequential transition times
Xu, Y., P. Müller, A. S. Wahed, and P. F. Thall (2016) · 2016
Later among the works it cites.
Reinforcement learning and dynamic programming using function approximators
Busoniu, L., R. Babuska, B. De Schutter, and D. Ernst (2017) · 2017
Later among the works it cites.
Functional feature construction for individualized treatment regimes
Laber, E. B. and A.-M. Staicu (2017) · 2017
Later among the works it cites.
From ads to interventions: Contextual bandits in mobile health
Tewari, A. and S. A. Murphy (2017) · 2017
Later among the works it cites.
Importance sampling policy evaluation with an estimated behavior policy
Hanna, J. P., S. Niekum, and P. Stone (2018) · 2018
Later among the works it cites.
High-dimensional a-learning for optimal dynamic treatment regimes
Shi, C., A. Fan, R. Song, and W. Lu (2018) · 2018
Later among the works it cites.
Reinforcement learning: An introduction
Sutton, R. S., A. G. Barto, et al. (2018) · 2018
Later among the works it cites.
Estimation and inference of heterogeneous treatment effects using random forests
Wager, S. and S. Athey (2018) · 2018
Later among the works it cites.
Interpretable dynamic treatment regimes
Zhang, Y., E. B. Laber, M. Davidian, and A. A. Tsiatis (2018) · 2018
Later among the works it cites.
Constructing dynamic treatment regimes in infinite-horizon settings
Ertefaie, A. (2019) · 2019
Later among the works it cites.
Entropy learning for dynamic treatment regimes
Jiang, B., R. Song, J. Li, and D. Zeng (2019) · 2019
Later among the works it cites.
Learning the dynamic treatment regimes from medical registry data through deep q-network
Liu, N., Y. Liu, B. Logan, Z. Xu, J. Tang, and Y. Wang (2019) · 2019
Later among the works it cites.
Estimating dynamic treatment regimes in mobile health using v-learning
Luckett, D. J., E. B. Laber, A. R. Kahkoska, D. M. Maahs, E. Mayer-Davis, and M. R. Kosorok (2019) · 2019
Later among the works it cites.
A sparse random projection-based test for overall qualitative treatment effects
Shi, C., W. Lu, and R. Song (2019) · 2019
Later among the works it cites.
Dynamic Treatment Regimes: Statistical Methods for Precision Medicine
Tsiatis, A., M. Davidian, S. Holloway, and E. Laber (2019) · 2019
Later among the works it cites.
Model selection for g-estimation of dynamic treatment regimes
Wallace, M. P., E. E. Moodie, and D. A. Stephens (2019) · 2019
Later among the works it cites.