Fetching the paper…
Reading the bibliography…
We consider the problem of approximating the stationary distribution of an ergodic Markov chain given a set of sampled transitions.
Dualdice: Behavior-agnostic estimation of discounted stationary distribution corrections
Ofir Nachum, Yinlam Chow, Bo Dai, and Lihong Li · 1906
Earlier work this paper cites.
Reinforcement Learning: An Introduction
R.S. Sutton and A.G. Barto · 1998
Earlier work this paper cites.
Statistical physics: statics, dynamics and renormalization
Leo P Kadanoff · 2000
Earlier work this paper cites.
Improving predictive inference under convariance shift by weighting the log-likelihood function
H. Shimodaira · 2000
Earlier work this paper cites.
Monte Carlo strategies in scientific computing
Jun S. Liu · 2001
Earlier work this paper cites.
Off-policy temporal-difference learning with function approximation
Doina Precup, Richard S Sutton, and Sanjoy Dasgupta · 2001
Earlier work this paper cites.
An introduction to mcmc for machine learning
Christophe Andrieu, Nando de Freitas, Arnaud Doucet, and Michael I. Jordan · 2003
Earlier work this paper cites.
Slice sampling
Radford M Neal · 2003
Earlier work this paper cites.
The discrete-time geo/geo/1 queue with negative customers and disasters
Ivan Atencia and Pilar Moreno · 2004
Earlier work this paper cites.
Phylogenetic comparative analysis: a modeling approach for adaptive evolution
Marguerite A Butler and Aaron A King · 2004
Earlier work this paper cites.
Monte Carlo Statistical Methods
C. Robert and G. Casella · 2004
Earlier work this paper cites.
Iterative kernel principal component analysis for image modeling
Kwang In Kim, Matthias O. Franz, and Bernhard Schölkopf · 2005
Earlier work this paper cites.
The tradeoffs of large scale learning
Léon Bottou and Olivier Bousquet · 2008
Earlier work this paper cites.
Analysis of the discrete time geo/geo/1 queue with single working vacation
Ji-hong Li and Nai-shuo Tian · 2008
Earlier work this paper cites.
Estimating divergence functionals and the likelihood ratio by penalized convex risk minimization
X.L. Nguyen, M. Wainwright, and M. Jordan · 2008
Earlier work this paper cites.
Direct importance estimation for covariate shift adaptation
Masashi Sugiyama, Taiji Suzuki, Shinichi Nakajima, Hisashi Kashima, Paul von Bünau, and Motoaki Kawanabe · 2008
Earlier work this paper cites.
Stochastic methods , volume 4
Crispin Gardiner · 2009
Earlier work this paper cites.
Covariate shift by kernel mean matching
Arthur Gretton, Alex Smola, Jiayuan Huang, Marcel Schmittfull, Karsten Borgwardt, and Bernhard Schölkopf · 2009
Earlier work this paper cites.
Queues–a course in queueing theory
Moshe Haviv · 2009
Earlier work this paper cites.
Probabilistic graphical models: principles and techniques
Daphne Koller and Nir Friedman · 2009
Earlier work this paper cites.
Markov Chains and Stochastic Stability
Sean Meyn, Richard L. Tweedie, and Peter W. Glynn · 2009
Cited alongside, same era.
Basics of applied stochastic processes
Richard Serfozo · 2009
Cited alongside, same era.
Mcmc using hamiltonian dynamics
Radford M Neal et al · 2011
Cited alongside, same era.
Bayesian learning via stochastic gradient langevin dynamics
Max Welling and Yee-Whye Teh · 2011
Cited alongside, same era.
Modeling stabilizing selection: expanding the ornstein–uhlenbeck model of adaptive evolution
Jeremy M Beaulieu, Dwueng-Chwuan Jhwueng, Carl Boettiger, and Brian C O’Meara · 2012
Cited alongside, same era.
Foundations of Machine Learning
Mehryar Mohri, Afshin Rostamizadeh, and Ameet Talwalkar · 2012
Cited alongside, same era.
Doubly robust off-policy value evaluation for reinforcement learning
Nan Jiang and Lihong Li · 2016
Later among the works it cites.
A kernelized stein discrepancy for goodness-of-fit tests
Qiang Liu, Jason Lee, and Michael Jordan · 2016
Later among the works it cites.
Simulation and the Monte Carlo method , volume 10
Reuven Y Rubinstein and Dirk P Kroese · 2016
Later among the works it cites.
Primer on monotone operator methods
Ernest K Ryu and Stephen Boyd · 2016
Later among the works it cites.
Data-efficient off-policy policy evaluation for reinforcement learning
Philip Thomas and Emma Brunskill · 2016
Later among the works it cites.
Learning from conditional distributions via dual embeddings
Bo Dai, Niao He, Yunpeng Pan, Byron Boots, and Le Song · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dietary hardness, loading behavior, and the evolution of skull form in bats
Sharlene E Santana, Ian R Grosse, and Elizabeth R Dumont · 2012
Cited alongside, same era.
Machine learning in non-stationary environments: Introduction to covariate shift adaptation
Masashi Sugiyama and Motoaki Kawanabe · 2012
Cited alongside, same era.
Density ratio estimation in machine learning
Masashi Sugiyama, Taiji Suzuki, and Takafumi Kanamori · 2012
Cited alongside, same era.
The fast convergence of incremental pca
Akshay Balsubramani, Sanjoy Dasgupta, and Yoav Freund · 2013
Cited alongside, same era.
Stochastic differential equations: an introduction with applications
Bernt Oksendal · 2013
Cited alongside, same era.
Breaking the curse of dimensionality with convex neural networks
Francis R. Bach · 2014
Cited alongside, same era.
Using options and covariance testing for long horizon off-policy policy evaluation
Zhaohan Guo, Philip S Thomas, and Emma Brunskill · 2017
Later among the works it cites.
Consistent on-line off-policy evaluation
Assaf Hallak and Shie Mannor · 2017
Later among the works it cites.
Markov chains and mixing times , volume 107
David A Levin and Yuval Peres · 2017
Later among the works it cites.
Black-box importance sampling
Qiang Liu and Jason Lee · 2017
Later among the works it cites.
Control functionals for monte carlo integration
Chris J Oates, Mark Girolami, and Nicolas Chopin · 2017
Later among the works it cites.
A-nice-mc: Adversarial training for mcmc
Jiaming Song, Shengjia Zhao, and Stefano Ermon · 2017
Later among the works it cites.
Optimal and adaptive off-policy evaluation in contextual bandits
Yu-Xiang Wang, Alekh Agarwal, and Miroslav Dudik · 2017
Later among the works it cites.
Online factorization and partition of complex networks from random walks
Lin F Yang, Vladimir Braverman, Tuo Zhao, and Mengdi Wang · 2017
Later among the works it cites.
Learning deep kernels for exponential family densities
Wenliang Li, Dougal Sutherland, Heiko Strathmann, and Arthur Gretton · 2018
Later among the works it cites.
Breaking the curse of horizon: Infinite-horizon off-policy estimation
Qiang Liu, Lihong Li, Ziyang Tang, and Dengyong Zhou · 2018
Later among the works it cites.
Off-policy deep reinforcement learning by bootstrapping the covariate shift
Carles Gelada and Marc G Bellemare · 2019
Later among the works it cites.
Adversarial learning of a sampler based on an unnormalized distribution
Chunyuan Li, Ke Bai, Jianqiao Li, Guoyin Wang, Changyou Chen, and Lawrence Carin · 2019
Later among the works it cites.
Towards optimal off-policy evaluation for reinforcement learning with marginalized importance sampling
Tengyang Xie, Yifei Ma, and Yu-Xiang Wang · 2019
Later among the works it cites.