Fetching the paper…
Reading the bibliography…
We propose State Matching Offline DIstribution Correction Estimation (SMODICE), a novel and versatile regression-based offline imitation learning (IL) algorithm derived via state-occupancy matching.
Dualdice: Behavior-agnostic estimation of discounted stationary distribution corrections
Nachum, O., Chow, Y., Dai, B., and Li, L · 1906
Earlier work this paper cites.
Residual algorithms: Reinforcement learning with function approximation
Baird, L · 1995
Earlier work this paper cites.
Convex optimization
Boyd, S., Boyd, S. P., and Vandenberghe, L · 2004
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning, 2011
Ross, S., Gordon, G. J., and Bagnell, J. A · 2011
Earlier work this paper cites.
Batch reinforcement learning
Lange, S., Gabel, T., and Riedmiller, M · 2012
Earlier work this paper cites.
Generative adversarial networks, 2014
Goodfellow, I. J., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
Puterman, M. L · 2014
Earlier work this paper cites.
Convex Analysis
Rockafellar, R. T · 2015
Earlier work this paper cites.
Learning from conditional distributions via dual embeddings, 2016
Dai, B., He, N., Pan, Y., Boots, B., and Song, L · 2016
Earlier work this paper cites.
Generative adversarial imitation learning, 2016
Ho, J. and Ermon, S · 2016
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor, 2018
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S · 2018
Earlier work this paper cites.
Kostrikov, I., Agrawal, K. K., Dwibedi, D., Levine, S., and Tompson, J · 2018
Earlier work this paper cites.
Behavioral cloning from observation, 2018
Torabi, F., Warnell, G., and Stone, P · 2018
Earlier work this paper cites.
A divergence minimization perspective on imitation learning methods, 2019
Ghasemipour, S. K. S., Zemel, R., and Gu, S · 2019
Earlier work this paper cites.
Relay policy learning: Solving long-horizon tasks via imitation and reinforcement learning
Gupta, A., Kumar, V., Lynch, C., Levine, S., and Hausman, K · 2019
Cited alongside, same era.
Stabilizing off-policy q-learning via bootstrapping error reduction
Kumar, A., Fu, J., Tucker, G., and Levine, S · 2019
Cited alongside, same era.
State alignment-based imitation learning
Liu, F., Ling, Z., Mu, T., and Su, H · 2019
Cited alongside, same era.
Generative adversarial imitation from observation, 2019
Torabi, F., Warnell, G., and Stone, P · 2019
Cited alongside, same era.
Imitation learning from observations by minimizing inverse dynamics disagreement
Yang, C., Ma, X., Huang, W., Sun, F., Liu, H., Huang, J., and Gan, C · 2019
Reinforcement learning via fenchel-rockafellar duality, 2020
Nachum, O. and Dai, B · 2020
Later among the works it cites.
State-only imitation learning for dexterous manipulation, 2020
Radosavovic, I., Wang, X., Pinto, L., and Malik, J · 2020
Later among the works it cites.
Mopo: Model-based offline policy optimization
Yu, T., Thomas, G., Yu, L., Ermon, S., Zou, J., Levine, S., Finn, C., and Ma, T · 2020
Later among the works it cites.
Gendice: Generalized offline estimation of stationary values
Zhang*, R., Dai*, B., Li, L., and Schuurmans, D · 2020
Later among the works it cites.
Off-policy imitation learning from observations
Zhu, Z., Lin, K., Dai, B., and Zhou, J · 2020
Later among the works it cites.
Offline learning from demonstrations and unlabeled experience
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Coindice: Off-policy confidence interval estimation
Dai, B., Nachum, O., Chow, Y., Li, L., Szepesvári, C., and Schuurmans, D · 2020
Cited alongside, same era.
State-only imitation with transition dynamics mismatch
Gangwani, T. and Peng, J · 2020
Cited alongside, same era.
Array programming with NumPy
Harris, C. R., Millman, K. J., van der Walt, S. J., Gommers, R., Virtanen, P., Cournapeau, D., Wieser, E., Taylor, J., Berg, S., Smith, N. J., Kern, R., Picus, M., Hoyer, S., van Kerkwijk, M. H., Brett, M., Haldane, A., del Río, J. F., Wiebe, M., Peterson, P., Gérard-Marchant, P., Sheppard, K., Reddy, T., Weckesser, W., Abbasi, H., Gohlke, C., and Oliphant, T. E · 2020
Cited alongside, same era.
Imitation learning as f f -divergence minimization, 2020
Ke, L., Choudhury, S., Barnes, M., Sun, W., Lee, G., and Srinivasa, S · 2020
Cited alongside, same era.
Morel: Model-based offline reinforcement learning
Kidambi, R., Rajeswaran, A., Netrapalli, P., and Joachims, T · 2020
Cited alongside, same era.
Domain adaptive imitation learning
Kim, K., Gu, Y., Song, J., Zhao, S., and Ermon, S · 2020
Cited alongside, same era.
Imitation learning via off-policy distribution matching
Kostrikov, I., Nachum, O., and Tompson, J · 2020
Cited alongside, same era.
Zolna, K., Novikov, A., Konyushkova, K., Gulcehre, C., Wang, Z., Aytar, Y., Denil, M., de Freitas, N., and Reed, S · 2020
Later among the works it cites.
Mitigating covariate shift in imitation learning via offline data without great coverage, 2021
Chang, J. D., Uehara, M., Sreenivas, D., Kidambi, R., and Sun, W · 2021
Later among the works it cites.
Replacing rewards with examples: Example-based policy search via recursive classification
Eysenbach, B., Levine, S., and Salakhutdinov, R · 2021
Later among the works it cites.
D4rl: Datasets for deep data-driven reinforcement learning, 2021
Fu, J., Kumar, A., Nachum, O., Tucker, G., and Levine, S · 2021
Later among the works it cites.
A minimalist approach to offline reinforcement learning
Fujimoto, S. and Gu, S. S · 2021
Later among the works it cites.
Optidice: Offline policy optimization via stationary distribution correction estimation
Lee, J., Jeon, W., Lee, B.-J., Pineau, J., and Kim, K.-E · 2021
Later among the works it cites.
Cross-domain imitation from observations, 2021
Raychaudhuri, D. S., Paul, S., van Baar, J., and Roy-Chowdhury, A. K · 2021
Later among the works it cites.
On pathologies in KL-regularized reinforcement learning from expert demonstrations
Rudner, T. G. J., Lu, C., Osborne, M., Gal, Y., and Teh, Y. W · 2021
Later among the works it cites.
DemoDICE: Offline imitation learning with supplementary imperfect demonstrations
Kim, G.-H., Seo, S., Lee, J., Jeon, W., Hwang, H., Yang, H., and Kim, K.-E · 2022
Closest in time.