Fetching the paper…
Reading the bibliography…
Learning to control an agent from data collected offline in a rich pixel-based visual observation space is vital for real-world applications of reinforcement learning (RL).
Way off-policy batch deep reinforcement learning of implicit human preferences in dialog
Jaques, N., Ghandeharioun, A., Shen, J. H., Ferguson, C., Lapedriza, À., Jones, N., Gu, S., and Picard, R. W · 1907
Earlier work this paper cites.
Error bounds for approximate policy iteration
Munos, R · 2003
Earlier work this paper cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Levine, S., Kumar, A., Tucker, G., and Fu, J · 2005
Earlier work this paper cites.
Learning near-optimal policies with bellman-residual minimization based fitted policy iteration and a single sample path
Antos, A., Szepesvári, C., and Munos, R · 2008
Earlier work this paper cites.
Finite-time bounds for fitted value iteration
Munos, R. and Szepesvári, C · 2008
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A. and Hinton, G · 2009
Earlier work this paper cites.
An analysis of single-layer networks in unsupervised feature learning
Coates, A., Ng, A., and Lee, H · 2011
Earlier work this paper cites.
Batch reinforcement learning
Lange, S., Gabel, T., and Riedmiller, M. A · 2012
Earlier work this paper cites.
Curiosity-driven exploration by self-supervised prediction
Pathak, D., Agrawal, P., Efros, A. A., and Darrell, T · 2017
Earlier work this paper cites.
Neural discrete representation learning
Van Den Oord, A., Vinyals, O., et al · 2017
Earlier work this paper cites.
Discovering and removing exogenous state variables and rewards for reinforcement learning
Dietterich, T., Trimponias, G., and Chen, Z · 2018
Earlier work this paper cites.
Visual foresight: Model-based deep reinforcement learning for vision-based robotic control
Ebert, F., Finn, C., Dasari, S., Xie, A., Lee, A. X., and Levine, S · 2018
Earlier work this paper cites.
Learning complex dexterous manipulation with deep reinforcement learning and demonstrations
Rajeswaran, A., Kumar, V., Gupta, A., Vezzani, G., Schulman, J., Todorov, E., and Levine, S · 2018
Earlier work this paper cites.
Information-theoretic considerations in batch reinforcement learning
Chen, J. and Jiang, N · 2019
Earlier work this paper cites.
Causal confusion in imitation learning
De Haan, P., Jayaraman, D., and Levine, S · 2019
Earlier work this paper cites.
Off-policy deep reinforcement learning without exploration
Fujimoto, S., Meger, D., and Precup, D · 2019
Earlier work this paper cites.
Dream to control: Learning behaviors by latent imagination
Hafner, D., Lillicrap, T., Ba, J., and Norouzi, M · 2019
Earlier work this paper cites.
Stabilizing off-policy q-learning via bootstrapping error reduction
Kumar, A., Fu, J., Soh, M., Tucker, G., and Levine, S · 2019
Earlier work this paper cites.
Safe policy improvement with baseline bootstrapping
Laroche, R., Trichelair, P., and des Combes, R. T · 2019
Earlier work this paper cites.
Safe policy improvement with soft baseline bootstrapping
Nadjahi, K., Laroche, R., and Tachet des Combes, R · 2019
Cited alongside, same era.
An optimistic perspective on offline reinforcement learning
Agarwal, R., Schuurmans, D., and Norouzi, M · 2020
Cited alongside, same era.
D4rl: Datasets for deep data-driven reinforcement learning
Fu, J., Kumar, A., Nachum, O., Tucker, G., and Levine, S · 2020
Cited alongside, same era.
Representations for stable off-policy reinforcement learning
Ghosh, D. and Bellemare, M. G · 2020
Cited alongside, same era.
Morel: Model-based offline reinforcement learning
Kidambi, R., Rajeswaran, A., Netrapalli, P., and Joachims, T · 2020
Cited alongside, same era.
Image augmentation is all you need: Regularizing deep reinforcement learning from pixels
An empirical investigation of representation learning for imitation
Chen, C., Chen, X., Toyer, S., Wild, C., Emmons, S., Fischer, I., Lee, K., Alex, N., Wang, S. H., Luo, P., Russell, S., Abbeel, P., and Shah, R · 2021
Later among the works it cites.
Decision transformer: Reinforcement learning via sequence modeling
Chen, L., Lu, K., Rajeswaran, A., Lee, K., Grover, A., Laskin, M., Abbeel, P., Srinivas, A., and Mordatch, I · 2021
Later among the works it cites.
Provably filtering exogenous distractors using multistep inverse dynamics
Efroni, Y., Misra, D., Krishnamurthy, A., Agarwal, A., and Langford, J · 2021
Later among the works it cites.
Offline reinforcement learning: Fundamental barriers for value function approximation
Foster, D. J., Krishnamurthy, A., Simchi-Levi, D., and Xu, Y · 2021
Later among the works it cites.
A minimalist approach to offline reinforcement learning
Fujimoto, S. and Gu, S. S · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kostrikov, I., Yarats, D., and Fergus, R · 2020
Cited alongside, same era.
Conservative q-learning for offline reinforcement learning
Kumar, A., Zhou, A., Tucker, G., and Levine, S · 2020
Cited alongside, same era.
Curl: Contrastive unsupervised representations for reinforcement learning
Laskin, M., Srinivas, A., and Abbeel, P · 2020
Cited alongside, same era.
Deep reinforcement and infomax learning
Mazoure, B., Tachet des Combes, R., Doan, T. L., Bachman, P., and Hjelm, R. D · 2020
Cited alongside, same era.
Kinematic state abstraction and provably efficient rich-observation reinforcement learning
Misra, D., Henaff, M., Krishnamurthy, A., and Langford, J · 2020
Cited alongside, same era.
Planning from pixels using inverse dynamics models
Paster, K., McIlraith, S. A., and Ba, J · 2020
Cited alongside, same era.
Data-efficient reinforcement learning with self-predictive representations
Schwarzer, M., Anand, A., Goel, R., Hjelm, R. D., Courville, A., and Bachman, P · 2020
Cited alongside, same era.
Provable representation learning for imitation with contrastive fourier features
Nachum, O. and Yang, M · 2021
Later among the works it cites.
Pretraining representations for data-efficient reinforcement learning
Schwarzer, M., Rajkumar, N., Noukhovitch, M., Anand, A., Charlin, L., Hjelm, R. D., Bachman, P., and Courville, A. C · 2021
Later among the works it cites.
Learning representations for pixel-based control: What matters and why?
Tomar, M., Mishra, U. A., Zhang, A., and Taylor, M. E · 2021
Later among the works it cites.
Representation learning for online and offline RL in low-rank mdps
Uehara, M., Zhang, X., and Sun, W · 2021
Later among the works it cites.
Representation matters: Offline pretraining for sequential decision making
Yang, M. and Nachum, O · 2021
Later among the works it cites.
Learning invariant representations for reinforcement learning without reconstruction
Zhang, A., McAllister, R. T., Calandra, R., Gal, Y., and Levine, S · 2021
Later among the works it cites.
BRAC+: improved behavior regularized actor critic for offline reinforcement learning
Zhang, C., Kuppannagari, S. R., and Prasanna, V. K · 2021
Later among the works it cites.
Sample-efficient reinforcement learning in the presence of exogenous information
Efroni, Y., Foster, D. J., Misra, D., Krishnamurthy, A., and Langford, J · 2022
Closest in time.
Byol-explore: Exploration by bootstrapped prediction
Guo, Z. D., Thakoor, S., Pîslar, M., Pires, B. A., Altché, F., Tallec, C., Saade, A., Calandriello, D., Grill, J.-B., Tang, Y., Valko, M., Munos, R., Azar, M. G., and Piot, B · 2022
Closest in time.
Uniqueness and complexity of inverse mdp models
Hutter, M. and Hansen, S · 2022
Closest in time.
Guaranteed discovery of controllable latent states with multi-step inverse models
Lamb, A., Islam, R., Efroni, Y., Didolkar, A., Misra, D., Foster, D., Molu, L., Chari, R., Krishnamurthy, A., and Langford, J · 2022
Closest in time.
Simsr: Simple distance-based state representations for deep reinforcement learning
Zang, H., Li, X., and Wang, M · 2022
Closest in time.
Behavior prior representation learning for offline reinforcement learning
Zang, H., Li, X., Yu, J., Liu, C., Islam, R., des Combes, R. T., and Laroche, R · 2023
Closest in time.