Fetching the paper…
Reading the bibliography…
Offline data are both valuable and practical resources for teaching robots complex behaviors.
Learning latent plans from play
Lynch, C., Khansari, M., Xiao, T., Kumar, V., Tompson, J., Levine, S., and Sermanet, P · 1903
Earlier work this paper cites.
The development and generalization of "contingency awareness" in early infancy: Some hypotheses
Watson, J. S · 1966
Earlier work this paper cites.
Gradient-based learning applied to document recognition
LeCun, Y., Bottou, L., Bengio, Y., and Haffner, P · 1998
Earlier work this paper cites.
Elements of information theory
Cover, T. M · 1999
Earlier work this paper cites.
Measuring information transfer
Schreiber, T · 2000
Earlier work this paper cites.
Causation, prediction, and search
Spirtes, P., Glymour, C., and Scheines, R · 2001
Earlier work this paper cites.
Information flows in causal networks
Ay, N. and Polani, D · 2008
Earlier work this paper cites.
Causality
Pearl, J · 2009
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Ross, S., Gordon, G., and Bagnell, D · 2011
Earlier work this paper cites.
Lower and upper bounds for approximation of the kullback-leibler divergence between gaussian mixture models
Durrieu, J.-L., Thiran, J.-P., and Kelly, F · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E · 2012
Earlier work this paper cites.
Batch reinforcement learning
Lange, S., Gabel, T., and Riedmiller, M · 2012
Earlier work this paper cites.
The local information dynamics of distributed computation in complex systems
Lizier, J. T · 2012
Earlier work this paper cites.
Identifiability of causal graphs using functional models
Peters, J., Mooij, J. M., Janzing, D., and Schölkopf, B · 2012
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Bandits with unobserved confounders: A causal approach
Bareinboim, E., Forney, A., and Pearl, J · 2015
Earlier work this paper cites.
Generating sentences from a continuous space
Bowman, S. R., Vilnis, L., Vinyals, O., Dai, A. M., Jozefowicz, R., and Bengio, S · 2015
Earlier work this paper cites.
Learning structured output representation using deep conditional generative models
Sohn, K., Lee, H., and Yan, X · 2015
Earlier work this paper cites.
Hindsight experience replay
Andrychowicz, M., Wolski, F., Ray, A., Schneider, J., Fong, R., Welinder, P., McGrew, B., Tobin, J., Pieter Abbeel, O., and Zaremba, W · 2017
Earlier work this paper cites.
Counterfactual data-fusion for online reinforcement learners
Forney, A., Pearl, J., and Bareinboim, E · 2017
Earlier work this paper cites.
Elements of Causal Inference: Foundations and Learning Algorithms
Peters, J., Janzing, D., and Schölkopf, B · 2017
Earlier work this paper cites.
Asymmetric actor critic for image-based robot learning
Pinto, L., Andrychowicz, M., Welinder, P., Zaremba, W., and Abbeel, P · 2017
Earlier work this paper cites.
Domain randomization for transferring deep neural networks from simulation to the real world
Tobin, J., Fong, R., Ray, A., Schneider, J., Zaremba, W., and Abbeel, P · 2017
Earlier work this paper cites.
Deep sets
Zaheer, M., Kottur, S., Ravanbakhsh, S., Poczos, B., Salakhutdinov, R. R., and Smola, A. J · 2017
Earlier work this paper cites.
Relational inductive biases, deep learning, and graph networks
Battaglia, P. W., Hamrick, J. B., Bapst, V., Sanchez-Gonzalez, A., Zambaldi, V., Malinowski, M., Tacchetti, A., Raposo, D., Santoro, A., Faulkner, R., et al · 2018
Earlier work this paper cites.
Woulda, coulda, shoulda: Counterfactually-guided policy search
Buesing, L., Weber, T., Zwols, Y., Racaniere, S., Guez, A., Lespiau, J.-B., and Heess, N · 2018
Cited alongside, same era.
Contingency-aware exploration in reinforcement learning
Choi, J., Guo, Y., Moczulski, M., Oh, J., Wu, N., Norouzi, M., and Lee, H · 2018
Cited alongside, same era.
Addressing function approximation error in actor-critic methods
Fujimoto, S., Hoof, H., and Meger, D · 2018
Cited alongside, same era.
Confounding-robust policy improvement
Kallus, N. and Zhou, A · 2018
Cited alongside, same era.
Deconfounding reinforcement learning in observational settings
Lu, C., Schölkopf, B., and Hernández-Lobato, J. M · 2018
Cited alongside, same era.
Mega-reward: Achieving human-level play without extrinsic rewards
Song, Y., Wang, J., Lukasiewicz, T., Xu, Z., Zhang, S., Wojcicki, A., and Xu, M · 2020
Later among the works it cites.
Fighting copycat agents in behavioral cloning from observation histories
Wen, C., Lin, J., Darrell, T., Jayaraman, D., and Gao, Y · 2020
Later among the works it cites.
Learning invariant representations for reinforcement learning without reconstruction
Zhang, A., McAllister, R. T., Calandra, R., Gal, Y., and Levine, S · 2020
Later among the works it cites.
A minimalist approach to offline reinforcement learning
Fujimoto, S. and Gu, S. S · 2021
Later among the works it cites.
Offline reinforcement learning with implicit q-learning
Kostrikov, I., Nair, A., and Levine, S · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The book of why: the new science of cause and effect
Pearl, J. and Mackenzie, D · 2018
Cited alongside, same era.
Multi-goal reinforcement learning: Challenging robotics environments and request for research
Plappert, M., Andrychowicz, M., Ray, A., McGrew, B., Baker, B., Powell, G., Schneider, J., Tobin, J., Chociej, M., Welinder, P., et al · 2018
Cited alongside, same era.
Mixmatch: A holistic approach to semi-supervised learning
Berthelot, D., Carlini, N., Goodfellow, I., Papernot, N., Oliver, A., and Raffel, C. A · 2019
Cited alongside, same era.
Monet: Unsupervised scene decomposition and representation, 2019
Burgess, C. P., Matthey, L., Watters, N., Kabra, R., Higgins, I., Botvinick, M., and Lerchner, A · 2019
Cited alongside, same era.
Quantifying generalization in reinforcement learning
Cobbe, K., Klimov, O., Hesse, C., Kim, T., and Schulman, J · 2019
Cited alongside, same era.
Causal confusion in imitation learning
De Haan, P., Jayaraman, D., and Levine, S · 2019
Cited alongside, same era.
Off-policy deep reinforcement learning without exploration
Fujimoto, S., Meger, D., and Precup, D · 2019
Cited alongside, same era.
Lyle, C., Zhang, A., Jiang, M., Pineau, J., and Gal, Y · 2021
Later among the works it cites.
Discovering and achieving goals via world models
Mendonca, R., Rybkin, O., Daniilidis, K., Hafner, D., and Pathak, D · 2021
Later among the works it cites.
Causal influence detection for improving efficiency in reinforcement learning
Seitzer, M., Schölkopf, B., and Martius, G · 2021
Later among the works it cites.
Risk-averse offline reinforcement learning
Urpí, N. A., Curi, S., and Krause, A · 2021
Later among the works it cites.
Task-independent causal state abstraction
Wang, Z., Xiao, X., Zhu, Y., and Stone, P · 2021
Later among the works it cites.
Human-to-robot imitation in the wild
Bahl, S., Gupta, A., and Pathak, D · 2022
Later among the works it cites.
Rt-1: Robotics transformer for real-world control at scale
Brohan, A., Brown, N., Carbajal, J., Chebotar, Y., Dabis, J., Finn, C., Gopalakrishnan, K., Hausman, K., Herzog, A., Hsu, J., et al · 2022
Later among the works it cites.
Generalizing goal-conditioned reinforcement learning with variational causal reasoning
Ding, W., Lin, H., Li, B., and Zhao, D · 2022
Later among the works it cites.
Mocoda: Model-based counterfactual data augmentation
Pitis, S., Creager, E., Mandlekar, A., and Garg, A · 2022
Later among the works it cites.
Latent plans for task agnostic offline reinforcement learning
Rosete-Beas, E., Mees, O., Kalweit, G., Boedecker, J., and Burgard, W · 2022
Later among the works it cites.
Bridging the gap to real-world object-centric learning
Seitzer, M., Horn, M., Zadaianchuk, A., Zietlow, D., Xiao, T., Simon-Gabriel, C.-J., He, T., Zhang, Z., Schölkopf, B., Brox, T., et al · 2022
Later among the works it cites.
Causal dynamics learning for task-independent state abstraction
Wang, Z., Xiao, X., Xu, Z., Zhu, Y., and Stone, P · 2022
Later among the works it cites.
Rt-2: Vision-language-action models transfer web knowledge to robotic control
Brohan, A., Brown, N., Carbajal, J., Chebotar, Y., Chen, X., Choromanski, K., Ding, T., Driess, D., Dubey, A., Finn, C., et al · 2023
Later among the works it cites.
Seeing is not believing: Robust reinforcement learning against spurious correlation
Ding, W., Shi, L., Chi, Y., and Zhao, D · 2023
Later among the works it cites.
Can active sampling reduce causal confusion in offline reinforcement learning?
Gupta, G., Rudner, T. G., McAllister, R. T., Gaidon, A., and Gal, Y · 2023
Later among the works it cites.
Learning agile skills via adversarial imitation of rough partial demonstrations
Li, C., Vlastelica, M., Blaes, S., Frey, J., Grimminger, F., and Martius, G · 2023
Later among the works it cites.
Efficient learning of high level plans from play
Urpí, N. A., Bagatella, M., Hilliges, O., Martius, G., and Coros, S · 2023
Later among the works it cites.
Diverse offline imitation learning, 2023
Vlastelica, M., Cheng, J., Martius, G., and Kolev, P · 2023
Later among the works it cites.
Object-centric learning for real-world videos by predicting temporal feature similarities, 2023
Zadaianchuk, A., Seitzer, M., and Martius, G · 2023
Later among the works it cites.
Regularized behavior cloning for blocking the leakage of past action information
Seo, S., Hwang, H., Yang, H., and Kim, K.-E · 2024
Closest in time.