Fetching the paper…
Reading the bibliography…
Recently, latent action learning, pioneered by Latent Action Policies (LAPO), have shown remarkable pre-training efficiency on observation-only data, offering potential for leveraging vast amounts of video available on the web for embodied AI.
Understanding intermediate layers using linear classifier probes
Alain, G · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Proximal policy optimization algorithms
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O · 2017
Earlier work this paper cites.
Neural discrete representation learning
Van Den Oord, A., Vinyals, O., et al · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A · 2017
Earlier work this paper cites.
Playing hard exploration games by watching youtube
Aytar, Y., Pfaff, T., Budden, D., Paine, T., Wang, Z., and De Freitas, N · 2018
Earlier work this paper cites.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Haarnoja, T., Zhou, A., Abbeel, P., and Levine, S · 2018
Earlier work this paper cites.
Behavioral cloning from observation
Torabi, F., Warnell, G., and Stone, P · 2018
Earlier work this paper cites.
Bhatt, A., Palenicek, D., Belousov, B., Argus, M., Amiranashvili, A., Brox, T., and Peters, J · 2019
Earlier work this paper cites.
Imitating latent policies from observation
Edwards, A., Sahni, H., Schroecker, Y., and Isbell, C · 2019
Earlier work this paper cites.
Recent advances in imitation learning from observation
Torabi, F., Warnell, G., and Stone, P · 2019
Earlier work this paper cites.
Leveraging procedural generation to benchmark reinforcement learning
Cobbe, K., Hesse, C., Hilton, J., and Schulman, J · 2020
Earlier work this paper cites.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Levine, S., Kumar, A., Tucker, G., and Fu, J · 2020
Earlier work this paper cites.
Reinforcement learning with videos: Combining offline observations with interaction
Schmeckpeper, K., Rybkin, O., Daniilidis, K., Levine, S., and Finn, C · 2020
Earlier work this paper cites.
Data-efficient reinforcement learning with self-predictive representations
Schwarzer, M., Anand, A., Goel, R., Hjelm, R. D., Courville, A., and Bachman, P · 2020
Earlier work this paper cites.
The pitfalls of simplicity bias in neural networks
Shah, H., Tamuly, K., Raghunathan, A., Jain, P., and Netrapalli, P · 2020
Earlier work this paper cites.
Learning invariant representations for reinforcement learning without reconstruction
Zhang, A., McAllister, R., Calandra, R., Gal, Y., and Levine, S · 2020
Earlier work this paper cites.
Exploring simple siamese representation learning
Chen, X. and He, K · 2021
Earlier work this paper cites.
Learning task informed abstractions
Fu, X., Yang, G., Agrawal, P., and Jaakkola, T · 2021
Earlier work this paper cites.
Generalization in reinforcement learning by soft data augmentation
Hansen, N. and Wang, X · 2021
Earlier work this paper cites.
Stabilizing deep q-learning with convnets and vision transformers under data augmentation
Hansen, N., Su, H., and Wang, X · 2021
Cited alongside, same era.
Dreaming: Model-based reinforcement learning by latent imagination without reconstruction
Okada, M. and Taniguchi, T · 2021
Cited alongside, same era.
Stable-baselines3: Reliable reinforcement learning implementations
Raffin, A., Hill, A., Gleave, A., Kanervisto, A., Ernestus, M., and Dormann, N · 2021
Cited alongside, same era.
The distracting control suite–a challenging benchmark for reinforcement learning from pixels
Stone, A., Ramirez, O., Konolige, K., and Jonschkowski, R · 2021
Cited alongside, same era.
Learning representations for pixel-based control: What matters and why?
Tomar, M., Mishra, U. A., Zhang, A., and Taylor, M. E · 2021
Cited alongside, same era.
Preventing mode collapse when imitating latent policies from observations, 2023
Struckmeier, O. and Kyrki, V · 2023
Later among the works it cites.
Semail: eliminating distractors in visual imitation via separated models
Wan, S., Wang, Y., Shao, M., Chen, R., and Zhan, D.-C · 2023
Later among the works it cites.
Simplified temporal consistency reinforcement learning
Zhao, Y., Zhao, W., Boney, R., Kannala, J., and Pajarinen, J · 2023
Later among the works it cites.
Semi-supervised offline reinforcement learning with action-free trajectories
Zheng, Q., Henaff, M., Amos, B., and Grover, A · 2023
Later among the works it cites.
Learning robust representation for reinforcement learning with distractions by reward sequence prediction
Zhou, Q., Wang, J., Liu, Q., Kuang, Y., Zhou, W., and Li, H · 2023
Later among the works it cites.
Repo: Resilient model-based reinforcement learning by regularizing posterior predictability
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Video pretraining (vpt): Learning to act by watching unlabeled online videos
Baker, B., Akkaya, I., Zhokov, P., Huizinga, J., Tang, J., Ecoffet, A., Houghton, B., Sampedro, R., and Clune, J · 2022
Cited alongside, same era.
Look where you look! saliency-guided q-networks for generalization in visual reinforcement learning
Bertoin, D., Zouitine, A., Zouitine, M., and Rachelson, E · 2022
Cited alongside, same era.
Dreamerpro: Reconstruction-free model-based reinforcement learning with prototypical representations
Deng, F., Jang, I., and Ahn, S · 2022
Cited alongside, same era.
Ego4d: Around the world in 3,000 hours of egocentric video
Grauman, K., Westbury, A., Byrne, E., Chavis, Z., Furnari, A., Girdhar, R., Hamburger, J., Jiang, H., Liu, M., Liu, X., et al · 2022
Cited alongside, same era.
Dungeons and data: A large-scale nethack dataset
Hambro, E., Raileanu, R., Rothermel, D., Mella, V., Rocktäschel, T., Küttler, H., and Murray, N · 2022
Cited alongside, same era.
Temporal difference learning for model predictive control
Hansen, N., Wang, X., and Su, H · 2022
Cited alongside, same era.
Agent-controller representations: Principled offline rl with rich exogenous information
Islam, R., Tomar, M., Lamb, A., Efroni, Y., Zang, H., Didolkar, A., Misra, D., Li, X., Van Seijen, H., Combes, R. T. d., et al · 2022
Cited alongside, same era.
Zhu, C., Simchowitz, M., Gadipudi, S., and Gupta, A · 2023
Later among the works it cites.
A recipe for unbounded data augmentation in visual reinforcement learning
Almuzairee, A., Hansen, N., and Christensen, H. I · 2024
Later among the works it cites.
Zero-shot generalization of vision-based rl without data augmentation
Batra, S. and Sukhatme, G. S · 2024
Later among the works it cites.
Genie: Generative interactive environments
Bruce, J., Dennis, M. D., Edwards, A., Parker-Holder, J., Shi, Y., Hughes, E., Lai, M., Mavalankar, A., Steigerwald, R., Apps, C., et al · 2024
Later among the works it cites.
Mudreamer: Learning predictive world models without reconstruction
Burchi, M. and Timofte, R · 2024
Later among the works it cites.
Dynamo: In-domain dynamics pretraining for visuo-motor control
Cui, Z. J., Pan, H., Iyer, A., Haldar, S., and Pinto, L · 2024
Later among the works it cites.
Droid: A large-scale in-the-wild robot manipulation dataset
Khazatsky, A., Pertsch, K., Nair, S., Balakrishna, A., Dasari, S., Karamcheti, S., Nasiriany, S., Srirama, M. K., Chen, L. Y., Ellis, K., et al · 2024
Later among the works it cites.
Multistep inverse is not all you need
Levine, A., Stone, P., and Zhang, A · 2024
Later among the works it cites.
Towards generalist robot learning from internet video: A survey
McCarthy, R., Tan, D. C., Schmidt, D., Acero, F., Herr, N., Du, Y., Thuruthel, T. G., and Li, Z · 2024
Later among the works it cites.
Towards principled representation learning from videos for reinforcement learning
Misra, D., Saran, A., Xie, T., Lamb, A., and Langford, J · 2024
Later among the works it cites.
Bridging state and history representations: Understanding self-predictive rl
Ni, T., Eysenbach, B., Seyedsalehi, E., Ma, M., Gehring, C., Mahajan, A., and Bacon, P.-L · 2024
Later among the works it cites.
Dmc-vb: A benchmark for representation learning for control with visual distractors
Ortiz, J., Dedieu, A., Lehrach, W., Guntupalli, S., Wendelken, C., Humayun, A., Zhou, G., Swaminathan, S., Lázaro-Gredilla, M., and Murphy, K · 2024
Later among the works it cites.
iqrl–implicitly quantized representations for sample-efficient reinforcement learning
Scannell, A., Kujanpää, K., Zhao, Y., Nakhaei, M., Solin, A., and Pajarinen, J · 2024
Later among the works it cites.
Ad3: Implicit action is the key for world models to distinguish the diverse visual distractors
Wang, Y., Wan, S., Gan, L., Feng, S., and Zhan, D.-C · 2024
Later among the works it cites.
Latent action pretraining from videos
Ye, S., Jang, J., Jeon, B., Joo, S., Yang, J., Peng, B., Mandlekar, A., Tan, R., Chao, Y.-W., Lin, B. Y., et al · 2024
Later among the works it cites.