Fetching the paper…
Reading the bibliography…
Multimodal pretraining is an effective strategy for the trinity of goals of representation learning in autonomous robots: 1) extracting both local and global task progressions; 2) enforcing temporal consistency of visual representation; 3) capturing trajectory-level language grounding.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Bradley, R. A. and Terry, M. E · 1952
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Ng, A. Y., Harada, D., and Russell, S · 1999
Earlier work this paper cites.
Preference-based policy learning
Akrour, R., Schoenauer, M., and Sebag, M · 2011
Earlier work this paper cites.
Sampling strategies for real-time action recognition
Shi, F., Petriu, E., and Laganiere, R · 2013
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
Puterman, M. L · 2014
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
The” something something” video database for learning and evaluating visual common sense
Goyal, R., Ebrahimi Kahou, S., Michalski, V., Materzynska, J., Westphal, S., Kim, H., Haenel, V., Fruend, I., Yianilos, P., Mueller-Freitag, M., et al · 2017
Earlier work this paper cites.
Model predictive path integral control: From theory to parallel computation
Williams, G., Aldrich, A., and Theodorou, E. A · 2017
Earlier work this paper cites.
Scaling egocentric vision: The epic-kitchens dataset
Damen, D., Doughty, H., Farinella, G. M., Fidler, S., Furnari, A., Kazakos, E., Moltisanti, D., Munro, J., Perrett, T., Price, W., et al · 2018
Earlier work this paper cites.
Spectral normalization for generative adversarial networks
Miyato, T., Kataoka, T., Koyama, M., and Yoshida, Y · 2018
Earlier work this paper cites.
Representation learning with contrastive predictive coding
Oord, A. v. d., Li, Y., and Vinyals, O · 2018
Earlier work this paper cites.
Time-contrastive networks: Self-supervised learning from video
Sermanet, P., Lynch, C., Chebotar, Y., Hsu, J., Jang, E., Schaal, S., Levine, S., and Brain, G · 2018
Earlier work this paper cites.
Lipschitz regularity of deep neural networks: analysis and efficient estimation
Virmaux, A. and Scaman, K · 2018
Earlier work this paper cites.
Temporal segment networks for action recognition in videos
Wang, L., Xiong, Y., Wang, Z., Qiao, Y., Lin, D., Tang, X., and Van Gool, L · 2018
Earlier work this paper cites.
Relay policy learning: Solving long-horizon tasks via imitation and reinforcement learning
Gupta, A., Kumar, V., Lynch, C., Levine, S., and Hausman, K · 2019
Earlier work this paper cites.
Scsampler: Sampling salient clips from video for efficient action recognition
Korbar, B., Tran, D., and Torresani, L · 2019
Earlier work this paper cites.
Towards explaining the regularization effect of initial large learning rate in training neural networks
Li, Y., Wei, C., and Ma, T · 2019
Earlier work this paper cites.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Sanh, V., Debut, L., Chaumond, J., and Wolf, T · 2019
Earlier work this paper cites.
A review on deep learning techniques for video prediction
Oprea, S., Martinez-Gonzalez, P., Garcia-Garcia, A., Castro-Vargas, J. A., Orts-Escolano, S., Garcia-Rodriguez, J., and Argyros, A · 2020
Earlier work this paper cites.
Pebble: Feedback-efficient interactive reinforcement learning via relabeling experience and unsupervised pre-training
Lee, K., Smith, L. M., and Abbeel, P · 2021
Earlier work this paper cites.
Rrl: Resnet as representation for reinforcement learning
Shah, R. M. and Kumar, V · 2021
Earlier work this paper cites.
Image augmentation is all you need: Regularizing deep reinforcement learning from pixels
Yarats, D., Fergus, R., and Kostrikov, I · 2021
Cited alongside, same era.
Learning invariant representations for reinforcement learning without reconstruction
Zhang, A., McAllister, R. T., Calandra, R., Gal, Y., and Levine, S · 2021
Cited alongside, same era.
Mgsampler: An explainable sampling strategy for video action recognition
Zhi, Y., Tong, Z., Wang, L., and Wu, G · 2021
Cited alongside, same era.
Rt-1: Robotics transformer for real-world control at scale
Brohan, A., Brown, N., Carbajal, J., Chebotar, Y., Dabis, J., Finn, C., Gopalakrishnan, K., Hausman, K., Herzog, A., Hsu, J., et al · 2022
Cited alongside, same era.
Can foundation models perform zero-shot task specification for robot manipulation?
Cui, Y., Niekum, S., Gupta, A., Kumar, V., and Rajeswaran, A · 2022
Cited alongside, same era.
Long-horizon video prediction using a dynamic latent hierarchy
Zakharov, A., Guo, Q., and Fountas, Z · 2022
Later among the works it cites.
Compositional foundation models for hierarchical planning
Ajay, A., Han, S., Du, Y., Li, S., Gupta, A., Jaakkola, T. S., Tenenbaum, J. B., Kaelbling, L. P., Srivastava, A., and Agrawal, P · 2023
Later among the works it cites.
Robotic offline rl from internet videos via value-function pre-training
Bhateja, C. A., Guo, D., Ghosh, D., Singh, A., Tomar, M., Vuong, Q., Chebotar, Y., Levine, S., and Kumar, A · 2023
Later among the works it cites.
Zero-shot robotic manipulation with pre-trained image-editing diffusion models
Black, K., Nakamoto, M., Atreya, P., Walke, H., Finn, C., Kumar, A., and Levine, S · 2023
Later among the works it cites.
Open X-Embodiment: Robotic learning datasets and RT-X models
Collaboration, O. X.-E., Padalkar, A., Pooley, A., Jain, A., Bewley, A., Herzog, A., Irpan, A., Khazatsky, A., Rai, A., Singh, A., Brohan, A., Raffin, A., Wahid, A., Burgess-Limerick, B., Kim, B., Schölkopf, B., Ichter, B., Lu, C., Xu, C., Finn, C., Xu, C., Chi, C., Huang, C., Chan, C., Pan, C., Fu, C., Devin, C., Driess, D., Pathak, D., Shah, D., Büchler, D., Kalashnikov, D., Sadigh, D., Johns, E., Ceola, F., Xia, F., Stulp, F., Zhou, G., Sukhatme, G. S., Salhotra, G., Yan, G., Schiavi, G., Su, H., Fang, H.-S., Shi, H., Amor, H. B., Christensen, H. I., Furuta, H., Walke, H., Fang, H., Mordatch, I., Radosavovic, I., Leal, I., Liang, J., Kim, J., Schneider, J., Hsu, J., Bohg, J., Bingham, J., Wu, J., Wu, J., Luo, J., Gu, J., Tan, J., Oh, J., Malik, J., Tompson, J., Yang, J., Lim, J. J., Silvério, J., Han, J., Rao, K., Pertsch, K., Hausman, K., Go, K., Gopalakrishnan, K., Goldberg, K., Byrne, K., Oslund, K., Kawaharazuka, K., Zhang, K., Majd, K., Rana, K., Srinivasan, K., Chen, L. Y., Pinto, L., Tan, L., Ott, L., Lee, L., Tomizuka, M., Du, M., Ahn, M., Zhang, M., Ding, M., Srirama, M. K., Sharma, M., Kim, M. J., Kanazawa, N., Hansen, N., Heess, N., Joshi, N. J., Suenderhauf, N., Palo, N. D., Shafiullah, N. M. M., Mees, O., Kroemer, O., Sanketi, P. R., Wohlhart, P., Xu, P., Sermanet, P., Sundaresan, P., Vuong, Q., Rafailov, R., Tian, R., Doshi, R., Martín-Martín, R., Mendonca, R., Shah, R., Hoque, R., Julian, R., Bustamante, S., Kirmani, S., Levine, S., Moore, S., Bahl, S., Dass, S., Song, S., Xu, S., Haldar, S., Adebola, S., Guist, S., Nasiriany, S., Schaal, S., Welker, S., Tian, S., Dasari, S., Belkhale, S., Osa, T., Harada, T., Matsushima, T., Xiao, T., Yu, T., Ding, T., Davchev, T., Zhao, T. Z., Armstrong, T., Darrell, T., Jain, V., Vanhoucke, V., Zhan, W., Zhou, W., Burgard, W., Chen, X., Wang, X., Zhu, X., Li, X., Lu, Y., Chebotar, Y., Zhou, Y., Zhu, Y., Xu, Y., Wang, Y., Bisk, Y., Cho, Y., Lee, Y., Cui, Y., hua Wu, Y., Tang, Y., Zhu, Y., Li, Y., Iwasawa, Y., Matsuo, Y., Xu, Z., and Cui, Z. J · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ego4d: Around the world in 3,000 hours of egocentric video
Grauman, K., Westbury, A., Byrne, E., Chavis, Z., Furnari, A., Girdhar, R., Hamburger, J., Jiang, H., Liu, M., Liu, X., et al · 2022
Cited alongside, same era.
Masked autoencoders are scalable vision learners
He, K., Chen, X., Xie, S., Li, Y., Dollár, P., and Girshick, R · 2022
Cited alongside, same era.
Bc-z: Zero-shot task generalization with robotic imitation learning
Jang, E., Irpan, A., Khansari, M., Kappler, D., Ebert, F., Lynch, C., Levine, S., and Finn, C · 2022
Cited alongside, same era.
Simple but effective: Clip embeddings for embodied ai
Khandelwal, A., Weihs, L., Mottaghi, R., and Kembhavi, A · 2022
Cited alongside, same era.
Pre-training for robots: Offline rl enables learning new tasks from a handful of trials
Kumar, A., Singh, A., Ebert, F., Nakamoto, M., Yang, Y., Finn, C., and Levine, S · 2022
Cited alongside, same era.
Phasic self-imitative reduction for sparse-reward goal-conditioned reinforcement learning
Li, Y., Gao, T., Yang, J., Xu, H., and Wu, Y · 2022
Cited alongside, same era.
Learning smooth neural functions via lipschitz regularization
Liu, H.-T. D., Williams, F., Jacobson, A., Fidler, S., and Litany, O · 2022
Cited alongside, same era.
Later among the works it cites.
Query-policy misalignment in preference-based reinforcement learning
Hu, X., Li, J., Zhan, X., Jia, Q.-S., and Zhang, Y.-Q · 2023
Later among the works it cites.
Language-driven representation learning for robotics
Karamcheti, S., Nair, S., Chen, A. S., Kollar, T., Finn, C., Sadigh, D., and Liang, P · 2023
Later among the works it cites.
Mind the gap: Offline policy optimization for imperfect rewards
Li, J., Hu, X., Xu, H., Liu, J., Zhan, X., Jia, Q.-S., and Zhang, Y.-Q · 2023
Later among the works it cites.
Mixmae: Mixed and masked autoencoder for efficient pretraining of hierarchical vision transformers
Liu, J., Huang, X., Zheng, J., Liu, Y., and Li, H · 2023
Later among the works it cites.
Structured world models from human videos
Mendonca, R., Bahl, S., and Pathak, D · 2023
Later among the works it cites.
Goal representations for instruction following: A semi-supervised language interface to control
Myers, V., He, A. W., Fang, K., Walke, H. R., Hansen-Estruch, P., Cheng, C.-A., Jalobeanu, M., Kolobov, A., Dragan, A., and Levine, S · 2023
Later among the works it cites.
R3m: A universal visual representation for robot manipulation
Nair, S., Rajeswaran, A., Kumar, V., Finn, C., and Gupta, A · 2023
Later among the works it cites.
Real-world robot learning with masked visual pre-training
Radosavovic, I., Xiao, T., James, S., Abbeel, P., Malik, J., and Darrell, T · 2023
Later among the works it cites.
Mutex: Learning unified policies from multimodal task specifications
Shah, R., Martín-Martín, R., and Zhu, Y · 2023
Later among the works it cites.
Perceiver-actor: A multi-task transformer for robotic manipulation
Shridhar, M., Manuelli, L., and Fox, D · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Later among the works it cites.
Bridgedata v2: A dataset for robot learning at scale
Walke, H. R., Black, K., Zhao, T. Z., Vuong, Q., Zheng, C., Hansen-Estruch, P., He, A. W., Myers, V., Kim, M. J., Du, M., et al · 2023
Later among the works it cites.
Openchat: Advancing open-source language models with mixed-quality data
Wang, G., Cheng, S., Zhan, X., Li, X., Song, S., and Liu, Y · 2023
Later among the works it cites.
A closer look at video sampling for sequential action recognition
Zhang, Y., Zhao, J., Chen, Z., Mi, S., Zhu, H., and Geng, X · 2023
Later among the works it cites.
Rt-2: Vision-language-action models transfer web knowledge to robotic control
Zitkovich, B., Yu, T., Xu, S., Xu, P., Xiao, T., Xia, F., Wu, J., Wohlhart, P., Welker, S., Wahid, A., et al · 2023
Later among the works it cites.