Fetching the paper…
Reading the bibliography…
Visual robotic manipulation research and applications often use multiple cameras, or views, to better perceive the world.
Fusing monocular information in multicamera slam
Sola, J., Monin, A., Devy, M., and Vidal-Calleja, T · 2008
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Slam-based automatic extrinsic calibration of a multi-camera rig
Carrera, G., Angeli, A., and Davison, A. J · 2011
Earlier work this paper cites.
Auto-encoding variational bayes
Kingma, D. P. and Welling, M · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Facenet: A unified embedding for face recognition and clustering
Schroff, F., Kalenichenko, D., and Philbin, J · 2015
Earlier work this paper cites.
High-dimensional continuous control using generalized advantage estimation
Schulman, J., Moritz, P., Levine, S., Jordan, M., and Abbeel, P · 2015
Earlier work this paper cites.
Embed to control: A locally linear latent dynamics model for control from raw images
Watter, M., Springenberg, J., Boedecker, J., and Riedmiller, M · 2015
Earlier work this paper cites.
Representation learning with contrastive predictive coding
Oord, A. v. d., Li, Y., and Vinyals, O · 2018
Earlier work this paper cites.
Sim2real viewpoint invariant visual servoing by recurrent control
Sadeghi, F., Toshev, A., Jang, E., and Levine, S · 2018
Earlier work this paper cites.
Time-contrastive networks: Self-supervised learning from video
Sermanet, P., Lynch, C., Chebotar, Y., Hsu, J., Jang, E., Schaal, S., and Levine, S · 2018
Earlier work this paper cites.
Solving rubik’s cube with a robot hand
Akkaya, I., Andrychowicz, M., Chociej, M., Litwin, M., McGrew, B., Petron, A., Paino, A., Plappert, M., Powell, G., Ribas, R., et al · 2019
Earlier work this paper cites.
A theoretical analysis of contrastive unsupervised representation learning
Arora, S., Khandeparkar, H., Khodak, M., Plevrakis, O., and Saunshi, N · 2019
Earlier work this paper cites.
Robonet: Large-scale multi-robot learning
Dasari, S., Ebert, F., Tian, S., Nair, S., Bucher, B., Schmeckpeper, K., Singh, S., Levine, S., and Finn, C · 2019
Earlier work this paper cites.
Deepmdp: Learning continuous latent space models for representation learning
Gelada, C., Kumar, S., Buckman, J., Nachum, O., and Bellemare, M. G · 2019
Earlier work this paper cites.
Learning latent dynamics for planning from pixels
Hafner, D., Lillicrap, T., Fischer, I., Villegas, R., Ha, D., Lee, H., and Davidson, J · 2019
Cited alongside, same era.
Learning precise 3d manipulation from multiple uncalibrated cameras
Akinola, I., Varley, J., and Kalashnikov, D · 2020
Cited alongside, same era.
RLBench: The robot learning benchmark & learning environment
James, S., Ma, Z., Arrojo, D. R., and Davison, A. J · 2020
Cited alongside, same era.
A framework for efficient robotic manipulation
Zhan, A., Zhao, P., Pinto, L., Abbeel, P., and Laskin, M · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al · 2021
Cited alongside, same era.
Tip-adapter: Training-free clip-adapter for better vision-language modeling
Zhang, R., Fang, R., Zhang, W., Gao, P., Li, K., Dai, J., Qiao, Y., and Li, H · 2021
Later among the works it cites.
Masked autoencoders as spatiotemporal learners
Feichtenhofer, C., Fan, H., Li, Y., and He, K · 2022
Later among the works it cites.
Multimodal masked autoencoders learn transferable representations
Geng, X., Liu, H., Lee, L., Schuurams, D., Levine, S., and Abbeel, P · 2022
Later among the works it cites.
Instruction-driven history-aware policies for robotic manipulations
Guhur, P.-L., Chen, S., Garcia, R., Tapaswi, M., Laptev, I., and Schmid, C · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Gao, P., Geng, S., Zhang, R., Ma, T., Fang, R., Zhang, Y., Li, H., and Qiao, Y · 2021
Cited alongside, same era.
Mastering atari with discrete world models
Hafner, D., Lillicrap, T., Norouzi, M., and Ba, J · 2021
Cited alongside, same era.
Masked autoencoders are scalable vision learners
He, K., Chen, X., Xie, S., Li, Y., Dollár, P., and Girshick, R · 2021
Cited alongside, same era.
Perceiver: General perception with iterative attention
Jaegle, A., Gimeno, F., Brock, A., Vinyals, O., Zisserman, A., and Carreira, J · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Cited alongside, same era.
Pretraining representations for data-efficient reinforcement learning
Schwarzer, M., Rajkumar, N., Noukhovitch, M., Anand, A., Charlin, L., Hjelm, D., Bachman, P., and Courville, A · 2021
Cited alongside, same era.
Cliport: What and where pathways for robotic manipulation
Shridhar, M., Manuelli, L., and Fox, D · 2021
Cited alongside, same era.
Hsu, K., Kim, M. J., Rafailov, R., Wu, J., and Finn, C · 2022
Later among the works it cites.
Q-attention: Enabling efficient learning for vision-based robotic manipulation
James, S. and Davison, A. J · 2022
Later among the works it cites.
Coarse-to-Fine Q-attention: Efficient learning for visual robotic manipulation via discretisation
James, S., Wada, K., Laidlow, T., and Davison, A. J · 2022
Later among the works it cites.
Look closer: Bridging egocentric and third-person views with transformers for robotic manipulation
Jangir, R., Hansen, N., Ghosal, S., Jain, M., and Wang, X · 2022
Later among the works it cites.
R3m: A universal visual representation for robot manipulation
Nair, S., Rajeswaran, A., Kumar, V., Finn, C., and Gupta, A · 2022
Later among the works it cites.
The unsurprising effectiveness of pre-trained vision models for control
Parisi, S., Rajeswaran, A., Purushwalkam, S., and Gupta, A · 2022
Later among the works it cites.
Real world robot learning with masked visual pre-training
Radosavovic, I., Xiao, T., James, S., Abbeel, P., Malik, J., and Darrell, T · 2022
Later among the works it cites.
Perceiver-actor: A multi-task transformer for robotic manipulation
Shridhar, M., Manuelli, L., and Fox, D · 2022
Later among the works it cites.
Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training
Tong, Z., Song, Y., Wang, J., and Wang, L · 2022
Later among the works it cites.
Masked visual pre-training for motor control
Xiao, T., Radosavovic, I., Darrell, T., and Malik, J · 2022
Later among the works it cites.
Lossless adaptation of pretrained vision models for robotic manipulation
Sharma, M., Fantacci, C., Zhou, Y., Koppula, S., Heess, N., Scholz, J., and Aytar, Y · 2023
Closest in time.